Skip to content
←Back to Open Source

OPEN SOURCE DEEP DIVE

Agent HarnessHeadless BrowserBrowser Automation

Moli: A Rust Headless Browser Built for AI Agents

Moli is Lexmount open-source headless browser: a standalone Rust kernel rather than a Chromium wrapper, built for AI agents (9k+ stars, dual Apache-2.0 / MIT). Its core design is on-demand layout and paint — by default it reads page structure without triggering layout or paint; –layout enables real geometry, hit-testing and screenshots, with layout frozen into a one-shot FrozenLayoutTree and paint state discarded after use. CDP, WebDriver Classic and WebDriver BiDi share one kernel, so no separate driver or browser install is needed. Official self-tests: 73 MiB median RSS on a 192-site crawl (headless Chrome: 773 MiB) and 81.88% success over 1,308 comparable Lexbench tasks (Chrome 99.85%). No benchmarks reproduced; recorded as unverified.

lexmount/moli9.1kRustApache-2.07 min read

What it is

Moli is an open-source headless browser written in Rust by Lexmount, built for AI agents rather than for humans. The repository went public in August 2026, has passed nine thousand stars, and ships under a dual Apache 2.0 / MIT license. The point that matters most: it is not a wrapper around Chromium. It is a standalone browser kernel — streaming HTML parsing, a native DOM, script execution, style cascade, networking and painting all live in its own code, with its own ownership and lifecycle rules.

It is usable three ways: a command line for fetching and extraction, a resident automation server, and a target for existing automation clients. All three share the same kernel.

On-demand layout and paint: two cost tiers

The design centers on its own motto — structure first, pixels on demand. Moli splits the question of when a browser should do work. The default is LayoutPolicy::Mock: geometry is deterministic and format-compatible, but no real layout and no paint run. Adding --layout switches to LayoutPolicy::OnDemand, which is when real layout, hit-testing, coordinate input, screenshots and low-frequency screencast frames become available.

The reasoning: most automation requests need page structure, not a continuously rendered visual world. The official mapping:

Agent requestWhat Moli does
Extract HTML/Markdown, query the DOM, run JS, inspect network and storageReads runtime state directly — no layout, no paint
Read an element box, hit-test coordinates, send coordinate inputRuns one layout pass and keeps only the latest frozen layout tree
Capture a screenshotRebuilds from the current DOM and style, replaces the frozen tree, renders one frame, discards it
Poll a screencastCompares generation metadata only; unchanged state emits no frame

Concretely, the first geometry request builds a working tree from the current DOM and style, freezes canonical geometry into an immutable, DOM-independent FrozenLayoutTree, retains only that tree, and discards the working tree, style borrows, layout caches, diagnostics and paint state. A screencast subscription remembers only an opaque visual-state token; paint results are never reused. There is no incrementally maintained layout tree, no damage graph, no retained display list, no GPU compositor, no persistent window.

This cost model suits crawling, retrieval pipelines, evaluation environments and reinforcement-learning workloads — places where one request in ten thousand genuinely needs pixels.

Three protocols, one kernel, no drivers

CDP, WebDriver Classic and WebDriver BiDi share the same kernel and scheduler, so no separate driver and no separate browser installation is required. Existing automation libraries connect straight over CDP. On the command line, moli serve starts a basic server, --layout enables real geometry and screenshot surfaces, and --resource additionally fetches optional image, font, audio, video and media resources.

On the CLI side, moli fetch produces HTML, Markdown, JSON and semantic text trees directly, with selector, script and response waits plus network tracing. Visual output is equally on demand: viewport images, full-document images and paginated documents are only produced when asked for.

It stands on: libcurl for network transport, html5ever for HTML parsing, rusty_v8 bindings to V8 for JavaScript, Stylo from Servo for selectors and computed style, Taffy and Parley for box and text layout, and AnyRender with Vello on the CPU for software rendering. Document and style have a single source of truth: the native DOM and its style integration.

Official benchmarks: read what was measured first

Three sets of numbers follow, all from Moli itself or from a benchmark repo under the same organization. We have not reproduced any of them; treat this tier as unverified. The most legible set is a mixed crawl of 192 real URLs, where a page only counts if meaningful content survives after JavaScript runs.

BrowserUseful pagesSuccess rateMedian timeMedian RSS
Moli10353.6%1.43 s73 MiB
Chrome Headless10152.6%1.43 s773 MiB
Lightpanda8544.3%0.97 s40 MiB
Obscura5729.7%1.30 s39 MiB

The takeaway is not who is faster — all four sit in the same time band — but that the same usefulness costs an order of magnitude less memory. Moli beats headless Chrome on success rate by one percentage point while using roughly a tenth of the resident memory.

The second set covers a single agent workload: CDP ready in 34.85 ms versus 169.37 ms, median episode activity 33.40 ms versus 57.13 ms, peak PSS 102.46 MiB versus 348.82 MiB, and 1 process with 24 threads versus 11 processes with 123 threads. That process count explains where the memory difference comes from: it is one process.

The third set is capability coverage. On 1,308 comparable tasks from the Lexbench headless-browser benchmark, Moli 0.1.1 passed 1,071 tasks — 81.88% — ahead of Kitesurf at 62.08%, Lightpanda at 53.29% and Obscura at 44.88%, with headless Chrome as the reference engine at 99.85%. A separate 557-task resource run put its median CPU time at 100.6 ms and median peak memory at 92 MiB, against Chrome's 687 ms and 697 MiB — roughly 15% of the CPU time and 13% of the memory. One full web-platform-tests run passed 1.612 million tests.

Read together, the claim must be stated precisely. Moli is not more accurate than Chrome — 81.88% against 99.85% is an eighteen-point gap, meaning roughly one task in five still fails, and that figure is Moli's own. What it sells is near-Chrome usefulness at about a tenth of the resource cost. Whether that trade is worth it depends on what a retry costs you.

The boundaries it states honestly

The README carries an explicit section on intentional limits: no GUI browser, no persistent window, no GPU compositor, no retained multi-frame paint architecture; no pursuit of pixel-for-pixel parity with Chrome, and no high-fidelity Canvas, WebGL or media playback; --layout supports software screenshots and raster-backed document export, but not every Chrome screenshot or print mode.

More creditably, it is explicit about failure: unsupported protocol paths return explicit errors, and Moli never pretends that a browser action, event, network observation or visual result occurred. For agents this matters more than any performance number — automation that fails silently produces confidently wrong conclusions, whereas an explicit error can be caught and retried upstream.

Where it sits in this section

A browser is the tool surface of an agent. Elsewhere in this section, Pi answers how a runtime should be layered, DeepSeek Harness answers how tools and interfaces should hang off a plugin system, and Hermes answers how skills should improve themselves. Moli answers one specific branch of that: when an agent reaches for the web, how that branch stays cheap and controllable. It is downstream, not a framework itself — but without it, the layers above pay ten times the memory for the same reach.

It also ships in a form agents can consume directly: the repository carries a skills/ directory, and the README simply says to hand your agent a prompt, letting the agent download the prebuilt binary and complete a first fetch on its own. A bundled WebMCP playground lists and calls a site's native web-model-context-protocol tools from the command line.

The commercial relationship is stated too: Moli is the open-source kernel, Lexmount Browser is the managed cloud runtime and control plane around it, and the project states plainly that the open-source headless browser is fully usable without the latter.

Unverified

We ran no benchmarks and no reproduction. Every figure above comes from the project's own repository and the Lexbench benchmark repo under the same organization. Stars and activity — public for two months, still receiving daily commits — make it a layer worth cataloguing on this toolchain, but its capability tier is recorded as unverified.

Related Projects

NousResearch/hermes-agentMIT

Hermes Agent: the open-source agent that turns self-improvement into a closed learning loop

The self-improving agent from Nous Research (248k+ stars). Its learning loop has five mechanisms: periodic memory nudges, autonomous skill creation after tasks, skills that self-improve in use, FTS5 session search with LLM summarization, and Honcho user modeling. One TUI, seven terminal backends (local/Docker/SSH/Singularity/Modal/Daytona/Vercel Sandbox), six chat platforms from a single gateway, no model lock-in. It doubles as a trajectory data-production apparatus. Self-improvement is falsifiable and no controlled experiment has been run, so this is graded as needing reproduction.

Agent HarnessSelf-Improving AgentsNous Research
Python248k52k956
deepseek-ai/deepseek-harnessMIT

DeepSeek Harness: the open-source agent runtime where everything is a plugin

DeepSeek AI's open-source agent harness (CLI dsh, MIT, 233k+ stars). The kernel provides only plugin-system semantics; tools, surfaces and model access all attach as plugins, on top of Cordis' paradigm for spatiotemporal composability (arXiv:2608.25512). It is in developer preview with an explicit warning about compatibility-breaking changes, and SAFETY.md is required reading before you run it. Not benchmarked yet, so this entry is graded as needing reproduction.

Agent HarnessPlugin ArchitectureDeepSeek
TypeScript234k28k1.0k
earendil-works/piMIT

Pi: a layered agent harness monorepo

Pi Agent Harness (formerly badlogic/pi-mono; 108k+ stars, MIT/TypeScript). Not one CLI but layered packages usable on their own: pi-coding-agent (interactive coding agent), pi-agent-core (runtime with tool calling and state), pi-ai (unified multi-provider LLM API), plus chord composition runtime, pi-telemetry contracts, pi-durable persistence and pi-tui differential terminal rendering. The README states plainly that there is no built-in permission system and offers micro-VM / Docker / sandbox paths instead. Not independently benchmarked; graded needs-reproduction.

Agent HarnessCoding AgentTypeScript
TypeScript109k14k334
JuliusBrussee/cavemanNOASSERTION

Caveman: The Token-Cutting Skill, Proxy and Middleware for Coding Agents

A MIT rule file makes agents talk like cavemen (code and errors never shortened), a local Go proxy compresses what agents read before each call with byte-exact originals recoverable, and middleware brings the same to your own app. JetBrains' 86-task A/B: -8.5% output tokens, quality flat; the repo's pinned 54-run suite: -33.2% input tokens, 18/18 checks passed.

token optimizationAgent HarnessLLM cost
Go108k6.3k244

As an Amazon Associate, we earn from qualifying purchases.