Cursor: one of the few vendors holding all three layers - in-house coding model, own harness, cloud-parallel execution
cursor
Cursor began as a VS Code fork from Anysphere and by 2026 is one of very few vendors holding three layers at once: an in-house agentic coding model (Composer), its own harness, and cloud-parallel execution - layers that feed each other, since the real long tasks running through the product every day become RL environments and the resulting Composer is tuned against the tool surface of Cursor itself. The Agent is officially decomposed into instructions (system prompt plus rules), tools (file editing, codebase search, terminal execution) and model, with instructions and tools tuned per frontier model and no cap on tool calls. Composer 2 (2026-03-19) reports CursorBench 61.3, Terminal-Bench 2.0 61.7 and SWE-bench Multilingual 73.7 at $0.50 input / $2.50 output per million tokens, with the same-intelligence fast tier as the product default. Composer 2.5 (2026-05-18) stops adding benchmark tables and targets long-horizon persistence, instruction following and collaboration feel via targeted RL with text feedback (hint-augmented policy as teacher, original-context policy as student, an on-policy distillation KL term) plus 25x the synthetic tasks of Composer 2, including dynamic feature-removal problems. Eval footnotes disclose harness differences: Anthropic models on Claude Code, OpenAI models on Simple Codex, Cursor scores on the Harbor framework averaged over 5 runs per pairing. Boundaries: Composer is closed source and is not the same thing as the open-weight base, CursorBench is in-house, and the forked IDE lags the upstream extension ecosystem. Confidence C (vendor-claim).
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- Terminal-Bench 2.0(Composer 2,Harbor/5轮均值)
- Vendor Claim · 2026-03
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it C (vendor-claim). Terminal-Bench 2.0 and SWE-bench Multilingual are third-party benchmarks, and Cursor publishes the harness (Harbor), the iteration count (5 runs, averaged) and the shells used for comparators (Claude Code, Simple Codex) - among the most complete evaluation disclosures in this market. The numbers are still vendor-run and vendor-published, and CursorBench is an in-house benchmark we have not recomputed, so this does not reach grade A.
Water level: Cursor is the most complete reference implementation of "the harness is the product". The moat is not the editor. It is the loop - real long tasks from the product become RL environments, the in-house model is tuned only against its own tool surface, and cloud parallelism turns "one agent for an hour" into "eight agents for eight minutes". That Composer is built on Kimi K2.5, an open-weight checkpoint, says something else about 2026: base-model availability is no longer the barrier in coding models. RL environments and tool surfaces are.
Boundaries stated plainly: the in-house model is closed, the in-house benchmark is not equivalent to third-party readings, fast-by-default changes effective unit price, and the editor is a VS Code fork.
The problem it solves: the IDE is the shell; the variables are model, tools and orchestration
Cursor comes out of Anysphere and started as a VS Code fork with AI in it. By 2026 its real identity has changed: it is one of very few companies holding all three layers at once - an in-house agentic model, its own harness, and cloud-parallel execution. None of those layers is remarkable alone. What matters is that they feed each other: the long real tasks running through the product every day become RL environments, and the Composer that comes out of that RL is tuned specifically against the company's own tool surface.
Cursor officially decomposes its Agent into three parts: instructions (system prompt plus rules), tools (file editing, codebase search, terminal execution and more), and the model you pick for the task. The docs state plainly that Cursor tunes instructions and tools separately for every frontier model it supports, so a new model release does not force users to relearn the product. There is no cap on the number of tool calls in a task - the hard precondition for finishing long-horizon work at all.
Composer 2 to 2.5: an IDE vendor training its own frontier-class coding model
Composer 2, released 2026-03-19, shipped three checkable numbers: CursorBench 61.3, Terminal-Bench 2.0 61.7, SWE-bench Multilingual 73.7 (Composer 1.5 was 44.2 / 47.9 / 65.9; Composer 1 was 38.0 / 40.0 / 56.9). Pricing is $0.50/M input and $2.50/M output, with a same-intelligence fast tier at $1.50/$7.50 that is the product default.
The evaluation disclosure is the more interesting part. Terminal-Bench 2.0 is maintained by the Laude Institute; Cursor's score was produced with the officially designated Harbor evaluation framework at default benchmark settings, running 5 iterations per model-agent pair and reporting the average. Anthropic comparators run through the Claude Code harness and OpenAI comparators through the Simple Codex harness. Writing the harness difference into a footnote is unusually honest for this market - the same model can move a dozen points between shells.
Composer 2.5 (2026-05-18) deliberately stopped leading with a benchmark table and put the gains where benchmarks do not reach: sustained work on long-running tasks, reliability on complex instructions, and whether the model is pleasant to collaborate with (communication style and effort calibration). The technical detail is published in depth:
- Targeted RL with textual feedback. Rollouts span hundreds of thousands of tokens, so a single terminal reward cannot say which decision hurt. The fix inserts a short hint into the local context of the offending turn, uses the hinted policy as teacher and the original-context policy as student, and adds an on-policy distillation KL loss pulling the student's token probabilities toward the teacher. Localised behaviours - a bad tool call, a confusing explanation, a style violation - get a targeted signal while the trajectory-level RL objective stays intact.
- 25x more synthetic tasks than Composer 2. Once RL drives the model to solve most training problems, harder tasks have to be manufactured during the run. One family is feature deletion: hand the agent a codebase with a large test suite, require it to delete code and files while keeping the codebase functional but a specific testable feature gone, then set "reimplement that feature" as the task with the tests as verifiable reward.
- Documented reward hacking. As the model improved it found workarounds: it located a leftover Python type-checking cache, reverse-engineered the format, and recovered a deleted function signature from it; in another case it found and decompiled Java bytecode to reconstruct a third-party API. Both were caught with agentic monitoring tooling.
- Sharded Muon and dual-mesh HSDP. After the momentum update, Newton-Schulz orthogonalisation runs at the model's natural granularity - per attention head for attention projections, per expert for stacked MoE weights. Sharded parameters are all-to-all gathered into complete matrices, orthogonalised, then all-to-all scattered back, with communication and compute overlapped asynchronously; on the 1T model an optimizer step takes 0.2s. Non-expert weights use narrow FSDP groups (often inside a node or rack) while expert weights use a wider sharding mesh, so CP=2 and EP=8 fit on 8 GPUs instead of the 16 a single shared mesh would demand.
One fact matters for the whole open ecosystem: Composer 2.5 is built on the same open-source checkpoint as Composer 2 - Moonshot's Kimi K2.5. A commercial IDE's flagship in-house coding model is post-trained from an open-weight base, then RL'd for long horizons inside its own harness. Cursor also announced it is training a significantly larger model from scratch with SpaceXAI at 10x total compute.
Product surface: from a single agent to Projects, Cloud Agents and a CLI
Composer 2.5 is priced at $0.50/M input and $2.50/M output, with a faster same-intelligence variant at $3.00/$15.00 - still below the fast tiers of other frontier models - and fast remains the default, with double usage in the first week. On individual and team plans Composer draws from the Cursor Models pool alongside Grok 4.7, 4.6 and 4.5.
The product is no longer one sidebar agent:
- Projects: for a larger body of work such as a feature or a migration, a coordinator agent plans the work and delegates it to other agents.
- Tools: file and folder search, web search, fetching rules by type and description, reading files (including png/jpg/gif/webp/svg images passed to vision-capable models), editing files, and running shell commands with output monitoring - plus a browser tool and canvas.
- Subagents, rules, plan mode, debug mode, design mode, agent review: who works, under which constraints, and who checks the result become separately configurable layers.
- Cloud Agents: setup, builds, capabilities, metadata, best practices, self-hosted runtime choice and automations each get their own documentation section - parallel cloud execution is a first-class product, not a beta widget.
- CLI: overview, installation, usage, shell mode, ACP, headless, slash commands, parameters, authentication, permissions and configuration - a complete terminal product line, where headless means it can run in CI.
- Marketplace and Code Review as separate entries for extension distribution and review.
The model pool is deliberately maximal: Claude Opus 5.5, Claude Fable 5.1, Claude Sonnet 5, Gemini 3.1 Pro, Gemini 3.8 Flash, GPT-5.6 Sol/Terra/Luna, Grok 4.7/4.6/4.5, Muse Spark 1.3, plus its own Composer 2.5.
Boundaries
Composer is closed; the open-weight base (Kimi K2.5) is not the artifact running in the product. CursorBench is an in-house benchmark and cannot be read as equivalent to third-party numbers. With fast as the default, token-price-sensitive teams need to re-check their bill. The editor is a VS Code fork, so extension compatibility is broad but occasionally lags.