Cursor: one of the few vendors holding all three layers - in-house coding model, own harness, cloud-parallel execution
Cursor began as a VS Code fork from Anysphere and by 2026 is one of very few vendors holding three layers at once: an in-house agentic coding model (Composer), its own harness, and cloud-parallel execution - layers that feed each other, since the real long tasks running through the product every day become RL environments and the resulting Composer is tuned against the tool surface of Cursor itself. The Agent is officially decomposed into instructions (system prompt plus rules), tools (file editing, codebase search, terminal execution) and model, with instructions and tools tuned per frontier model and no cap on tool calls. Composer 2 (2026-03-19) reports CursorBench 61.3, Terminal-Bench 2.0 61.7 and SWE-bench Multilingual 73.7 at $0.50 input / $2.50 output per million tokens, with the same-intelligence fast tier as the product default. Composer 2.5 (2026-05-18) stops adding benchmark tables and targets long-horizon persistence, instruction following and collaboration feel via targeted RL with text feedback (hint-augmented policy as teacher, original-context policy as student, an on-policy distillation KL term) plus 25x the synthetic tasks of Composer 2, including dynamic feature-removal problems. Eval footnotes disclose harness differences: Anthropic models on Claude Code, OpenAI models on Simple Codex, Cursor scores on the Harbor framework averaged over 5 runs per pairing. Boundaries: Composer is closed source and is not the same thing as the open-weight base, CursorBench is in-house, and the forked IDE lags the upstream extension ecosystem. Confidence C (vendor-claim).