Skip to content
← Tags

#Code Harness (7)

AI CodingDevelopmentTopC

Trae: ByteDance two-form-factor coding line - TraeCode (IDE + SOLO) and TraeWork, plus a stalled MIT-open shell

Trae is the AI coding line from ByteDance, and the official docs split it into two clearly different things. TraeCode is a development tool with AI deeply integrated, in two modes: IDE mode keeps the editor, terminal, debugger, extensions and source control for scenarios that need fine-grained control over code changes and execution, while SOLO mode hands the lead to AI - describe the requirement in natural language, by voice or by uploading local files, and it decomposes the task itself and runs code generation, testing, preview, change summary and deployment, with a task panel, an AI conversation panel and a tool panel (built-in editor, doc viewer, browser) from left to right. TraeWork is the AI-native workspace grown out of the SOLO mode of TraeCode, on web, desktop and mobile, in Work / Code / Design modes, aimed at product managers, data analysts, operations and designers rather than developers only. The long-task ceilings are stated concretely: Max mode extends the context window to 1M, allows up to 200 tool-call rounds per task and reads up to 750 lines per file read to cut down on chunking. bytedance/trae-agent is MIT-licensed but has not been updated for about eight months and cannot stand for the current engineering of Trae. Boundaries: no third-party benchmark readings are published (no SWE-bench or Terminal-Bench self-evidence), so every capability claim comes from feature docs and vendor descriptions; clients are closed source; 1M / 200 rounds / 750 lines are ceilings rather than typical experience and the vendor itself says they raise cost substantially; model availability is doubly limited by region (Seed, MiniMax and GLM unavailable to US users) and by plan tier, so teams across regions do not share one model set; concurrent cloud tasks are a hard tier difference and Free excludes SOLO, so evaluation cannot look at monthly price alone. Confidence C (vendor-claim).

1M · 200 轮Max 模式上限(上下文 / 单任务工具轮次)Vendor Claim · 2026-09
ProductByteDanceSiteRepo
Trae: ByteDance two-form-factor coding line - TraeCode (IDE + SOLO) and TraeWork, plus a stalled MIT-open shell
AI CodingDevelopmentTopC

Qoder: Alibaba agentic platform for real work - nine product lines on one knowledge engine

Qoder is the agentic coding platform from Alibaba, positioned officially as an agentic platform for real work: not an editor but an end-to-end loop - understand the task and context, plan, call tools, verify results, iterate toward the deliverable - resting on three stated principles (context engineering, agent autonomy, goal-directed loops), with nine product lines sharing one knowledge engine (desktop Qoder and Qoder IDE coexisting rather than replacing each other, Editor and Quest forms, a JetBrains plugin, Qoder CLI, cloud agents and more); session history and memory are stored separately but can be imported from the IDE. Repo Wiki is generated locally by multiple agents, never uploads the codebase, is off by default and supports Auto Update, Auto Export and Citation back to source locations. Quest has four drives - Agent, Experts, Goal and Spec (convertible to scheduled tasks): Spec runs requirement clarification (multiple choice, with Recommend / Continue / Skip), a structured Spec covering requirements, design, task breakdown and acceptance criteria, human review, execution, then Review/Commit/Push, while Goal takes only the desired outcome and evaluates progress at the end of every round, continuing automatically until met. Two scaled cases: building Qoder with Qoder (10 people, 3 weeks, 500,000 lines of agent code merged into a 4-million-line legacy system, 99% agent-generated, still in production at v1.4.0 with zero incidents; the method is a cognitive base plus Ultra Spec plus Experts cross-review plus a verifier agent filtering hallucinated issues, with humans only deciding SLO definitions and irreversible operations, and each person driving 20-plus Experts tasks a day); and AutoSDK for AMap in-car systems across 20-plus repositories and over a million lines, where the strict first-pass rate went from 37.3% to 61.5% (problem framing cites KoCo-Bench / arXiv:2601.13240v3: general coding reaches 90% Pass@1 while domain code generation reaches only 8.9%). Boundaries: the client and knowledge engine are closed; every scaled number comes from official cases and vendor self-reporting, and self-evidence from a product about itself carries methodological self-interest, none of it independently reproduced; Experts cost per unit is clearly above a single agent (median about 75 versus about 50 Credits) and Credits reset each cycle rather than accumulating. Confidence C (vendor-claim).

500,000 行 · 99% agent 生成10 人 3 周并入 400 万行遗留系统(厂商案例)Vendor Claim · 2026
ProductAlibabaSiteRepo
Qoder: Alibaba agentic platform for real work - nine product lines on one knowledge engine
AI CodingDevelopmentTopC

Kimi Code: a model lab that builds its own harness, then opens it to every other shell over two protocols

Kimi Code is the developer coding service from Moonshot AI at kimi.com/code, and it sits in an unusual spot in this cohort: most harnesses are shell vendors plugging into models, while Kimi Code is a model lab building its own shell and then opening the API over two protocols at once - OpenAI-compatible at api.kimi.com/coding/v1 (China) and api.kimi.ai/coding/v1 (overseas), Anthropic-compatible at api.kimi.com/coding/ and api.kimi.ai/coding/ - with dedicated integration guides for Claude Code, OpenCode, Codex and Hermes Agent, which amounts to opening its own models to every other shell. Three clients run in parallel: Desktop (released 2026-09-17 for macOS Apple Silicon/Intel and Windows, moving the CLI agent core into a GUI), the CLI (the kimi command, widest feature surface, install.sh verifies checksums; Windows relies on Git Bash from Git for Windows with KIMI_SHELL_PATH for non-standard bash.exe) and a VS Code extension. The model surface is three models across four model IDs - k3, k3-256k, kimi-for-coding, kimi-for-coding-highspeed - and the docs insist on the model ID rather than a version name (writing K3 or K2.8 Preview fails outright), since a misspelled HighSpeed ID falls back silently instead of erroring. K2.7 Code HighSpeed runs about 180 tokens/s (up to 260 on short context) at 5-6x speed for 3x quota, and it only accelerates model output, so rounds dominated by tool calls feel little faster. K2.7 Code was released and open-sourced on 2026-06-12 with official deltas over K2.6 of Program-Bench +10.4%, MCP Mark Verified +11.4%, SWE Marathon +76.2% and 30% fewer reasoning tokens. Boundaries: all readings come from the vendor or vendor-relayed external benchmarks and were not independently reproduced; the K3 technical report is not public alongside the weights; quota is tightly bound to membership tiers (the 1M tier costs about twice the 256K tier, HighSpeed three times); older models retire fast (kimi-k2 on 2026-05-25, kimi-latest on 2026-01-28, kimi-k2.5 and moonshot-v1 on 2026-08-31). Confidence C (vendor-claim).

4 个 model ID · 1M 上下文内置模型面(K3 / K2.8 Preview / K2.7 Code HighSpeed)Vendor Claim · 2026-09
ProductMoonshot AISite
Kimi Code: a model lab that builds its own harness, then opens it to every other shell over two protocols
AI CodingDevelopmentTopC

ZCode: the official GLM-5.3 harness, defined as an ADE rather than an editor

ZCode is the official harness Z.ai built for GLM-5.3, and it defines itself as neither an AI editor nor a CLI but an ADE (Agentic Development Environment): it turns the 1M context window and long-horizon capability of GLM-5.3 into a stable desktop experience covering planning, coding, review and iteration, keeping goal, files, terminal output, browser context, execution mode and Git state inside one task so continuity survives from plan through implementation to verification, with model capability, tool calling and the execution chain tuned over multiple rounds against GLM-5.3. The most distinctive design is /goal mode: after a goal is set, every round ends with an automatic check of whether the goal is met, and the agent continues into the next round on its own until completion is confirmed (the documented example is a research task running twenty-plus rounds, with a right-hand panel showing what each round did). One session holds one goal at a time; /goal shows it, /goal sets it (replacing any existing goal), /goal replace swaps it explicitly, /goal pause suspends it, and the docs are honest about the fit - work that is easy to state in one sentence but takes many rounds to finish. The card metric keeps the verifiable price fact layer: GLM-5.3 API at $1.4 input / $4.4 output per million tokens. Boundaries: the client is not open source and the deep tuning only covers the GLM-5.3 family; a stable 1M context is a vendor claim and mid-context recall is an industry-wide weakness, so stuffing a whole repo is no substitute for retrieval; idle-time tasks are being rolled out to subscribers rather than always available; the free quota is a 5-day window, not a lasting benefit. Confidence C (vendor-claim).

$1.4 / $4.4GLM-5.3 API 单价(输入/输出,每百万 token)Vendor Claim · 2026-09
ProductZ.aiSite
ZCode: the official GLM-5.3 harness, defined as an ADE rather than an editor
AI CodingDevelopmentTopA

Claude Code: the agentic coding tool that lives in the terminal, and the reference harness for Anthropic models

Claude Code is the agentic coding tool from Anthropic, and its own definition is deliberately modest: it lives in your terminal, understands your codebase, and executes routine tasks, explains complex code and handles git workflows through natural language. The official docs name three surfaces - terminal, IDE, and tagging @claude on GitHub. anthropics/claude-code stood at 148,438 stars / 24,902 forks on 2026-09-28, among the highest in this class; stated plainly, the repository has no LICENSE file (the GitHub API returns null), so it is publicly readable but not openly licensed - a material difference from the Apache-2.0 core of Codex CLI that must not be blurred during evaluation. Installation moved from npm to native installers (claude.ai/install.sh, install.ps1, brew cask, winget) and npm install -g @anthropic-ai/claude-code is officially deprecated, usually a way to escape the Node version matrix and global-permission support burden. The plugins/ directory is not a shell: it ships 13 officially maintained capability packs - code-review, pr-review-toolkit, commit-commands, feature-dev, frontend-design, security-guidance, agent-sdk-dev, plugin-dev, hookify, the explanatory and learning output styles, the opus 4.5 migration and ralph-wiggum - each a bundle of custom commands plus agents, turning personal prompts into installable team assets. It is also the reference harness for Anthropic models in third-party benchmarks (the Cursor footnote routes Anthropic model scores through Claude Code). Boundaries: no open license, closed models with unavoidable data egress, a terminal-first form factor that costs non-CLI users a learning curve, and a ceiling that depends on how well any third-party model adapts to the Anthropic protocol and tool calling when swapped in. The card metric is the verifiable fact (stars). Confidence: confirmed.

148,438GitHub stars(仓库无 LICENSE)Confirmed · 2026-09
ProductAnthropicSiteRepo
Claude Code: the agentic coding tool that lives in the terminal, and the reference harness for Anthropic models
AI CodingDevelopmentTopC

Codex: an Apache-2.0 core, one agent in four form factors, and the reference harness for OpenAI models

Codex is the coding agent from OpenAI, and in 2026 it is not one product but four entry points onto one core: the CLI in your terminal, an IDE extension (officially for VS Code, Cursor and Windsurf), a desktop app launched with a single codex app command, and Codex Web running tasks remotely at chatgpt.com/codex. The openai/codex repository is Apache-2.0 and stood at 126,918 stars / 19,819 forks when fetched on 2026-09-28 - one of the few frontier-vendor coding harnesses whose harness is itself open source, so context assembly, tool ordering and approval gating can all be read. Safety is a first-class citizen: docs/ carries sandbox.md, execpolicy.md, exec.md, agents_md.md, skills.md and slash_commands.md, where execpolicy turns which commands may run unattended into an auditable, versionable, org-distributable policy instead of model goodwill, exec lets Codex run in CI, scripts and batch jobs rather than only waiting in a terminal, and AGENTS.md has become a de facto cross-tool standard. Four install channels (install.sh, install.ps1, npm, brew) and two billing models: usage included in ChatGPT plans versus per-token API keys, a cost question to settle before choosing. Codex is also the reference harness for OpenAI models - the Cursor Terminal-Bench 2.0 footnote states that OpenAI model scores use the Simple Codex harness, and the vendor-published Terminal-Bench 4.0 reading of 58.18% is already carried on the GPT-6 Astra card because it measures the model rather than the shell, so this entry does not claim it. Boundaries: models are closed and data egress cannot be avoided; developers.openai.com returns 403 to automated fetching, so the fact layer rests on the repository README/docs and the GitHub API. The card metric keeps only the verifiable fact layer (license plus stars). Confidence C (vendor-claim).

Apache-2.0 · 126,918★Harness 内核许可 + GitHub starsVendor Claim · 2026-09
ProductOpenAISiteRepo
Codex: an Apache-2.0 core, one agent in four form factors, and the reference harness for OpenAI models
AI CodingDevelopmentTopC

Cursor: one of the few vendors holding all three layers - in-house coding model, own harness, cloud-parallel execution

Cursor began as a VS Code fork from Anysphere and by 2026 is one of very few vendors holding three layers at once: an in-house agentic coding model (Composer), its own harness, and cloud-parallel execution - layers that feed each other, since the real long tasks running through the product every day become RL environments and the resulting Composer is tuned against the tool surface of Cursor itself. The Agent is officially decomposed into instructions (system prompt plus rules), tools (file editing, codebase search, terminal execution) and model, with instructions and tools tuned per frontier model and no cap on tool calls. Composer 2 (2026-03-19) reports CursorBench 61.3, Terminal-Bench 2.0 61.7 and SWE-bench Multilingual 73.7 at $0.50 input / $2.50 output per million tokens, with the same-intelligence fast tier as the product default. Composer 2.5 (2026-05-18) stops adding benchmark tables and targets long-horizon persistence, instruction following and collaboration feel via targeted RL with text feedback (hint-augmented policy as teacher, original-context policy as student, an on-policy distillation KL term) plus 25x the synthetic tasks of Composer 2, including dynamic feature-removal problems. Eval footnotes disclose harness differences: Anthropic models on Claude Code, OpenAI models on Simple Codex, Cursor scores on the Harbor framework averaged over 5 runs per pairing. Boundaries: Composer is closed source and is not the same thing as the open-weight base, CursorBench is in-house, and the forked IDE lags the upstream extension ecosystem. Confidence C (vendor-claim).

61.7Terminal-Bench 2.0(Composer 2,Harbor/5轮均值)Vendor Claim · 2026-03
ProductAnysphereSite
Cursor: one of the few vendors holding all three layers - in-house coding model, own harness, cloud-parallel execution