Qoder: Alibaba agentic platform for real work - nine product lines on one knowledge engine
qoder
Qoder is the agentic coding platform from Alibaba, positioned officially as an agentic platform for real work: not an editor but an end-to-end loop - understand the task and context, plan, call tools, verify results, iterate toward the deliverable - resting on three stated principles (context engineering, agent autonomy, goal-directed loops), with nine product lines sharing one knowledge engine (desktop Qoder and Qoder IDE coexisting rather than replacing each other, Editor and Quest forms, a JetBrains plugin, Qoder CLI, cloud agents and more); session history and memory are stored separately but can be imported from the IDE. Repo Wiki is generated locally by multiple agents, never uploads the codebase, is off by default and supports Auto Update, Auto Export and Citation back to source locations. Quest has four drives - Agent, Experts, Goal and Spec (convertible to scheduled tasks): Spec runs requirement clarification (multiple choice, with Recommend / Continue / Skip), a structured Spec covering requirements, design, task breakdown and acceptance criteria, human review, execution, then Review/Commit/Push, while Goal takes only the desired outcome and evaluates progress at the end of every round, continuing automatically until met. Two scaled cases: building Qoder with Qoder (10 people, 3 weeks, 500,000 lines of agent code merged into a 4-million-line legacy system, 99% agent-generated, still in production at v1.4.0 with zero incidents; the method is a cognitive base plus Ultra Spec plus Experts cross-review plus a verifier agent filtering hallucinated issues, with humans only deciding SLO definitions and irreversible operations, and each person driving 20-plus Experts tasks a day); and AutoSDK for AMap in-car systems across 20-plus repositories and over a million lines, where the strict first-pass rate went from 37.3% to 61.5% (problem framing cites KoCo-Bench / arXiv:2601.13240v3: general coding reaches 90% Pass@1 while domain code generation reaches only 8.9%). Boundaries: the client and knowledge engine are closed; every scaled number comes from official cases and vendor self-reporting, and self-evidence from a product about itself carries methodological self-interest, none of it independently reproduced; Experts cost per unit is clearly above a single agent (median about 75 versus about 50 Credits) and Credits reset each cycle rather than accumulating. Confidence C (vendor-claim).
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- 10 人 3 周并入 400 万行遗留系统(厂商案例)
- Vendor Claim · 2026
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it C (vendor-claim). A fair amount is independently checkable: nine product lines each with published documentation, the Cloud Agents four-primitive model and its TypeScript/Python/Go SDKs in a public repo (
github.com/QoderAI/cloud-agents), plan tiers plus a median consumption table on the pricing page, Repo Wiki's verifiable behaviour (generated locally, no codebase upload, off by default), and the/better-harnessreview loop with its four questions in public docs. But the reading we adopt as the metric - a 10-person team merging 500,000 lines of agent-written code into a 4-million-line legacy system in three weeks, 99% agent-generated, with no production incidents - comes from an official case study whose subject is the vendor's own team building Qoder with Qoder. It is simultaneously vendor-reported and self-proving, with no third-party audit and no reproducible definition (what counts as a line, which changes are included, how an incident is defined are all unstated), and we have not reproduced it. Per our contract that lands at vendor-claim. The AMAP AutoSDK 37.3% to 61.5% figure is the same case: only the KoCo-Bench measurement it cites (arXiv:2601.13240v3 - 90% Pass@1 on general programming vs 8.9% on domain code generation) is third-party and checkable.Water level: Qoder's strategic position is not the editor. It is the only vendor treating harness governance as the product's main line. Three pieces of evidence interlock. First, it refuses to bet reliability on longer context - the case study states outright that even with a 1M-token window, attention dilution and context decay still occur and a larger window does not mean every token receives equal attention, so consistency is preserved by Experts decomposition, independent contexts and a shared knowledge engine instead. That is a public falsification of the long-context myth. Second,
/better-harnessturns "audit your own engineering environment" into a slash command, conceding that the bottleneck has moved from model to harness, and it requires the fix to be the smallest durable one (explicitly arguing against hardening transient failures into permanent rules). Third, the knowledge engine is treated as an engineering asset that can be produced, tuned, refreshed and consumed rather than a one-off index; AMAP's core conclusion - the ceiling on AI coding is domain knowledge, not the model - and the 90% vs 8.9% KoCo-Bench gap are two independent expressions of the same judgement. Read together, Qoder is betting on organization-level cognitive infrastructure, not single-developer productivity.The directly transferable lessons: an Experts Team Lead that does not write code but decomposes work into a DAG while refusing to co-schedule tasks that touch the same file is a scheduling constraint worth copying; "spend the first two days refreshing the knowledge layer, highest ROI of the whole three weeks" is the most valuable empirical data point for any team about to run parallel agents over an existing codebase; and the median cost table (Editor Agent ~12, Quest Agent ~50, Experts ~75 Credits) is a rare cost baseline you can align against.
Boundaries stated plainly: the client and knowledge engine are closed; every scale number is vendor-reported and partly self-proving; Experts costs roughly 50% more per request than a single agent, so parallelism is bought with quota; Credits expire each cycle and cannot be stockpiled; and QoderWake's "digital employee" is closer to long-running responsibility orchestration today than to unattended delivery of production changes.
What it is: Alibaba's "agentic platform for real work" - a product family, not an editor
Qoder is Alibaba's agentic coding platform, and the official positioning is deliberately broad: an agentic platform for real work, bringing AI into software development, terminal workflows, managed cloud execution, everyday productivity and long-running digital roles. The design starting point is not "completion plus chat" but an end-to-end loop: understand the task and its context, plan the work, use tools to execute, verify the result, iterate toward the deliverable. Three foundations are called out explicitly: context engineering (give agents persistent context from code, knowledge, rules, tools, files and the working environment), agent autonomy (decide and execute multi-step work independently while retaining necessary review points), and a goal-oriented loop (from a stated goal through planning, execution, verification and iteration until the deliverable is ready).
The family currently has nine lines, and this is not marketing-speak: each has its own documentation and its own form factor.
- Qoder (desktop): grew out of Quest mode in Qoder IDE and is now an independent desktop product. The docs state plainly that the two coexist and do not replace each other - use the IDE if you prefer writing and debugging code directly, use Qoder if you prefer delegating an end-to-end task. Conversation history and memory are stored separately per product, and you can import IDE history and memories on first run.
- Qoder IDE: a dedicated agentic development workspace split into Editor (in-flow assistance) and Quest (long-running, multi-step delegation).
- JetBrains Plugin: code suggestions, Ask, Agent, MCP and project rules inside supported JetBrains IDEs.
- Qoder CLI: a coding agent in the terminal, extendable into scripts and automation.
- Cloud Agents: a managed agent API surface built on four primitives - Agent / Environment / Session / Event. Configure agents and environments, start sessions, stream results. SDKs in TypeScript, Python and Go (
github.com/QoderAI/cloud-agents), with separate data-ownership, access and billing rules for personal vs enterprise spaces. - QoderWork: delegate document, spreadsheet, research, browser and desktop tasks and receive usable local deliverables.
- QoderWake: create digital employees called Wakers for ongoing responsibilities, conversations, automations and multi-stage flows - Groups with a Leader, Autonomous Work, WakerFlow, and
@Wakerdispatch from IM. - Mobile & Web: monitor IDE and CLI tasks, review plans and handle approvals away from your machine.
- Enterprise: centralized purchasing, members and identity, policies, knowledge, models, marketplace and audit.
The knowledge engine: Repo Wiki + code knowledge graph + Memory is the actual foundation
The most distinctive thing about Qoder is that it treats "make the agent understand an existing repository" as a first-principles problem rather than solving it by stretching the context window.
Repo Wiki generates repository-level documentation automatically from code, comments, documents and commit history; the code knowledge graph models module dependencies, interface contracts and data flow; Memory and Knowledge Cards turn past runs into reusable assets. The official case study calls these three layers the cognitive foundation and makes a hard claim: when dozens of agents write code in parallel, every knowledge gap gets amplified. In a fixed three-week schedule they spent the first two days entirely on refreshing the knowledge layer (full Repo Wiki refresh, a deep knowledge-graph update targeting modules added in the previous month and recent interface changes, plus consolidating historical Specs) - and in retrospect call those two days the highest-ROI part of the whole three weeks.
The engineering boundaries are stated too: Repo Wiki is generated locally by multiple agents without uploading the codebase, is off by default, and supports Auto Update, Auto Export and Citation back to source locations.
Four ways to drive Quest: Agent / Experts / Goal / Spec (schedulable)
Spec-driven: turn on the Spec toggle in the Quest input "+" - requirement clarification (multiple-choice, with Recommend / Continue / Skip), then a structured Spec (requirement description, design plan, task breakdown, acceptance criteria), human review, execution against the Spec, then Review / Commit / Push. A Spec can be converted into a scheduled task or into a Goal.
Goal-driven: describe only the outcome you want; Quest decomposes it, plans the path and keeps iterating. After each round it evaluates progress and judges whether the goal is met, and if not it moves on to the next iteration automatically. Control stays available: pause, edit the goal or delete the task mid-run, pausing preserves full context and intermediate artifacts, and editing the goal continues from the next round against the new goal without resetting existing progress. Documented fits are tasks with clear success criteria and a verification loop: raising test coverage, performance tuning, cross-file consistency fixes, whole-repo framework upgrades, batch API replacements.
Experts is the most recognizable layer. A Team Lead expert does not write code: it reads the Spec's task list, identifies dependencies, breaks the work into a DAG and delegates - dependent tasks do not run in parallel, and tasks likely to edit the same file do not run simultaneously. Backend, frontend, testing and review experts each take assignments in their own independent context. Each expert can be pinned to its own model, given a dedicated prompt of up to 10,000 characters, and attached to its own Skills and MCP servers, with an Expert Team Canvas for the global view.
The case study offers an unusually honest reason for splitting: even though Qoder IDE supports a 1M-token context window, attention dilution and context decay still occur as context grows - a larger window does not mean every token receives equal attention. One agent running a task end to end forgets earlier agreements in the second half. So consistency is guaranteed by the shared knowledge engine and the same original decision records, not by brute-forcing one very long context.
/better-harness: treating the harness itself as an auditable object
This is the feature we think peers should copy. The docs first define an Agent Harness: an agent works in a loop of understand, act, inspect, adjust, and a reliable loop needs more than a capable model - clear project context, usable tools, operating boundaries, validation methods and a way to retain lessons. Those supporting mechanisms are the harness: repository instructions, rules, skills, hooks, plugins, connectors, scripts, test commands, release checks, human review steps. Its purpose is to let the agent answer four practical questions: what outcome is expected and what is out of scope; how should the project be operated and changed; what evidence proves the result is correct; what should happen when an operation or validation fails.
Harness gaps are invisible in a single successful task and surface over time: the same instruction repeated, conventions never written down, validation skipped, review feedback never reaching the next task. /better-harness audits exactly those recurring patterns through a loop of map the current harness, find the breakpoint, choose the smallest durable fix, verify the improvement. Every finding carries its evidence, the affected part of the workflow, the proposed durable fix (rule / skill / hook / script / human approval step) and how the improvement will be verified later. The docs also include two explicit anti-patterns: do not turn transient failures or one-off preferences into permanent project rules, and treat the report as a decision aid to confirm against your repository and team process before repairing.
Shipping "audit your own engineering environment" as a slash command means Qoder believes the reliability bottleneck has moved from the model to the harness. That matches what we see elsewhere (Anthropic's skills, OpenAI's AGENTS.md, Kimi's tool-layer guardrails), but Qoder is the only one that productized the audit itself.
Public case studies: two scale numbers you can actually check
- Building Qoder with Qoder: 10 people, 3 weeks, 500,000 lines of agent-written code merged into a 4-million-line legacy system. V1.0 had to add four major modules (a standalone Quest view, the knowledge engine, multi-workspace parallelism, Experts collaboration) on a fixed three-week deadline. The frontend is a VS Code extension; the backend is Go services for agent orchestration, the knowledge engine and multi-model invocation across two repositories, grown from v0.1 to roughly four million lines in nine months. The method is a three-part stack: cognitive foundation + Ultra Spec + Experts collaboration. An Ultra Spec differs from an ordinary Spec in that multiple agents first research broadly and deeply in parallel, then merge and converge into one directly executable plan; the finished Spec is then cross-reviewed by multiple agents - an architect on module boundaries, a security expert on permission holes, a performance expert on bottlenecks, a legacy-system expert wired into the knowledge graph on compatibility - and a verifier agent works backward to filter out false issues created by hallucination. Humans made only the final calls, focused on SLO definitions and irreversible operations. During execution three groups ran in parallel, each person leading 20-plus Experts tasks per day and thousands of subtasks over three weeks; the vendor states overall execution efficiency was an order of magnitude higher than single-agent mode. All 500,000 lines reached the main branch, 99% generated by agents, and by v1.4.0 they were still in production with no online incidents.
- AMAP (Gaode) automotive AutoSDK: strict one-shot pass rate from 37.3% to 61.5%. AutoSDK spans more than a million lines across twenty-odd Git repositories. The case first frames the problem with KoCo-Bench measurements: general-purpose programming clears 90% Pass@1 while domain code generation reaches only 8.9%; adding domain-knowledge retrieval under an agent paradigm lifts that to 34.2%, peaking at 62.5% on the best track (arXiv:2601.13240v3). One-shot failures were classified into four types - exploration drift, generation deviation, architecture violation, constraint omission - and root-cause analysis converged on one place: business-term meanings, module responsibility boundaries and the trade-offs behind historical decisions had never been structured into a form AI could consume. The knowledge engine was then used to build a production / tuning / refresh / consumption loop so the same class of error does not happen twice.
Two smaller public cases: retail chain Kidswant compressed a week-long full-stack requirement into one day and built "Zhishu", an AI platform running 400 business scenarios and 300 agents; an individual developer reports delivering in two months an MES product that previously took 20 people three months.
Extension and security surface
Extension points cover the four current standards - Skills / Plugins / Connectors / Hooks - plus subagents, AppShot and Computer control (desktop and browser operation). On security there are three escalating levels of code safety checking, Static Check / Lightweight Scan / Deep Scan, with a remediation flow. Desktop tasks split into Coding (grouped by workspace, with execution-mode and branch controls) and General (grouped by folder), and Automations start agent work on a schedule with every run reviewable.
Commercial terms: they actually publish a median-cost table
Four individual plans: Free (2-week Pro trial, 300 Credits, limited completions and NES, BYOK supported), Pro $20/mo with 2,000 Credits, Pro+ $60/mo with 6,000, Ultra $200/mo with 20,000. Paid-plan quota is stated as equivalent in value to the subscription fee, Credits are valid only within the current subscription period and reset to zero when it ends; running out downgrades you to basic models with a daily limit, and failed requests are not deducted. More useful is the published median consumption table (200K context): Editor Ask ~4, Editor Agent ~12, Quest Agent ~50, Quest Experts ~75, Repo Wiki ~50 per repository (at 50K context: Editor Ask ~3, Editor Agent ~7). Publishing medians instead of just saying "usage-based" is rare in this market.
Boundaries
The client and knowledge engine are not open source, and Repo Wiki / knowledge-graph quality depends on the quality of the repository's own comments, docs and commit history. Every scale number (500,000 lines, 99%, an order of magnitude, 37.3% to 61.5%) comes from official case studies and vendor reporting; "building Qoder with Qoder" is a vendor proving its own product and carries methodological self-interest, and we have not recomputed it. The KoCo-Bench readings come from arXiv:2601.13240v3 - third-party and checkable, but still paper-reported. Experts mode is materially more expensive per request than a single agent (median ~75 vs ~50 Credits), so the parallelism gain is bought with quota and is easy to underestimate when planning capacity per head. Credits expiring each cycle means they cannot be stockpiled as prepaid balance. And QoderWake's "digital employee" form is closer to long-running responsibility orchestration today than to unattended delivery of production changes.