Kimi Code: a model lab that builds its own harness, then opens it to every other shell over two protocols
kimi-code
Kimi Code is the developer coding service from Moonshot AI at kimi.com/code, and it sits in an unusual spot in this cohort: most harnesses are shell vendors plugging into models, while Kimi Code is a model lab building its own shell and then opening the API over two protocols at once - OpenAI-compatible at api.kimi.com/coding/v1 (China) and api.kimi.ai/coding/v1 (overseas), Anthropic-compatible at api.kimi.com/coding/ and api.kimi.ai/coding/ - with dedicated integration guides for Claude Code, OpenCode, Codex and Hermes Agent, which amounts to opening its own models to every other shell. Three clients run in parallel: Desktop (released 2026-09-17 for macOS Apple Silicon/Intel and Windows, moving the CLI agent core into a GUI), the CLI (the kimi command, widest feature surface, install.sh verifies checksums; Windows relies on Git Bash from Git for Windows with KIMI_SHELL_PATH for non-standard bash.exe) and a VS Code extension. The model surface is three models across four model IDs - k3, k3-256k, kimi-for-coding, kimi-for-coding-highspeed - and the docs insist on the model ID rather than a version name (writing K3 or K2.8 Preview fails outright), since a misspelled HighSpeed ID falls back silently instead of erroring. K2.7 Code HighSpeed runs about 180 tokens/s (up to 260 on short context) at 5-6x speed for 3x quota, and it only accelerates model output, so rounds dominated by tool calls feel little faster. K2.7 Code was released and open-sourced on 2026-06-12 with official deltas over K2.6 of Program-Bench +10.4%, MCP Mark Verified +11.4%, SWE Marathon +76.2% and 30% fewer reasoning tokens. Boundaries: all readings come from the vendor or vendor-relayed external benchmarks and were not independently reproduced; the K3 technical report is not public alongside the weights; quota is tightly bound to membership tiers (the 1M tier costs about twice the 256K tier, HighSpeed three times); older models retire fast (kimi-k2 on 2026-05-25, kimi-latest on 2026-01-28, kimi-k2.5 and moonshot-v1 on 2026-08-31). Confidence C (vendor-claim).
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- 内置模型面(K3 / K2.8 Preview / K2.7 Code HighSpeed)
- Vendor Claim · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it C (vendor-claim). Plenty here is hard and checkable: K3's weights are open, the 2.8T parameter count and the 16-of-896 expert ratio are written into official docs, the four model IDs and the two compatible-protocol base URLs can be called and verified directly, the CLI version line is traceable release by release from v0.4.0 (plugin system, 2026-05-27) to v2.1.0 (2026-09-23), and the dangerous-command guard and workspace trust boundaries are observable runtime behaviour. But the reading we adopted as the metric - K2.7 Code's +76.2% on SWE Marathon over K2.6, plus Program-Bench +10.4%, MCP Mark Verified +11.4% and 30% lower reasoning tokens - comes entirely from the official changelog's paraphrase of "external benchmark evaluations". No comparator harness, no iteration count, no evaluation framework is given, we have not recomputed any of it, and K3's technical report still has not been published alongside the weights. Under the taxonomy's confidence rule that is vendor-claim, and open weights do not upgrade it: what is open is the model, not the evaluation.
Water level: Kimi Code's strategic position is not in its clients. It is the only Chinese lab that has simultaneously pulled off three things - its own model, its own harness, and being the model inside other people's harnesses. Three pieces of evidence corroborate each other. Cursor states officially that Composer 2 and Composer 2.5 are built on Kimi K2.5, an open-weight checkpoint, meaning a US commercial IDE's flagship in-house coding model has its base here. Kimi Code itself opens
/codingendpoints over both the OpenAI and Anthropic protocols and writes a dedicated integration guide for Claude Code, OpenCode, Codex and Hermes Agent - deliberately going after the role of model inside someone else's shell. And/import-from-cc-codexreduces migrating from competitors to one slash command. Together these say something specific about 2026: base-model availability has commoditised in coding, and distribution now comes from being directly installable into other people's harnesses.What peers should copy is the defensive detail in the tool layer: Write and Edit require a prior Read and reject on any on-disk change (stale-overwrite guard); Grep keeps filtering
.envand private keys even withinclude_ignored=true; Bash converts a timed-out foreground command to background instead of killing it, with two-phase SIGTERM / 5s / SIGKILL termination; cron applies deterministic jitter and coalesces missed fires with acoalescedCount; Anthropic-compatible providers do not read ambient credentials. None of these are marketing bullets. They are the preconditions for a long-running agent not destroying something, and they are all in public docs.Boundaries stated plainly: every capability reading is vendor-sourced; K3's architecture rests on a blog and docs with no technical report to check against; HighSpeed is described by the vendor itself as resource-limited and fluctuating; quota is bound to membership tiers (the 1M tier costs about double, HighSpeed triple), so capacity planning off API unit prices will be wrong; and model retirement is fast, so workflows pinned to something like
kimi-k2.5need a budgeted migration path.
The problem it solves: a model lab that builds its own harness, then opens it to everyone else's
Kimi Code is Moonshot AI's developer-facing coding service, entry point at kimi.com/code. Its position in this field is unusual. Most harnesses are shell vendors wiring up somebody else's model. Kimi Code is a model lab building its own shell, and simultaneously publishing the API over both the OpenAI and the Anthropic protocol so other people's shells can carry its model. The docs list four base URLs plainly - OpenAI-compatible at https://api.kimi.com/coding/v1 (China) and https://api.kimi.ai/coding/v1 (overseas), Anthropic-compatible at https://api.kimi.com/coding/ and https://api.kimi.ai/coding/ - and ship dedicated integration guides for Claude Code, OpenCode, Codex and Hermes Agent.
Three client surfaces run in parallel: Desktop (generally available 2026-09-17, macOS Apple Silicon / Intel and Windows, carrying the CLI's agent core into a graphical interface), the CLI (the kimi command, with the widest capability surface), and a VS Code extension. The CLI installs via curl -fsSL https://code.kimi.com/kimi-code/install.sh | bash with checksum verification; on Windows it relies on the Git Bash bundled with Git for Windows as its shell environment, and KIMI_SHELL_PATH can point at a non-standard bash.exe.
The model surface: K3 / K2.8 Preview / K2.7 Code HighSpeed across four model IDs
Kimi Code currently offers three models under four model IDs (k3, k3-256k, kimi-for-coding, kimi-for-coding-highspeed). The docs stress filling in the model ID, not the model version name - passing K3 or K2.8 Preview fails the call outright. A misspelled HighSpeed ID, by contrast, silently falls back to the standard tier without an error, which presents as "high speed is on but nothing got faster".
- K3: the flagship, 2.8T parameters, 1M context, native visual understanding, described as the first open-source model in the 3-trillion-parameter class, with weights released. Architecturally it uses Kimi Delta Attention (KDA, a hybrid linear attention) plus Attention Residuals, pushes MoE sparsity further, and activates 16 of 896 experts under the Stable LatentMoE framework; together with training-method and data-recipe changes this yields roughly 2.5x the overall scaling efficiency of K2.
k3-256kis the 256K-context version, and the docs state directly thatk3(1M) consumes about twice the quota ofk3-256k. - K2.8 Preview (model ID
kimi-for-coding): rolled out in place inside Kimi Code on 2026-09-11, ID unchanged, zero configuration changes for clients and third-party tools. Positioned as close to K3 in performance with significantly more efficient thinking, it opens 1M context to every membership tier and supports three thinking-effort levelslow/high/max(defaultmax). With thinking turned off, requests for both the K3 series and K2.8 Preview are served by K2.8 Preview without thinking. - K2.7 Code HighSpeed (
kimi-for-coding-highspeed): the same model as K2.7 Code at roughly 180 tokens/s output, up to 260 tokens/s in short-context scenarios - about 5-6x the speed for 3x the quota. The docs are honest about the limit: HighSpeed only accelerates model output, so tool calls and script execution are unaffected, and when they dominate a turn the felt speedup is small.
K2.7 Code itself shipped and was open-sourced on 2026-06-12 with three published deltas against K2.6: Program-Bench +10.4%, MCP Mark Verified +11.4%, SWE Marathon +76.2%, alongside 30% lower reasoning-token usage (officially, less overthinking). The effort mapping for third-party tools is documented too: ultra/max/xhigh to max, high/medium to high, low/minimum/light to low, none disables thinking, and any other unknown value returns HTTP 400.
The tool surface: built-in tools are part of the runtime, not MCP attachments
Kimi Code CLI publishes its built-in tools in seven groups and draws the distinction from MCP tools explicitly: built-in tools are managed directly by the runtime, their lifecycle is bound to the session, and no external process is required. Approval semantics are uniform - read-only tools (Read, Grep, Glob, ReadMediaFile) are auto-allowed by default, while write and execute tools (Write, Edit, Bash) require approval.
- Files.
Readtakesline_offset(negative counts from the end),column_offset,n_linesandmax_chars, defaulting to 100,000 characters with up to 500,000 requestable; a line too long for one page is returned in fragments together withNext Readarguments, and UTF-16 LE/BE files that fail strict decoding return readable text with a lossy-decoding warning on every page.WriteandEditrequire a prior Read of that file in the session and reject the operation if the file changed on disk since - a hard guard against stale overwrites.Grepruns ripgrep and filters sensitive files such as.envand private keys even wheninclude_ignored=true. - Shell.
Bashforeground defaults to 60s with a 5-minute maximum. A foreground command that hits its timeout is not killed by default; it converts to a background task under the 600s background default, andbash_auto_background_on_timeout=falserestores kill-on-timeout. Stopping uses two-phase termination: SIGTERM, a 5-second grace period, then SIGKILL. stdin is always closed, so interactive commands receive EOF immediately. - Web, plan, state.
WebSearchandFetchURL(which extracts body text from HTML rather than returning the page);EnterPlanMode/ExitPlanModeform a constrained Plan mode whereWriteandEditmay only touch the current plan file andTaskStopis blocked entirely, with the option to present 1-3 alternative approaches at exit;TodoListkeeps a visible subtask list across multi-step work. - Collaboration.
Agent(subagents time out after 2 hours by default),AgentSwarm(a shared prompt template plus an items array, up to 128 subagents, ramping without a cap by default - 5 immediately then one more every 700 ms - withKIMI_CODE_AGENT_SWARM_MAX_CONCURRENCYavailable as a ceiling),AskUserQuestion(1-4 structured multiple-choice questions, optionallybackground=trueso the turn is not blocked),NotifyUser(short mid-turn progress updates into an Updates panel paged with Ctrl-P / Ctrl-N), andSkill(onlytype="inline"skills, maximum nesting depth 3). - Background and scheduled work.
TaskList,TaskOutput(inline preview of the most recent 32 KB, full log on disk with anoutput_path),TaskStopandWaitFor(waits inside the current turn, up to 600s, with Ctrl-S steering to end the wait early);CronCreate,CronListandCronDeletere-inject a prompt into the current session on a 5-field cron schedule, with at most 50 active scheduled tasks per session.
The scheduling detail is worth isolating. To stop every user firing on the hour, the scheduler applies deterministic jitter: recurring tasks shift forward by min(10% of the period, 15 minutes), and one-shot tasks landing exactly on :00 or :30 move forward by up to 90 seconds. If the scheduler misses several fire times (a laptop asleep, say) it fires once on wake, wrapping the prompt in a <cron-fire> envelope carrying a coalescedCount. Recurring tasks alive for more than 7 days fire a final time with stale="true" and are then deleted automatically.
Permissions and long-horizon control: three modes, a dangerous-command guard, Goal and Tower
v0.40.0 on 2026-09-02 renamed the former YOLO / Auto / Manual modes to Never Ask / Ask When Needed / Always Ask and shipped a built-in dangerous-command guard: commands such as shutdown, reboot or rm -rf are blocked outright in Never Ask mode and always require confirmation elsewhere (disable with [permission] dangerous_command_guard=false). The direction of travel is a tightening even as automation increases - irreversible operations get their own hard boundary.
Several long-horizon controls interlock. Goal mode (/goal <objective>) pushes across turns until the objective is met or a decision point needs a human; the goal queue (/goal next) lines up subsequent objectives and picks them up automatically; v0.43.0 changed the Goal time budget to exclude time while the session is closed and removed the 24-hour cap. Tower multi-agent orchestration (/tower <base-branch>) and AgentSwarm are two distinct parallel shapes. /btw opens a side conversation without interrupting the main turn. v0.6.0 simply removed the 1,000-step per-turn cap.
Context management has explicit mechanics: micro compaction is on by default and trims older oversized tool results; a provider 413 context overflow recovers by compacting and retrying first; compaction output is capped at 128k tokens by default to avoid provider max_tokens errors; once accumulated media in a session exceeds 20 MB the oldest images and videos are omitted with a warning; and switching model or reasoning effort invalidates the built-up context cache, so the official advice is to start a new session rather than toggle inside a long one.
The ecosystem surface: Plugins, Skills, Hooks, MCP, ACP - and a migration command
The plugin system launched 2026-05-27 (v0.4.0). /plugins is a tabbed panel: Installed, Official (Kimi-maintained), Third-party, and Custom (GitHub URL, zip or local path, with a confirmation prompt before third-party installs). Two official plugins stand out: Kimi Computer Use (Windows x64 support since 2026-08-06, Windows 10 1903+ / Windows 11) and Kimi WebBridge, which lets the agent drive the user's real browser. Kimi Datasource is the official data plugin spanning five domains - stock market, macroeconomics, corporate registry, academic literature and legal - and on 2026-08-20 it added 13 sources at once: Chinese government data (NDA/NBS) and standards (GB/HB/DB/TT), eight international-organisation datasets (WHO, FAO, UNSD, ECB, Eurostat, UNICEF, OECD, FRED), Xinhua Finance and Caixin.
All four standard extension interfaces are present: MCP (stdio, streamable HTTP and legacy SSE), Agent Skills (with experimental Sub-Skills - sub-skill.review audits existing skills, sub-skill.consolidate merges them into hierarchical groups), Hooks (beyond TurnStarted, UserPromptQueued, TaskStarted and SessionHeartbeat there is a dedicated Interrupt event that fires instead of Stop when the user breaks a turn with Esc, so external tooling no longer mistakes an interrupted turn for a running one), and ACP (the kimi acp subcommand lets Zed, JetBrains AI Chat and other editors drive Kimi's sessions and tool calls directly).
One command says a lot about the competitive landscape: /import-from-cc-codex imports instructions, Skills and MCP settings from Claude Code and Codex in a single step. Migration cost has been reduced to a slash command.
Across devices: Remote Control, kimi web, kimi vis
Remote Control graduated from experimental in v0.42.0 (2026-09-09): take over a local session from a phone or another computer, no flag required. kimi web (v0.17.0, an alias for kimi server run --open) continues the current session in a browser chat interface, and --host exposes that server beyond the local machine, hardened with token authentication and rate limiting. kimi vis is a session visualizer supporting --port, --host, --no-open and kimi vis <sessionId> deep links. Shell mode (type !) runs terminal commands without leaving the conversation and writes the output back into context for later turns, with Ctrl+B moving long commands to the background.
The Desktop client turns process visibility into product language: tool calls, thinking and the files changed in each turn are all rendered; sensitive operations raise approval prompts; the built-in browser sits in the right panel with tabs that follow the session and page context shared with the agent, plus screenshot annotation, text-selection comments, @ file mentions and element picking.
Commercials and safety boundaries
Kimi Code is part of Kimi membership benefits and shares one quota with the membership subscription; calls from the CLI, VS Code, Desktop and third-party tools all draw from it. New plans removed the weekly quota cap and keep only the rolling 5-hour rate window. The Go tier has no coding quota: Plus and above can use Kimi Code, K3 needs Plus and above, and K3's 1M context plus HighSpeed need Pro and above. Once subscription quota runs out, Extra Usage acts as a fallback (a wallet shared with Kimi on the web, minimum ¥25 per top-up, up to 10 top-ups and ¥3,000 per day, ¥10,000 balance cap, optional monthly spending cap). The official order-of-magnitude examples: a simple request costs about ¥0.03, a complex multi-step task about ¥1.6.
On safety there are concrete hardenings. v2.1.0 tightened workspace trust boundaries so file tools can no longer reach outside the working directory through symbolic links, and project-local config only takes effect once the workspace is trusted. Since v0.36.0 untrusted workspaces cannot plant same-named fd or stty executables. Since v0.16.0 Anthropic-compatible providers no longer read ambient Anthropic shell credentials and custom headers, avoiding accidental credential leakage.
Boundaries
The K3 / K2.8 / K2.7 Code numbers are vendor-published or vendor-relayed from "external benchmark evaluations" and we have not recomputed them. K3's technical report has not been released alongside the weights, so architecture detail currently rests on the official blog and docs. HighSpeed is explicitly described by the vendor as resource-limited with a fluctuating experience while capacity is ramped. The quota system is bound to membership tiers, so capacity planning off API unit prices will be wrong - the 1M tier costs roughly twice the 256K tier, and HighSpeed three times. Model retirement is fast: the kimi-k2 series was discontinued 2026-05-25, kimi-latest on 2026-01-28, and kimi-k2.5 with the moonshot-v1 series on 2026-08-31.