Claude Code: the agentic coding tool that lives in the terminal, and the reference harness for Anthropic models
claude-code
Claude Code is the agentic coding tool from Anthropic, and its own definition is deliberately modest: it lives in your terminal, understands your codebase, and executes routine tasks, explains complex code and handles git workflows through natural language. The official docs name three surfaces - terminal, IDE, and tagging @claude on GitHub. anthropics/claude-code stood at 148,438 stars / 24,902 forks on 2026-09-28, among the highest in this class; stated plainly, the repository has no LICENSE file (the GitHub API returns null), so it is publicly readable but not openly licensed - a material difference from the Apache-2.0 core of Codex CLI that must not be blurred during evaluation. Installation moved from npm to native installers (claude.ai/install.sh, install.ps1, brew cask, winget) and npm install -g @anthropic-ai/claude-code is officially deprecated, usually a way to escape the Node version matrix and global-permission support burden. The plugins/ directory is not a shell: it ships 13 officially maintained capability packs - code-review, pr-review-toolkit, commit-commands, feature-dev, frontend-design, security-guidance, agent-sdk-dev, plugin-dev, hookify, the explanatory and learning output styles, the opus 4.5 migration and ralph-wiggum - each a bundle of custom commands plus agents, turning personal prompts into installable team assets. It is also the reference harness for Anthropic models in third-party benchmarks (the Cursor footnote routes Anthropic model scores through Claude Code). Boundaries: no open license, closed models with unavoidable data egress, a terminal-first form factor that costs non-CLI users a learning curve, and a ceiling that depends on how well any third-party model adapts to the Anthropic protocol and tool calling when swapped in. The card metric is the verifiable fact (stars). Confidence: confirmed.
- CONFIDENCE
- Confirmed
- Two or more independent sources, or reproduced by our harness
- KEY METRIC
- GitHub stars(仓库无 LICENSE)
- Confirmed · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it A (confirmed), and grade A here covers the factual layer only: repository stars and forks, the absence of a LICENSE, the four install channels and the npm deprecation, the names of the thirteen first-party plugins, the data-collection and "feedback is not used for training" clauses in the README, and Cursor's footnote naming it the comparator harness for Anthropic models. All of these come from checkable primary sources (repo README, docs tree, GitHub API, a third-party evaluation footnote). We assert no SOTA capability reading for it, so there is no claim that would force a downgrade to C.
Water level: Claude Code's value has two layers. The first is that it became the industry's default shell for terminal agents - third-party evaluations use it as the execution environment for Anthropic models, while third-party model vendors (Kimi, Z.ai) use it as a host for theirs. Both holding true at once means its tool surface and protocol have leaked out into infrastructure. The second is that its plugin system productises engineering process itself: code review, a PR toolkit, commit commands, security guidance and model migration each ship as a pack, so teams distribute packages rather than folklore prompts.
Two things stated plainly: the repo is public but not openly licensed, which is not the same kind of reusability as Apache-2.0 Codex CLI; and the models are closed with data leaving the perimeter, a hard boundary for compliance-sensitive teams.
What it is: an agentic coding tool that lives in the terminal, and Anthropic's reference harness
Claude Code's own definition is deliberately modest: an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code and handling git workflows through natural language. The official docs name three surfaces: terminal, IDE, and tagging @claude on GitHub.
anthropics/claude-code stood at 148,438 stars / 24,902 forks when fetched on 2026-09-28, among the highest for tools in this class. One thing must be stated plainly: the repository has no LICENSE file (the GitHub API returns null), so it is publicly readable but not openly licensed - a material difference from Codex CLI's Apache-2.0 that should not be blurred during evaluation.
Install: native installers first, npm officially deprecated
The recommended install path has shifted to native installers: curl -fsSL https://claude.ai/install.sh | bash on macOS/Linux, irm https://claude.ai/install.ps1 | iex on Windows, with package managers via brew install --cask claude-code and winget install Anthropic.ClaudeCode. npm install -g @anthropic-ai/claude-code is explicitly marked deprecated. Then run claude in the project directory. A product that actively deprecates its npm channel is usually escaping the support burden of Node version matrices and global install permissions.
Plugins: thirteen first-party engineering packs shipped in the repo
The plugins/ directory is not a stub. It is a maintained set of capability packs, each combining custom commands with agents: code-review, pr-review-toolkit, commit-commands, feature-dev, frontend-design, security-guidance, agent-sdk-dev, plugin-dev, hookify, explanatory-output-style, learning-output-style, claude-opus-4-5-migration, ralph-wiggum.
The list is itself a statement about the product's shape. Code review, a PR review toolkit, commit commands, feature development, frontend design and security guidance are each their own pack. hookify implies hooks are a first-class extension point that tooling can generate. agent-sdk-dev and plugin-dev are extensions for people who build extensions. Two output styles (explanatory, learning) mean even response register is a swappable part. And claude-opus-4-5-migration turns "migrate to a new model" into a plugin - a clear signal of engineering process being productised.
Why it is the standard shell for Anthropic models in evaluations
Cursor's official Composer 2 evaluation footnote is explicit: on Terminal-Bench 2.0, Anthropic model scores use the Claude Code harness, OpenAI models use the Simple Codex harness, and Cursor itself runs the designated Harbor framework averaged over 5 iterations per model-agent pair. When a third party wants to measure the agent ceiling of a Claude-family model, Claude Code is the default shell.
Paired with Claude Opus 5.5 (released 2026-09-22, top of the Artificial Analysis intelligence index at 57.62, already catalogued here), that harness is the actual execution environment behind the reading. Coding agent performance is the product of model and harness, not the model alone - the precondition for reading any leaderboard in this space.
It is also a host for third-party models
Claude Code speaks the Anthropic Messages protocol, which by 2026 has become one of the de facto standards. Kimi's official documentation ships a dedicated guide for using Kimi models inside Claude Code and publishes Anthropic-compatible base URLs (https://api.kimi.com/coding/ in China, https://api.kimi.ai/coding/ overseas); Z.ai's coding plans likewise list Claude Code among the primary shells. The implication: part of Claude Code's moat is not the model but the fact that its tool surface and protocol are widely accepted - users can run someone else's model in its shell.
Data and privacy: the policy is written into the README
The README states collection scope directly: usage data (such as code acceptances or rejections), associated conversation data, and feedback submitted via /bug. It then lists three safeguards - limited retention periods for sensitive information, restricted access to user session data, and an explicit policy against using feedback for model training. Putting "we do not train on your feedback" in the repository README rather than burying it in a privacy policy gives enterprise procurement a clause it can quote. Feedback flows through the in-product /bug command plus an official Discord.
Boundaries
The repository is public but not openly licensed; models are closed and data leaves the perimeter; a terminal-first shape carries a learning cost for non-CLI users; and when a third-party model is substituted for Anthropic's, the ceiling depends on how well that model adapts to the Anthropic protocol and its tool-calling conventions.