Codex: an Apache-2.0 core, one agent in four form factors, and the reference harness for OpenAI models
openai-codex
Codex is the coding agent from OpenAI, and in 2026 it is not one product but four entry points onto one core: the CLI in your terminal, an IDE extension (officially for VS Code, Cursor and Windsurf), a desktop app launched with a single codex app command, and Codex Web running tasks remotely at chatgpt.com/codex. The openai/codex repository is Apache-2.0 and stood at 126,918 stars / 19,819 forks when fetched on 2026-09-28 - one of the few frontier-vendor coding harnesses whose harness is itself open source, so context assembly, tool ordering and approval gating can all be read. Safety is a first-class citizen: docs/ carries sandbox.md, execpolicy.md, exec.md, agents_md.md, skills.md and slash_commands.md, where execpolicy turns which commands may run unattended into an auditable, versionable, org-distributable policy instead of model goodwill, exec lets Codex run in CI, scripts and batch jobs rather than only waiting in a terminal, and AGENTS.md has become a de facto cross-tool standard. Four install channels (install.sh, install.ps1, npm, brew) and two billing models: usage included in ChatGPT plans versus per-token API keys, a cost question to settle before choosing. Codex is also the reference harness for OpenAI models - the Cursor Terminal-Bench 2.0 footnote states that OpenAI model scores use the Simple Codex harness, and the vendor-published Terminal-Bench 4.0 reading of 58.18% is already carried on the GPT-6 Astra card because it measures the model rather than the shell, so this entry does not claim it. Boundaries: models are closed and data egress cannot be avoided; developers.openai.com returns 403 to automated fetching, so the fact layer rests on the repository README/docs and the GitHub API. The card metric keeps only the verifiable fact layer (license plus stars). Confidence C (vendor-claim).
- CONFIDENCE
- Vendor Claim
- Official model card or keynote only, no independent re-test
- KEY METRIC
- Harness 内核许可 + GitHub stars
- Vendor Claim · 2026-09
- MATURITY
- Product
- research → demo → product → production
Our takeWe grade it C (vendor-claim). Terminal-Bench 4.0 is hosted by Stanford, Harbor and the Laude Institute and is a credible third-party benchmark, but the 58.18% reading was published vendor-side and we have not recomputed it;
developers.openai.comalso returns 403 to automated fetching, so the detailed capability surface can only be inferred from the open repository. What is confirmable is hard: Apache-2.0 licensing, 126,918 stars, four install channels, the sandbox / execpolicy / exec / skills / AGENTS.md concepts in the docs tree, and Cursor's footnote naming it the comparator harness for OpenAI models.Water level: Codex's strategic position is not "another CLI". Two things matter. It open-sourced the harness, which makes its sandbox and execution policy a readable reference implementation for the industry. And competitors use it as their baseline comparator, which is the field's way of saying it represents the real ceiling of OpenAI's models. For teams, four surfaces on one core means local, IDE, desktop and cloud-hosted work can share one AGENTS.md and one policy set - a far cheaper migration than switching vendors.
Costs stated plainly: closed models, data leaves the perimeter, and the two billing models are easy to conflate.
What it is: one Apache-2.0 coding agent shipped as four surfaces
Codex is OpenAI's coding agent, and in 2026 it is not one product but four entry points onto the same agent: the Codex CLI (runs locally; the repo describes it as "a lightweight coding agent that runs in your terminal"), an IDE extension (officially for VS Code, Cursor and Windsurf), a desktop app (launched with a single codex app), and Codex Web (the cloud agent at chatgpt.com/codex, where tasks run remotely). All four share one core, which is the real difference from products that ship a plugin and stop there.
openai/codex is Apache-2.0 licensed and stood at 126,918 stars / 19,819 forks when fetched on 2026-09-28. That matters: this is one of the few frontier-vendor coding harnesses whose harness is itself open source, so anyone can read how it assembles context, orders tools, and gates approvals.
Install and auth: four install channels, two billing models
Install coverage is unusually complete: one line on macOS/Linux (curl -fsSL https://chatgpt.com/codex/install.sh | sh), one line on Windows (irm https://chatgpt.com/codex/install.ps1 | iex), plus npm install -g @openai/codex and brew install --cask codex. Standalone installers download from releases.openai.com/codex by default and fall back to GitHub Releases when metadata or assets are unavailable, with CODEX_INSTALLER_USE_RELEASES_OPENAI_COM=false to force the GitHub path. Treating distribution-source failure as an explicitly configurable engineering problem is the kind of detail that only appears after being hit at real scale.
Authentication has two shapes. The recommended path is Sign in with ChatGPT, using Codex as part of a Plus, Pro, Business, Edu or Enterprise plan; an API key also works but needs additional setup. The first is subscription-inclusive usage, the second is per-token billing - the same agent under two entirely different cost models, and that is the first calculation a team should do.
The security model is a first-class citizen: sandbox, approvals, execpolicy
The fastest way to judge an agent harness is to read the nouns in its docs directory. openai/codex's docs/ contains sandbox.md (sandboxing and approvals), execpolicy.md (execution policy), exec.md (non-interactive execution), config.md and example-config.md, agents_md.md (AGENTS.md), skills.md, slash_commands.md, authentication.md, install.md and open-source-fund.md. The directory is itself a capability map:
- sandbox + approvals: the filesystem and network isolation boundary, plus human approval before crossing it. This is the precondition for letting an agent run on a machine that holds real credentials.
- execpolicy: which commands may run unattended and which must be asked about, expressed as policy rather than left to model discretion. Policy is auditable, versionable, and distributable across an organisation.
- exec: a non-interactive execution entry point, meaning Codex can live in CI, in scripts and in batch jobs rather than only in a terminal waiting for a human.
- AGENTS.md: the repository-level brief for agents. It has become a de facto cross-tool standard - Codex is not the only harness that reads it.
- skills / slash commands: reusable capability packaging and shortcuts; the layer that turns personal prompts into team assets.
Why it functions as a reference harness in evaluations
Terminal-Bench 4.0 is hosted by Stanford, Harbor and the Laude Institute and is one of the main 2026 benchmarks for long-horizon agent work in a terminal. Any credible cross-vendor comparison has to pin the harness, and Cursor's official Composer 2 evaluation footnote says so explicitly: Anthropic model scores use the Claude Code harness and OpenAI model scores use the Simple Codex harness, while Cursor's own score came from Harbor, the designated framework for Terminal-Bench 2.0, averaged over 5 iterations per model-agent pair.
The signal is that Codex is now treated by peers as the standard shell for OpenAI-family models. In coding agents, harness and model multiply rather than add: the same model can move a dozen points on long-task completion between shells. A harness that competitors use as their baseline comparator is, by implication, the industry's estimate of that model's real ceiling.
Paired with GPT-6 Astra (already catalogued on this site), the vendor-published Terminal-Bench 4.0 reading is 58.18%, ranked first. That number was published vendor-side, we have not recomputed it, and it already sits on the GPT-6 Astra card (it measures the model, not the shell), so this entry does not claim it as its own metric and is graded C (vendor-claim); the card metric is limited to the verifiable fact layer - core license and repository stars.
Boundaries
The core is Apache-2.0, but the models are closed and reachable only through OpenAI's services, so data leaving the perimeter is unavoidable. Usage included in a ChatGPT plan and usage billed against an API key follow very different accounting, and mixing them distorts cost analysis. developers.openai.com returns 403 to automated fetching; the factual layer of this entry therefore rests on the GitHub repository README and docs directory, GitHub API metadata, and third-party evaluation footnotes.