OPEN SOURCE DEEP DIVE
Caveman: The Token-Cutting Skill, Proxy and Middleware for Coding Agents
A MIT rule file makes agents talk like cavemen (code and errors never shortened), a local Go proxy compresses what agents read before each call with byte-exact originals recoverable, and middleware brings the same to your own app. JetBrains' 86-task A/B: -8.5% output tokens, quality flat; the repo's pinned 54-run suite: -33.2% input tokens, 18/18 checks passed.
The bill is per word, so make it say less
A token is the unit AI billing counts, roughly three quarters of a word. Your agent pays for every token it writes and every token it reads, and most agents write like a cover letter and read like a firehose. Caveman attacks both ends: one rule file shrinks what it says, a local proxy shrinks what it reads, and a middleware layer brings the same capability into code you ship yourself. The project started as a joke on a Friday in April 2026, hit 4,000 stars in a week, and is past 100,000 stars at ingest time, with an Adobe Research citation, a JetBrains lab test and a ThePrimeagen reaction video attached.
The README's opening comparison is blunt: the same React re-render diagnosis costs 69 tokens from a normal agent and 19 from a caveman agent. Same diagnosis, same fix, same useMemo; the only thing that died was the throat-clearing. The discipline is equally explicit: code, commands, file paths and exact error messages are never cavemanned. Security warnings and "are you sure?" confirmations come back in full sentences, and only then does the register resume.
Three pieces: a rule file, a local proxy, a middleware layer
The skill is the smallest entry point: one MIT rule file, free forever, working in 30+ agents (Claude Code, Codex, Gemini, Cursor, Windsurf, Cline, Copilot and more), installed with a single npx skills add JuliusBrussee/caveman -g. Intensity is a dial: /caveman lite|full|ultra, plus wenyan-lite|wenyan-full|wenyan-ultra for classical Chinese, because someone actually asked. /caveman-commit and /caveman-review carry the register into commit messages and review comments; /caveman-compress shrinks CLAUDE.md-style memory files by 46% on average with headings, code, paths and URLs verified intact and a backup kept for rollback; pixel mode goes further and renders the skill itself as PNG pages the model reads as images, cutting 1,069 estimated tokens to 415, a 61% reduction.
The proxy is the heavyweight: a Go program (565 .go files, roughly 139k lines) that runs on your machine between the agent and the AI provider. Before every request it compresses what the agent is about to read: detect() classifies incoming content into json, log, code, diff, search and text, each type with its own keep-rules that preserve 70-95% of the savings; every squeezed byte keeps a byte-exact original in local SQLite behind one recovery handle, so the agent can always pull the full text back. contextwindow.Pack selects relevant originals with BM25 plus recency plus error signals while preserving chronology. The default wrap hands the agent five MCP tools, the browse server when Chrome resolves, command-output shrink on Claude, opencode, Gemini, Hermes and OpenClaw, and pixel mode on new skill installs; Codex skips the shrink hook because its runtime rejects the rewrite (openai/codex#18491). Wrapping covers 10 agents without editing their config files, and provider passthrough includes Claude Pro/Max logins. learn and learn implement scan past sessions for token sinks and propose fixes, auto-reverting when a change measures no savings; trial runs A/B comparisons.
The middleware (alpha) moves the same compression into your own application: on the TypeScript side it wraps Vercel AI SDK, OpenAI, Anthropic, Google, LangChain, Strands, Mastra and MCP; on the Python side it adds LangGraph, LiteLLM, Agno, CrewAI, PydanticAI, AutoGen, LlamaIndex and FastAPI. Tool results are shrunk before the model sees them, the original stays in your history, and the model can fetch it back.
Evidence: three number sets, and the red row stays red
Adobe Research's CAVEWOMAN paper (arXiv 2606.24083) measured the register across eight models, five datasets and five compression levels: output-side caveman style cuts realized cost 1.4 to 2.4x per model, up to 3x in the best case. The same paper carries the other half of the finding: compressing the human's prompt into caveman-speak makes models answer longer and worse. Caveman therefore never rewrites your prompts. Only the agent's mouth.
JetBrains ran a paired A/B on 86 real coding tasks (Claude Code 2.1.200, skill only, July 2026, before the proxy existed): 8.5% fewer output tokens, about 10% of cost, with no detectable quality change (sign test p = 0.82). Their conclusion, that an agent's bill is mostly reading rather than writing and no talking style fixes that, is precisely why the proxy was built.
The repo's own pinned 54-run Claude Code suite (three runs per case, provider-reported input tokens, every answer checked against an exact oracle) takes input tokens from 885,793 to 591,673, -33.2%, with 18 of 18 answer checks passed and a case-clustered 95% interval of 14.6% to 48.5%. The Dashboard HTML row is +9.9%, it is red, and the README says it stays red: that case had no compression transform, so the proxy paid its own overhead and won nothing back. In the maintainer's words, the day the red row is hidden is the day you should stop trusting the green ones. In the same suite, Headroom's wrap saved 6.7% and failed 3 of 18 checks.
The committed eval snapshot closes the writing side: ten dev questions against an already-terse "Answer concisely." control still cut output tokens by 50% at the median, measuring length only, not correctness, as the README labels it. Read together, the three sets give the honest picture: chat-style Q&A cuts deep; agentic coding sessions, where most tokens are code and tool calls the skill never touches, cut high single digits on output with quality flat.
The reading side: 129.8x and a 219k-character tax
The browse server measures web pages as an input surface: a focused question against a 200-row table costs 121 tokens through Caveman versus 15,704 for the Playwright ARIA snapshot, a 129.8x gap; tiny forms lose only 2.3x, and the benchmark document says so.
subagent-tax measures a different hidden read: on one real machine, 219k of a 267k-character request was tool schemas, the full tool surface every subagent re-sends before doing any work. The command exists so you can measure your own tax on your own machine instead of trusting someone else's number.
Same-layer comparison: what each tool touches, and who checked
The comparison table quotes each competitor's README as of 2026-09-19 and labels who verified what. RTK touches shell command output only (Read and Grep tool calls bypass it); JetBrains, same lab and same method across 425 billed trials, measured +7.6% median cost per task at low reasoning effort (p = 0.004). Headroom touches tool output, logs, files and history through a reversible cache but ships telemetry on by default. context-mode runs in a sandbox and returns searchable sections rather than the whole text. pxpipe re-renders context as images, admits it is lossy, and fails silently. Caveman's differentiators: originals are always recoverable (byte-exact, local SQLite, one recovery handle); the CLI's anonymous counts are on by default but one caveman telemetry off disables them, while the skill and hooks never phone home.
Licensing, telemetry, and when to skip
The license is a split: the skill and CLI are MIT; the proxy runtime is BSL-1.1, converting to Apache-2.0 on the earlier of 2030-06-21 or the fourth anniversary of first public release, with third-party hosting requiring a commercial license. The telemetry posture is stated in the comparison table: CLI anonymous counts on by default and switchable off, skill and hooks never sending anything.
The README also ships a "when to skip" list: per-request billing saves nothing because there is no per-token meter, and pure code generation outputs mostly code, which the skill never touches. Publishing the boundaries on the marketing page is the same value system as keeping the red row.
Community record: #1 on Hacker News with 904 points and 366 comments, #1 on GitHub Trending overall in July 2026, Trendshift Repository of the Day in April and #1 Go repo of the month in July, #8 Product of the Day on Product Hunt. The New Stack's skeptical piece ("might not save as many tokens as you think") is linked from the README verbatim, with a one-line note: Fair headline. We link it anyway. For a project whose entire pitch is saving tokens, that self-auditing posture is more worth copying than any single number.