DocsCutting agent costs
Tools that make coding agents cheaper, compared
The tools that cut coding agent spend, sorted by what they act on.
The layers
Start with where the tokens are. Claude Code sends the whole conversation with every request, and on our benchmark run 94 percent of the plain agent's tokens were that resend. Much of that resend is tool output, so a tool that shrinks tool output acts on a large share of the bill. A tool that shortens the model's own writing acts on a much smaller share.
| Layer | What it acts on | Tools |
|---|---|---|
| Shell output | Command results before the agent reads them | RTK |
| A proxy | Tool results and history in the outgoing request | tamp, pxpipe, Caveman proxy, Edgee, Headroom |
| The model's writing | The agent's own prose and code volume | Caveman skill, Ponytail |
| Retrieval | Which files the agent reads at all | claude-context |
| The client | Each tool result as it arrives, plus a local map of the code | StateSync |
| Measurement | Reports spend, changes nothing | ccusage |
Each tool
RTK
Filters shell output before it reaches the agent, with rules for git, grep, test runners and package managers. Wired in through agent hooks, so git status becomes rtk git status. Claims "60-90%" on common commands, and says itself that the saving dilutes across the whole bill and that Claude Code's built-in Read, Grep and Glob skip the hook. Free, Apache 2.0.
Headroom
A compression library with a proxy, an agent wrapper and an MCP server. Claims "20% fewer tokens for coding agents". It is the one tool here that publishes a quality check beside its saving, on small GSM8K and TruthfulQA samples. Free, Apache 2.0, with an enterprise tier.
tamp
A proxy in front of the provider API, started with npx @sliday/tamp. Claims "52.6% fewer input tokens", without describing the measurement. It leaves error results, paths, URLs and version strings alone because compressing them can corrupt them. Free, MIT.
Caveman
A skill that shortens the agent's own prose, claiming "65% output tokens reduced". Its README says input and reasoning tokens are untouched and the skill adds about a thousand input tokens a turn. A separate proxy trims what the agent reads. Free skill, source available runtime.
Ponytail
A ruleset that pushes the agent to write the least code that works. Reports 22 percent fewer tokens and 20 percent lower cost on a twelve-task run, and corrected its own earlier, larger figure in public. Free, MIT.
pxpipe
A proxy that turns bulky context into images, which are priced differently. Claims "59-70%" lower billing, and publishes the cost: exact strings like identifiers can be lost. Free, MIT.
claude-context
Semantic code search over MCP, so the agent finds code instead of reading widely. Claims about 40 percent fewer tokens. Needs a hosted vector database and an embedding API key. Free, MIT.
Edgee
A gateway that trims tool results and shortens output. Claims "-50% tokens on a typical session", shown on one example session. License and price are not on the page.
ccusage
Measurement only. Reads local session files and reports tokens and cost. Run it before installing anything else here, so you know what your bill is made of. Free.
StateSync
A desktop app that runs your coding agent on your own account: Claude Code, Codex, OpenCode, Cursor CLI, Copilot, Pi, Auggie, Factory Droid and Grok. As each tool result arrives, it carries forward the part the agent will use, and the whole result is one call away if it needs it. Install it, open a folder, pick an agent. Nothing to wire in per agent. In our run of the TokenBench 55-task set it cut cost 46% and was 48% faster across 55 tasks on Claude Sonnet 4.6, with quality scored on every task: higher on 15, level on 34, lower on 6. Every task is public, including the ones it lost. It is a paid tool, with a free trial.
Side by side
| Tool | Layer | Setup | Claim | Basis stated |
|---|---|---|---|---|
| RTK | Shell output | Hook per agent | 60-90% of command output | Vendor says it dilutes across the bill |
| Headroom | Proxy or library | Wrapper, proxy or MCP | 20% for coding agents | Traces, plus a small quality check |
| tamp | Proxy | npx, set base URL | 52.6% input tokens | Not described |
| Caveman | Model's writing, plus proxy | Skill, proxy setup | 65% output tokens | Vendor scopes the skill to output only |
| Ponytail | Model's writing | Rule file per agent | 22% tokens, 20% cost | 12 tasks, own harness |
| pxpipe | Proxy, images | npx, set base URL | 59-70% billing | Own instrumentation, lossy on exact strings |
| claude-context | Retrieval | MCP plus hosted database | About 40% tokens | Details not published |
| Edgee | Gateway | Install script | 50% on a session | One example session |
| ccusage | None | npx | None | Measurement only |
| StateSync | Client | Install once | 46% lower cost | 55 tasks, quality scored, harness public |
How to read the numbers
- What is the percentage of? Ninety percent of shell output and twenty percent of the bill can be the same result.
- Was quality measured beside it? A saving with no quality score is half a result.
- What does it take to keep running? Every hook, proxy, skill and rule file is one more thing per agent that can quietly break after an update.
- What happens to exact strings? Paths, hashes and version numbers are where lossy compression hurts.
We have not run these tools against each other under one method, so nothing here says one produces worse work than another. Start with ccusage and the free fixes, then decide which layer the rest of your bill is at.
Sources
- Headroom, README and product pagegithub.com/headroomlabs-ai/headroom
- RTK, READMEgithub.com/rtk-ai/rtk
- tamp, READMEgithub.com/sliday/tamp
- Caveman, READMEgithub.com/JuliusBrussee/caveman
- Ponytail, READMEgithub.com/DietrichGebert/ponytail
- pxpipe, READMEgithub.com/teamchong/pxpipe
- claude-context, READMEgithub.com/zilliztech/claude-context
- Edgee Token Compression, product pagewww.edgee.ai/token-compression
- ccusage, READMEgithub.com/ryoppippi/ccusage
- Anthropic, Claude Code: manage costs effectivelycode.claude.com/docs/en/costs
- TokenBench harness, our 55-task rungithub.com/Silverspine1/TokenBench