DocsCutting agent costs
How to reduce Claude Code costs
The free fixes that keep a Claude Code bill down, and what StateSync adds on top.
Where the money goes
Claude Code works in a loop. It reads the conversation, calls a tool, adds the result, then sends everything again. A file it read early on is paid for on every call after that.
On the run behind our numbers, 94 percent of the plain agent's tokens were that resend, billed at the cache read price. What the model actually wrote was a small share of the bill.
The free fixes
All built in. Do these first, whatever else you use.
Clear between tasks
/clearstops the next task paying to carry the last one. It is the biggest free saving there is.Compact at a switch, not mid-task
Run
/compactwhen a plan is agreed or a change lands, and say what to keep:/compact keep the failing test names. Mid-task it drops detail the agent still needs.Match the model to the job
Small fixes and tests rarely need the largest model.
/modelswitches mid-session.Keep CLAUDE.md short
It loads into every session, so every line is paid for on every call. Keep commands and conventions. Move long explanations into files the agent can read when it needs them.
Ask for narrow reads
Name the file and the function. Ask for the failing test, not the whole run, and
git diffon a path, not the tree. Every line returned rides along on every later call.Use subagents on purpose
A subagent keeps exploration out of the main thread, but it pays for its own reads. Use one for research, not for every small task.
Keep the start of the session stable
Prompt caching makes the resend cheaper. Editing CLAUDE.md mid-session throws the cache away and the next call pays full price.
Check /cost
/costshows what a session spent on API billing. When one feels expensive, find the call that made it so.
What habits cannot do
Habits work between tasks, or when you remember them. None of them can look at each result as it comes back and keep only what the next step needs. That happens hundreds of times in a task.
Where StateSync fits
StateSync is a desktop app that runs Claude Code on your own account. As each result arrives, it carries forward the part the agent will use, and the whole result is one call away if it needs it. A local map of your code means the agent starts near the right file. Same model, same account, nothing to configure.
Short tasks save the least.
| Task length | Tasks | Cost saving | Cheaper |
|---|---|---|---|
| Under 2 minutes | 18 | 27% | 17 of 18 |
| 2 to 4 minutes | 27 | 50% | 27 of 27 |
| 4 minutes and up | 10 | 50% | 9 of 10 |
Measured by us on the TokenBench 55-task set, on Claude Code with Claude Sonnet 4.6, with one StateSync run per task against the plain agent's earlier runs. See every task, losses included.
On a Max plan
Your bill is flat, so the saving shows up as room. Fewer requests and tokens per task means more work before you reach a session or weekly limit. On the run that was 29 percent fewer requests and 45 percent fewer tokens for the same work.
Where to start
- Today:
/clearbetween tasks,/compactat a switch, a shorter CLAUDE.md. - This week: a default model for each kind of task, and narrower reads.
- Then try StateSync: install it, open your project folder, start a chat and pick Claude Code. The Account page in the app shows what you saved.
Other tools cover parts of this. The comparison sorts them by what they act on.
Sources
- Anthropic, Claude Code: manage costs effectivelycode.claude.com/docs/en/costs
- Anthropic, Claude Code: memory and CLAUDE.mdcode.claude.com/docs/en/memory
- Anthropic, Claude Code: subagentscode.claude.com/docs/en/sub-agents
- Anthropic, prompt cachingdocs.anthropic.com/en/docs/build-with-claude/prompt-caching
- TokenBench harness, our 55-task rungithub.com/Silverspine1/TokenBench