DocsCutting agent costs
How to reduce Codex CLI token usage
The settings and habits that make each Codex message carry less, so a plan lasts longer.
How Codex is billed
On a ChatGPT plan you get a number of messages per five-hour window, and it changes with the model. After that you buy credits, priced per million tokens. With an API key you pay per token from the start. OpenAI's pricing page has the current figures.
A long session reaches the limit sooner, because every message carries everything read so far.
Settings worth knowing
Codex reads ~/.codex/config.toml. These keys control how big the conversation gets.
model_auto_compact_token_limit: the size at which history is compacted automatically.model_context_window: the context size Codex assumes. Lower means earlier compaction.tool_output_token_limit: the most each tool result can keep in history. Tool output is most of the bill, so this one matters.compact_prompt: your own instruction for what a compaction should keep.
Habits that matter more
One task per session
A session that has done four things carries all four into the fifth. Start fresh for unrelated work.
Compact at a switch, with an instruction
Run
/compactwhen a plan is agreed or a change lands. Setcompact_promptto keep what you always need, like failing test names and the files in play.Pick the model per task
Limits vary a lot between models on the same plan. A small fix does not need the model with the tightest limit.
Ask for narrow reads
Name the file and the function. Ask for the failing test, not the whole run, and
git diffon a path.
The trouble with a fixed cut
tool_output_token_limit keeps the start of a long result and drops the rest. A test runner often puts the lines that matter at the end. A size limit cannot tell which lines the next step needs.
Where StateSync fits
StateSync runs Codex in its chat and terminals on your ChatGPT plan or API key, and gives it StateSync's tools: reads, searches and command output come back sized to what the next step needs, instead of being cut at a fixed length. Our Codex measurements are early. The saving shows on longer work in existing codebases, and the settings and habits on this page still apply alongside it.
Where to start
- Today: one task per session,
/compactat a switch, and acompact_prompt. - This week: a default model for each kind of task, and narrower reads.
Sources
- OpenAI, Codex configuration referencelearn.chatgpt.com/docs/config-file/config-reference
- OpenAI, Codex basic configurationlearn.chatgpt.com/docs/config-file/config-basic
- OpenAI, Codex pricing and limitslearn.chatgpt.com/docs/pricing
- TokenBench harness, our 55-task rungithub.com/Silverspine1/TokenBench