StateSyncDocs

DocsCutting agent costs

How to reduce Codex CLI token usage

The settings and habits that make each Codex message carry less, so a plan lasts longer.

How Codex is billed

On a ChatGPT plan you get a number of messages per five-hour window, and it changes with the model. After that you buy credits, priced per million tokens. With an API key you pay per token from the start. OpenAI's pricing page has the current figures.

A long session reaches the limit sooner, because every message carries everything read so far.

Settings worth knowing

Codex reads ~/.codex/config.toml. These keys control how big the conversation gets.

  • model_auto_compact_token_limit: the size at which history is compacted automatically.
  • model_context_window: the context size Codex assumes. Lower means earlier compaction.
  • tool_output_token_limit: the most each tool result can keep in history. Tool output is most of the bill, so this one matters.
  • compact_prompt: your own instruction for what a compaction should keep.

Habits that matter more

  1. One task per session

    A session that has done four things carries all four into the fifth. Start fresh for unrelated work.

  2. Compact at a switch, with an instruction

    Run /compact when a plan is agreed or a change lands. Set compact_prompt to keep what you always need, like failing test names and the files in play.

  3. Pick the model per task

    Limits vary a lot between models on the same plan. A small fix does not need the model with the tightest limit.

  4. Ask for narrow reads

    Name the file and the function. Ask for the failing test, not the whole run, and git diff on a path.

The trouble with a fixed cut

tool_output_token_limit keeps the start of a long result and drops the rest. A test runner often puts the lines that matter at the end. A size limit cannot tell which lines the next step needs.

Where StateSync fits

StateSync runs Codex in its chat and terminals on your ChatGPT plan or API key, and gives it StateSync's tools: reads, searches and command output come back sized to what the next step needs, instead of being cut at a fixed length. Our Codex measurements are early. The saving shows on longer work in existing codebases, and the settings and habits on this page still apply alongside it.

Where to start

  • Today: one task per session, /compact at a switch, and a compact_prompt.
  • This week: a default model for each kind of task, and narrower reads.

Sources

  1. OpenAI, Codex configuration referencelearn.chatgpt.com/docs/config-file/config-reference
  2. OpenAI, Codex basic configurationlearn.chatgpt.com/docs/config-file/config-basic
  3. OpenAI, Codex pricing and limitslearn.chatgpt.com/docs/pricing
  4. TokenBench harness, our 55-task rungithub.com/Silverspine1/TokenBench

Something missing or wrong on this page? Write to support@statesync.net.