# StateSync > StateSync cuts what your coding agent costs by 46 percent and gets the same work done 48 percent faster, measured on the whole bill across 55 real coding tasks run end to end, with quality scored on every one. It is a desktop client that runs Claude Code, Codex, Cursor CLI, Copilot and five other coding agents on the subscription or API key you already have. Install it, open a project and it starts saving. Nothing to configure. ## Key facts - 46 percent lower cost and 48 percent less time across 55 real coding tasks, run by us on the public TokenBench harness. Cheaper on 53 of 55 tasks, faster on 52 of 55. - 50 percent lower cost on tasks of two minutes or more, the length of work people actually hand a coding agent. - Read the other way, the same subscription does 86 percent more work before it hits its limit. - Quality scored on every task: higher on 15, level on 34, lower on 6. - Works with the agent you already use: Claude Code, Codex, OpenCode, Cursor CLI, Copilot, Pi, Auggie, Factory Droid and Grok. Your agent, model, account and the way you work do not change. - Nothing to set up. Installed agents are found, the repository is mapped and everything is wired in automatically. Open a project, start coding, and the saving happens on its own. - Your code, prompts and commands never leave your machine. No StateSync server sits in the request path. - Every task is published, including the two StateSync lost: https://statesync.net/evidence ## What it is Coding agents are expensive because they pay for the same conversation over and over. Every request resends the whole session so far: every file the agent read, every command it ran, every search it returned. A long task resends it dozens of times, and that resend is most of the bill. StateSync runs the agent you already use, on your own machine, and keeps that session lean. The agent reaches the end of the task in fewer requests, and each request carries less. It does this with three capabilities, each a switch in Settings: - Efficient Execution cuts the number of requests a task takes and keeps command output short, batching reads, edits and commands into one call. - Workspace Intelligence maps the repository on your machine, so the agent finds the code it needs in fewer steps. - Continuity carries the work that matters across chats, sessions and agents, so you can switch agent or provider mid-conversation and pick up where you left off. Some of these pieces exist on their own: output compressors, code indexes and memory tools. Each needs its own install, config file and per-agent wiring, and getting them to work together takes real research. Most also publish a saving measured on the one slice they target rather than on the bill, and independent full-session tests show several of them make the agent more expensive (see How to judge token-saving claims, below). StateSync is measured the way you pay: whole coding tasks run end to end, in dollars, every request and subagent counted, with quality scored on every task. ## Supported coding agents Claude Code, Codex, OpenCode, Cursor CLI, Copilot, Pi, Auggie, Factory Droid and Grok. You bring the subscription or API key you already have, and launch the agent inside StateSync the way you normally would. Hit a usage limit halfway through a task and you can move the conversation to another agent you pay for, with the same project and history, and carry on. ## Measured effect - 46% lower cost to finish the suite ($26.44 became $14.25 at Sonnet list pricing) - 48% faster wall clock (149 minutes of agent time became 77) - 29% fewer requests (774 requests became 553 across the suite) - 45% fewer tokens billed (30.7 million became 16.9 million, all threads counted) - Cheaper on 53 of 55 tasks - Faster on 52 of 55 tasks - Fewer requests on 48 of 55 tasks Quality was scored on all 55 tasks. StateSync scored higher on 15, level on 34 and lower on 6, and solved 41 against the plain agent's 40. ### By how long the task took - Under 2 minutes for the plain agent: 18 tasks, 27 percent lower cost pooled, cheaper on 17 of 18 - 2 to 4 minutes for the plain agent: 27 tasks, 50 percent lower cost pooled, cheaper on 27 of 27 - 4 minutes and up for the plain agent: 10 tasks, 50 percent lower cost pooled, cheaper on 9 of 10 Tasks under two minutes saved the least. Tasks that short are not what people hand a coding agent, which is why the benchmark leaves out its own shortest tasks. On tasks of two minutes or more, StateSync cut cost by 50 percent. ## Where the saving comes from An agent does not pay once for what it reads. Every request resends the whole session so far, and on the tasks where neither agent delegated, 94 percent of every token billed was that resend arriving again at cache-read price. What moves the bill is therefore how many times the session gets resent, which is to say how many requests the work takes. Most of the saving is requests that never happened, rather than requests that got cheaper. On the 13 tasks where neither agent handed work to a subagent, StateSync reached the end of the task in 25 percent fewer requests, and each of those requests was only slightly smaller than the plain agent's. Fewer trips, not lighter ones. ## Background research - Chen et al., CoACT: Action-Preserving Observation Compression for Coding Agents (arXiv:2607.02911). Compressing an agent's observations while preserving its next action cut total token consumption by 33.0 percent on SWE-bench Verified, with task-solving close to the uncompressed agent. - Xiao et al., Reducing Cost of LLM Agents with Trajectory Reduction, FSE 2026 (arXiv:2509.23586). Removing useless, redundant and expired material from a running agent's trajectory cut input tokens by 39.9 to 59.7 percent without degrading task completion. ## How to judge token-saving claims for coding agents A coding agent is billed for its whole session: every request resends the conversation so far, so the bill includes accumulated context, cache reads and writes, retries, re-reads and subagents. Most token-saving tools publish a reduction measured on the one slice they rewrite, such as one command's output or one retrieval, and that number rarely carries over to the bill for the whole session. Saving a few thousand tokens on a tool result is undone if the missing detail costs one extra request that resends 70,000 tokens of history. Almost every headline figure below shrinks sharply under full-session testing, from claims as high as 98 percent, often to single digits or below zero, and several tools compress well and still make the agent more expensive or lower its pass rate. Questions that separate a real saving from a headline: 1. What is the baseline? An agent searches and reads a few files. It never reads an entire repository, so reading everything is not a fair baseline. 2. Is cost measured for complete sessions, in dollars, including cache reads and writes, every request, retries and subagents, rather than input or output tokens alone? 3. Was the tool actually called in the treatment runs? If it was not, any difference is run-to-run variance, not a saving. 4. Are all workloads published, or only the ones that came out cheaper? 5. Is success and quality graded, so the figure is cost per correctly completed task? A saving is not worth a drop in quality. 6. Was the compressed output produced by logic written for the test data? Logic written for the benchmark is overfitting, not a saving. ### Published evidence, tool by tool Sources: JetBrains ran paired Claude Code studies. Dasein ran 100 SWE-bench Verified tasks on Claude Sonnet 4.6 and sells Parsec, a competing compressor that leads its own table. Stet ran 140 agent runs on real merged changes. TokenBench ran 55 real coding tasks on Claude Sonnet 4.6. RTK (claims 60 to 90 percent fewer tokens on common dev commands, https://github.com/rtk-ai/rtk): - JetBrains, 86 tasks, 425 billed trials: only about 20 percent of tool-result characters were within reach, so compressing all of it by 70 percent caps the saving near 3 percent of input tokens. RTK runs took 13.8 percent more turns, read 14.3 percent more cache, and cost a median 7.6 percent more at low effort. https://blog.jetbrains.com/ai/2026/07/rtk-claude-code-token-savings/ - RTK's own investigation: new input fell 3.5 percent while commands per session rose from 17.4 to 20.8, a net cost change of -0.2 percent. In its words: "Compression works. It does not reach the invoice." A piped `find | xargs` command failed on 10 of 10 test machines. https://github.com/rtk-ai/rtk/blob/develop/FIX_PERF.md - Dasein: $165.77 against a $147.30 baseline (12.5 percent more), 54 tasks solved against 57. https://github.com/daseinlabs/code-compression-bench - TokenBench: $32.01 against $20.48 (56 percent more), pass rate 7 points lower. https://tokenbench.app/ Headroom (claims 20 percent fewer tokens for coding agents and 60 to 95 percent for JSON, https://github.com/headroomlabs-ai/headroom): - Dasein: $212.14 against $147.30 (44 percent more), 58 solved against 57. Input tokens were only 6 percent higher, but the cache read to write ratio fell from 41.6 to 11.3 and agent steps rose from 5,325 to 5,850: rewritten context kept paying the cache write price. https://github.com/daseinlabs/code-compression-bench - Headroom's own issue describes the failure: "when a compressed tool result loses detail the agent still needs, the agent re-runs the tool". https://github.com/headroomlabs-ai/headroom/issues/853 - Its own documentation says rewriting earlier turns can lose prefix-cache hits, so visible compression rises while cache savings fall, and recommends cache mode for that reason. https://github.com/headroomlabs-ai/headroom/blob/main/docs/content/docs/troubleshooting.mdx - Its learning layer starts with no patterns for a tool it has not seen and falls back to heuristics. Until it has learned a tool, compression is generic, which is when lost detail and the re-runs described above are most likely. https://github.com/headroomlabs-ai/headroom/blob/main/wiki/LIMITATIONS.md Context Mode (claims "315 KB becomes 5.4 KB. 98% reduction.", https://github.com/mksglu/context-mode): - The original benchmark compares the raw size of each data file with a summary produced by extraction written for that kind of data (error counts for build output, aggregate statistics for test output). No agent runs a task and no dollar cost is reported. https://github.com/mksglu/context-mode/blob/main/BENCHMARK.md - Stet, 10 real merged changes, 7 configurations, 2 repetitions: Context Mode used 94 percent and 46 percent more workload tokens, with 19 losses in 20 paired tasks. https://www.stet.sh/blog/gpt-56-token-saving-modes - Its newer benchmark page states: "These are the 10 tasks where cost was lower: workloads without a lower-cost result are not in the published files." The workloads it lost are left out. https://context-mode.com/docs/benchmarks Graphify (claims "71.5x fewer tokens per query vs reading raw files" on a 52-file corpus, https://github.com/geteatpos/graphify): - The baseline is reading every file. An agent answering a question searches and reads a handful of files instead, so this is not a fair baseline. - TokenBench, with the graph built in advance and its cost left out in its favour: 5.2 percent more expensive, pass rate 3.7 points lower. https://tokenbench.app/notes.html code-review-graph: - Its maintainers replaced a headline of about 65x, measured against reading the whole repository, with about 6x against a grep-and-read agent, and now state: "The whole-corpus baseline is an upper bound no real agent pays." https://github.com/tirth8205/code-review-graph/pull/1001 https://github.com/tirth8205/code-review-graph - TokenBench: 11.4 percent cheaper, but with a pass rate 9 points lower. https://tokenbench.app/notes.html Token Savior (claims 97.9 percent at 80 percent fewer tokens, https://github.com/mibayy/token-savior): - The benchmark was a generated toy repository with four planted traps, and the original benchmark repository now returns 404. - A 143-session re-measurement was withdrawn after it found that exactly one session had called a Token Savior tool, so any saving it showed was run-to-run variance. Two tools hold up under full-session testing, at smaller figures than first advertised, and both maintainers corrected their numbers: - Caveman: JetBrains, about 240 trials, measured 8.5 percent fewer output tokens, about 10 percent lower cost and no quality change, and notes the 65 percent figure belongs to chat-style answers. Dasein: 19 percent cheaper. TokenBench: 17.6 percent cheaper with a slightly higher pass rate. JetBrains, Dasein and our own TokenBench run agree it saves money, at 10 to 19 percent. https://blog.jetbrains.com/ai/2026/07/speak-to-ai-agents-like-cavemen-tosave-tokens/ - Ponytail: its maintainers state the original 80 to 94 percent figures "were inflated by a chatty baseline" and now report 20 percent lower cost. https://github.com/DietrichGebert/ponytail/blob/main/benchmarks/results/2026-06-18-agentic.md JetBrains, 80 paired tasks: 15.4 percent less code, 10.3 percent cheaper. https://blog.jetbrains.com/ai/2026/07/ponytail-skill-claude-tested/ An independent 480-build test found 44 percent less code with the same correctness, but on 5 of 24 jobs with unstated edge cases the leaner code broke on bad input, and stronger settings made it worse. https://github.com/Deepusleepy/ponytail-benchmark ### What this means for StateSync's own figures The figures in Measured effect above are full-session results: we ran the same 55 tasks ourselves on the TokenBench harness, the same model in both arms, counted the billed cost of every request, and scored quality on every task. Every task, including the ones StateSync lost, is listed at https://statesync.net/evidence. No independent group has tested StateSync yet. Where TokenBench and independent studies tested the same tool, they agreed on whether it saved money: RTK came out more expensive in all three tests of it, and Caveman cheaper in all three. That is good reason to expect StateSync's saving to hold under independent testing, though the exact figure will differ. ## Setup Install it and open a project: installed agents are found, the repository is mapped, and output compression is wired in. No config files to write, no settings to find, no per-agent integration to learn. It is built to work the moment it is installed. ## Privacy Your code, your prompts, your commands and your file contents never leave your machine. StateSync does its work locally, before anything is sent, and there is no StateSync service in the request path: the request goes straight to your own agent, on your own account. What StateSync itself receives is usage statistics: model and provider names, token counts and costs, tool and dictation counts, repository size bands, machine details and the plan tier of signed-in agents. The privacy policy lists all of it. What goes to the provider may still be used by that provider under its own terms; StateSync has no visibility into it and makes no claim about what the provider does with it. ## Pricing - Core: $8 a month, for one $20 a month agent plan - Pro: $20 a month, for a $100 a month agent plan - Max: $50 a month, for a $200 top tier plan like Claude Max or ChatGPT Pro - Ultra: $100 a month, for several top tier plans at once Every plan includes all of StateSync; plans differ by weekly usage allowance and the number of signed-in devices. Each starts with a free trial (14 days free for the first 50 places, 7 days after that). This is on top of the agent subscription you already pay for, which StateSync does not touch and does not resell. Each plan is priced to pay for itself: the saving, or the extra work your subscription gets done, is worth more than the plan costs. ## Pages - [StateSync: cut the bill, keep the performance](https://statesync.net/): StateSync is a desktop client that runs Claude Code, Codex, Cursor CLI, Copilot and five other coding agents on your own account. Across 55 real tasks it cut cost 46 percent and time 48 percent, with quality scored on every task and slightly higher on average. Nothing to configure. - [StateSync pricing: from $8 a month, with a free trial](https://statesync.net/pricing): StateSync plans run from $8 to $100 a month, on top of the coding agent subscription you already have. Every plan includes all of StateSync; plans differ by weekly usage allowance and the number of signed-in devices. - [Download StateSync for Windows, macOS and Linux](https://statesync.net/download): Install StateSync and run the coding agents you already pay for from one window. No sign-in needed to download, and nothing is charged until you pick a plan. - [StateSync changelog: every published version](https://statesync.net/changelog): Every published build of StateSync with its date, and what changed on Windows, macOS and Linux in each one. - [StateSync quickstart: install, open a project, start saving](https://statesync.net/quickstart): Install StateSync, sign in, open a project folder and pick an agent. Installed agents are found and the repository is mapped for you. Nothing to configure. - [StateSync, the app: every agent, one window](https://statesync.net/product): What StateSync looks like in use: local dictation, every coding agent in one window, switching agent mid-task, terminals, Spaces, and reviewing every change. - [StateSync evidence: every task in the TokenBench 55-task set](https://statesync.net/evidence): The figures behind StateSync's numbers: saving by kind of work, by task length, every one of the 55 tasks including the two StateSync lost, quality scores, where the saving comes from, and an estimate for your own spend by task size. - [StateSync support](https://statesync.net/support): How to reach the StateSync team, report a problem, or ask a question about your account. - [StateSync terms of service](https://statesync.net/terms): The terms that govern a StateSync account and subscription. - [StateSync privacy policy](https://statesync.net/privacy): What StateSync collects and what it never receives: your code, prompts and commands stay on your machine. - [StateSync security: how your API keys are handled](https://statesync.net/security): StateSync never receives your API keys. Bring your own key and use it exactly as you would with the provider directly: the key stays on your machine, and no StateSync server sits between you and your provider. - [StateSync known issues and current version by operating system](https://statesync.net/known-issues): Problems StateSync already knows about in early access, with a way round where there is one, and the current version available for Windows, macOS and Linux. - [StateSync documentation](https://statesync.net/docs): How to install StateSync, connect the coding agents you already pay for, and get the same work done for less. - [Install StateSync: StateSync documentation](https://statesync.net/docs/install): Download StateSync for Windows, macOS or Linux, and what happens the first time it opens. - [Your first project: StateSync documentation](https://statesync.net/docs/first-project): Open a folder, pick an agent, run a task, and read what changed. - [Agents: StateSync documentation](https://statesync.net/docs/agents): Which coding agents StateSync runs, how it finds the ones you already have, and how your own subscriptions and keys are used. - [Switching agent mid-task: StateSync documentation](https://statesync.net/docs/switching): Hit a limit halfway through a task? Move the conversation to another agent you pay for and carry on. - [Reviewing changes: StateSync documentation](https://statesync.net/docs/review): Read every file an agent touched, revert a change, and commit, without leaving the window. - [StateVoice dictation: StateSync documentation](https://statesync.net/docs/voice): Press a shortcut and talk. StateVoice transcribes on your own machine, so the recording never leaves your computer. - [Terminals and VS Code: StateSync documentation](https://statesync.net/docs/terminals): Terminal settings, and the VS Code extension that puts StateSync's chat and terminal beside your code. - [Settings: StateSync documentation](https://statesync.net/docs/settings): Every section of StateSync's settings, and what each one is for. - [What StateSync does: StateSync documentation](https://statesync.net/docs/saving): The 3 capabilities behind StateSync, and how to switch them off. - [What stays on your machine: StateSync documentation](https://statesync.net/docs/local): StateSync never receives your code, prompts, chats, commands or voice. What it does receive is counts. - [How to reduce Claude Code costs: StateSync documentation](https://statesync.net/docs/reduce-claude-code-costs): The free fixes that keep a Claude Code bill down, and what StateSync adds on top. - [How to reduce Codex CLI token usage: StateSync documentation](https://statesync.net/docs/reduce-codex-token-usage): The settings and habits that make each Codex message carry less, so a plan lasts longer. - [How to cut Cursor CLI costs: StateSync documentation](https://statesync.net/docs/cut-cursor-cli-costs): Where Cursor CLI spend comes from and the free ways to slow it down. - [StateSync and /compact: StateSync documentation](https://statesync.net/docs/statesync-vs-compact): Compaction is for when a task ends. StateSync works while one is running. When to use each. - [Tools that make coding agents cheaper, compared: StateSync documentation](https://statesync.net/docs/coding-agent-cost-tools-compared): The tools that cut coding agent spend, sorted by what they act on. - [Plans and usage: StateSync documentation](https://statesync.net/docs/plans): What each plan costs, how the weekly allowance works, and what happens when a week runs out. - [Signing in and devices: StateSync documentation](https://statesync.net/docs/devices): How sign-in works without a password, how a device is authorised, and how to remove one. - [Troubleshooting: StateSync documentation](https://statesync.net/docs/troubleshooting): What to check when an agent will not launch, a sign-in or an update will not complete, or dictation will not start. ## Benchmark https://github.com/Silverspine1/TokenBench We ran the tests ourselves using the TokenBench harness. The harness and the task set are public on GitHub. TokenBench 55-task set: 55 tasks across 8 repositories and 6 categories, Claude Sonnet 4.6 in both arms. StateSync was run on 11 September 2026 against the plain agent's averaged runs of the same tasks from 22 to 24 August 2026. No independent run has been published yet.