What we measured.
55 real coding tasks across 8 repositories, each run through StateSync and compared with a plain agent on the same model. Every task is on this page, wins and losses.
The 55 are the tasks in TokenBench's 73 where the plain agent made more than 8 model requests. The shortest tasks are left out: they are not the work people hand a coding agent, and the saving on them is the smallest. StateSync ran on 11 September 2026, and the plain agent's runs of the same tasks are from 22 to 24 August 2026.
- -46%
- cost
- -48%
- time to finish
- -45%
- tokens billed
- -29%
- requests to the model
Cheaper on 53 of 55, faster on 52.
$26.44 of work became $14.25, and 149 min of waiting became 77 min. Read the other way, the same money buys 86% more work.
By kind of work
Every kind of task saved. Building features in steps saved the most, and adding a feature the least. Kinds with fewer than 3 tasks are left out.
By how long the task took
Tasks under two minutes saved the least. Tasks that short are not what people hand a coding agent, which is why the benchmark leaves out its own shortest tasks. On tasks of two minutes or more, StateSync cut cost by 50 percent.
Every task
Cheaper on 53 of 55. The 2 that cost more are on the chart too.
Sorted largest saving first. The bar on the right is the task that cost 41% more.
Quality
Quality in our TokenBench run is measured with the harness's standardised scorer. StateSync scored higher on 15 tasks, level on 34 and lower on 6, and solved 41 of 55 against the plain agent's 40 on average across its runs.
- 15
- scored higher
- 34
- level
- 6
- scored lower
Where the saving comes from
Most of the saving is requests that never had to happen, rather than requests that got cheaper. Read on the tasks where neither agent handed work to a subagent.
Every 100 tokens the plain agent was billed for, on the 13 tasks where neither agent handed work to a subagent
- still billed
- still billed
- smaller requests
- saved by smaller requests
- fewer requests
- saved by fewer requests
Watch a task run.
The same tasks, recorded end to end with and without StateSync, both agents on Claude Sonnet 5. These are single runs rather than the pooled figures above, so watch them for what the work looks like, not for a number.
Both agents were given the same prompt
One data table
The task and user tables moved onto one shared table.
Claude Sonnet 5 on both sides, the same starting code and the same settings.
- 1Move the Tasks and Users tables onto one shared, typed data-table module.
- 2Change no behaviour: search, filters, sorting, bulk delete and row actions all keep working.
- 3Add column visibility per table that survives a reload, with a Reset columns action.
- 4Keep the existing tests green and add tests for the new module.
One data tableThe task and user tables moved onto one shared table.
Estimate your own saving.
Say what you spend and how long your work runs. The projection comes from the results above.
How long a task takes your agent without StateSync.
Projected past the longest measured task (5 minutes).
Roughly, on the same subscription
+89%
more work before you reach a limit. Your bill does not change.
- Plan for that
- Max, $50 a month
- Same work would cost
- $106 of subscription
An estimate from the measured tasks, not a quote. Your models, your work and your repositories will give their own numbers.
See the plansStateSync ran each task once, against the plain agent's earlier runs of the same task, averaged. That is enough for a total but not for any one task: a single run of a non-deterministic agent can land well either side of its own average. One model, run by us. Other models, tasks and repositories will give other numbers.