What we measured.

55 real coding tasks across 8 repositories, each run through StateSync and compared with a plain agent on the same model. Every task is on this page, wins and losses.

The 55 are the tasks in TokenBench's 73 where the plain agent made more than 8 model requests. The shortest tasks are left out: they are not the work people hand a coding agent, and the saving on them is the smallest. StateSync ran on 11 September 2026, and the plain agent's runs of the same tasks are from 22 to 24 August 2026.

TokenBench harness

-46%
cost
-48%
time to finish
-45%
tokens billed
-29%
requests to the model

Cheaper on 53 of 55, faster on 52.

$26.44 of work became $14.25, and 149 min of waiting became 77 min. Read the other way, the same money buys 86% more work.

By kind of work

Every kind of task saved. Building features in steps saved the most, and adding a feature the least. Kinds with fewer than 3 tasks are left out.

Building features in steps17 tasks-50%
Restructuring8 tasks-49%
Interface work5 tasks-47%
Bug fixes20 tasks-41%
Adding a feature4 tasks-36%

By how long the task took

Tasks under two minutes saved the least. Tasks that short are not what people hand a coding agent, which is why the benchmark leaves out its own shortest tasks. On tasks of two minutes or more, StateSync cut cost by 50 percent.

Under 2 minutes17 of 18 cheaper-27%
2 to 4 minutes27 of 27 cheaper-50%
4 minutes and up9 of 10 cheaper-50%

Every task

Cheaper on 53 of 55. The 2 that cost more are on the chart too.

Sorted largest saving first. The bar on the right is the task that cost 41% more.

Quality

Quality in our TokenBench run is measured with the harness's standardised scorer. StateSync scored higher on 15 tasks, level on 34 and lower on 6, and solved 41 of 55 against the plain agent's 40 on average across its runs.

15
scored higher
34
level
6
scored lower

Read the scorer on GitHub

Where the saving comes from

Most of the saving is requests that never had to happen, rather than requests that got cheaper. Read on the tasks where neither agent handed work to a subagent.

Every 100 tokens the plain agent was billed for, on the 13 tasks where neither agent handed work to a subagent

70525
still billed
still billed
smaller requests
saved by smaller requests
fewer requests
saved by fewer requests

Watch a task run.

The same tasks, recorded end to end with and without StateSync, both agents on Claude Sonnet 5. These are single runs rather than the pooled figures above, so watch them for what the work looks like, not for a number.

Both agents were given the same prompt

One data table

The task and user tables moved onto one shared table.

Claude Sonnet 5 on both sides, the same starting code and the same settings.

  1. 1Move the Tasks and Users tables onto one shared, typed data-table module.
  2. 2Change no behaviour: search, filters, sorting, bulk delete and row actions all keep working.
  3. 3Add column visibility per table that survives a reload, with a Reset columns action.
  4. 4Keep the existing tests green and add tests for the new module.

One data tableThe task and user tables moved onto one shared table.

Estimate your own saving.

Say what you spend and how long your work runs. The projection comes from the results above.

How you pay for coding agents
15 minutes

How long a task takes your agent without StateSync.

Projected past the longest measured task (5 minutes).

Roughly, on the same subscription

+89%

more work before you reach a limit. Your bill does not change.

Plan for that
Max, $50 a month
Same work would cost
$106 of subscription

An estimate from the measured tasks, not a quote. Your models, your work and your repositories will give their own numbers.

See the plans

StateSync ran each task once, against the plain agent's earlier runs of the same task, averaged. That is enough for a total but not for any one task: a single run of a non-deterministic agent can land well either side of its own average. One model, run by us. Other models, tasks and repositories will give other numbers.