The Durability CurveInstrument · Marathon Gap
The Marathon Calculator

Where the sprint score stops predicting the marathon

A long agent task succeeds only if its steps survive in sequence, so a few points of per-step reliability you cannot see on a sprint benchmark decide the whole run. Set two models and your task length. Then get the routing call: cost per finished task, once you price in the frontier's premium.

Your model · per-step reliability 95.0%
The share of steps it clears without a wrong turn you have to undo.
Don't know your number? Calibrate from a real run
Of my recent runs, % finished, over steps each.
That implies 97.6% per step. Set above.
Frontier model · per-step reliability 98.0%
A 3.0-point edge over your model. The gap a sprint benchmark barely registers.
Task length · steps in the chain 40steps
Count the chain, not the prompt. A multi-hour run is hundreds of steps; 40 is a conservative marathon.
Tasks finished vs chain length your modelfrontier
Your model finishes
13%
of tasks at 40 steps
Frontier finishes
45%
of the same tasks
The routing call
Marathon-bound
Tokens per step 50k
Frontier price per token ×6

Get the next instrument.

One structural lens at a time, with the sourcing shown. Essays and runnable instruments, free.

Subscribe · free
Runs entirely in your browser. Nothing leaves the page.
← Read the essay: the marathon gap

The exact exponent is a diagnostic, not a law of physics: real agents recover, which simply raises the effective per-step reliability you set above, and a finished run can retry rather than restart whole. The point survives the caveat. Reliability gaps that round to nothing on a short task go nonlinear once the work has to survive many handoffs.