A long agent task succeeds only if its steps survive in sequence, so a few points of per-step reliability you cannot see on a sprint benchmark decide the whole run. Set two models and your task length. Then get the routing call: cost per finished task, once you price in the frontier's premium.
One structural lens at a time, with the sourcing shown. Essays and runnable instruments, free.
Subscribe · freeThe exact exponent is a diagnostic, not a law of physics: real agents recover, which simply raises the effective per-step reliability you set above, and a finished run can retry rather than restart whole. The point survives the caveat. Reliability gaps that round to nothing on a short task go nonlinear once the work has to survive many handoffs.