Resources / agentic system design

Rework Cost Check

What does one task really cost once limits, judges, validators and reviewers send work back?

Run the Rework Cost Check 

Prefer a spreadsheet? Download the Excel version

Built by Sunil Mathew, co-authored with Claude

What to bring
One agent workflow. Token counts from a handful of traces, or estimates. How often each check or limit sends work back. Your peak tasks per minute and your tokens-per-minute quota.
What to do
List the paths that send work back, cost one round of each, set a cap and a final outcome for each, and add your per-task budget.
What you leave with
Two numbers — typical cost per task and worst case at peak — plus a Pass, Attention or Fail on five checks, and the next fix to make.

Your numbers stay in your browser. Nothing is sent or stored, and there is nothing to sign up for.

Typical cost per task — tokens — Unit economics: what one task costs on a normal day, counting the rounds that get sent back.
Worst case at peak — of your quota — Availability: the share of your tokens-per-minute quota this workflow would ask for if every loop ran to its cap during your busiest minute.
A The workflow

Four numbers about a normal task

B Paths that send work back

Cost one round of each

Whatever sent the work back, a round costs the same three things: the steps that run again, the check that re-reads the new attempt, and the extra context the redo carries. Add only the paths your workflow actually has. Zero is a real answer; blank is not.

Add a path only if it makes the model do work a second time. That is what this tool counts. A send-back that spends no further tokens — a person editing the output themselves, say — belongs in your process notes, not here.

C Budget

One budget for the whole task

No price entered yet, so this stays in tokens. The scheduling example leaves the price blank on purpose: it is the one figure nobody else can supply, because it depends on your provider, your tier and your own mix of input and output tokens. Put your number in Your price per million tokens above and the cost per task appears here. At that rate, a typical task costs —, and a thousand of them cost —. That is the everyday figure, not the worst case.

For scale: the same task at frontier rates

Your typical cost per task, priced at what three frontier models charge. It is a range because input and output are priced differently: the low figure is all input, the high figure is all output, and real work sits between them, usually nearer the low end. Checked 28 September 2026.

Model Input
per 1M
Output
per 1M
Your typical task Worth knowing
Claude Opus 5.5 Anthropic 4.00 20.00 — US-only inference costs 1.1×. Fast mode costs 2×.
GPT-6 Astra OpenAI 10.00 50.00 — Long context costs $20 in and $75 out, so a growing redo can double this.
Gemini 3.1 Pro Google 2.00 12.00 — Prompts over 200k cost $4 in and $18 out.

Figures are US dollars, which is how all three publish. Rates change, and two of these carry a condition that doubles them, so check the provider before you quote one: Anthropic · OpenAI · Google .

The five checks

Pass, attention or fail

Each check reads your figures and says where you stand, then what to do about it. Work down the list; the one to fix first is called out at the bottom.

Pass
Nothing to do here.
Attention
It works today, and something about it will surprise you. Worth fixing.
Fail
This is the one that stops you scaling. Fix it first.
Not answered
The check needs a figure you have not entered yet.
  1. Not answered

    1. Every path that sends work back is listed

    A path you have not listed is spend you cannot see. Most teams count their retries and miss the judge entirely, because a judge rejection looks like normal operation rather than a failure.

    No path is marked as present yet.

  2. Not answered

    2. Each round is costed as repeated, review and growth

    A round is three things and teams usually count only the first. The review call and the growing context are what make a judge loop expensive, and neither appears in a retry counter.

    No path is marked as present yet.

  3. Not answered

    3. Every loop has a limit, a decided outcome, and an owner

    A loop with no cap has no worst case. This is the check that turns "it depends" into a number, and an uncapped path is the single most common finding here.

    No path is marked as present yet.

  4. Not answered

    4. Both numbers fit: what you pay daily, and your busiest minute

    One number is not enough, because the two ways this hurts you are different. The typical cost is what you pay every day, and it decides whether the product makes money. The worst case at peak decides whether one bad minute takes the whole system down. A workflow can be fine on one and fatal on the other.

    Peak tasks per minute and the quota are needed for this one.

  5. Not answered

    5. One token budget for the whole task, not one per call

    A task that goes round three times spends three times. A budget checked on each call sees the attempts one at a time and never sees the total, so it cannot stop a task that is running away. To do that, every round has to count against one limit that belongs to the task.

    Choose how your budget is enforced.

How to find these numbers

Where each one lives, in your own system

Tokens for one clean run
  • Open your tracing tool, for example Langfuse, LangSmith or Arize Phoenix, and filter to tasks that finished with nothing sent back.
  • Take 20 to 50 of them and use the median of total input plus output tokens.
  • No tracing? Every model API response carries a usage field with input and output token counts. Log it with a task id for a day, then add up the calls per task.
Always-on checks
  • In the same traces, find the judge, validator or guardrail calls that ran on the first attempt, and take their median tokens.
  • If the check is code rather than a model, a JSON schema check for instance, its token cost is 0.
Average rounds per task, per path
  • Take a normal week and count send-backs divided by tasks.
  • Judge: failed verdicts in your judge or eval logs.
  • Validation: parse errors, schema failures and guardrail blocks that triggered a redo.
  • Tool error: tool calls that returned an error and made the agent re-plan or retry.
  • Human: items sent back divided by items reviewed, from your review queue.
  • Transport refusal: 429, 503 and timeout counts from your client or gateway logs. Use a normal hour, which is usually close to 0. The peak matters for the worst case, not for the typical cost.
Work repeated
  • Find out what your code does after a send-back: does it restart the task, redo the step, or retry the call only? The answer is in your orchestrator or retry code.
  • Then add up, from traces, the tokens of the steps that run again.
  • Hidden repeat — model SDK retries. Many SDKs retry some errors automatically, often twice by default. Those stack on top of any retries your own code makes.
  • Hidden repeat — graph re-entry. A graph framework looping back to an earlier node repeats every node in between.
Review call
  • The judge's or checker's token count on a redo, from the same trace.
  • A judge usually reads the whole output plus a rubric, so its input is large.
Context growth
  • In one trace with a send-back, compare the input tokens of the first attempt with those of the redo. The difference is the growth: the note, the error or the rejected draft carried forward.
  • If each round keeps all earlier notes, growth rises every round. Enter the average growth, or the last round’s growth for a cautious number.
Max rounds per task
  • Your retry settings, the loop limit in your orchestrator (graph frameworks usually have a recursion or step limit), and your SDK’s retry setting.
  • If a path has no explicit limit in your code, leave the cap blank. The tool will show the worst case as unbounded, and that is the finding.
Peak tasks per minute
  • The busiest minute in the last 30 days from your app metrics. Use the 99th-percentile minute, not the average.
  • Pre-launch: use your launch estimate and mark it as an estimate.
Tokens-per-minute quota
  • The limits page in your model provider’s console, for your account tier and model.
  • Most providers also return rate-limit headers on each response showing the limit and what remains.
  • Note who else draws on the same key or organisation. The quota is usually shared.
Per-task budget
  • Your own code or gateway config.
  • If nothing stops a task at a token total, the answer is "Not enforced".
Price per million tokens
  • Your provider’s pricing page. Input and output are priced differently.
  • Blend them by your own ratio of input to output tokens, taken from traces.
No data yet?
  • Use estimates and label them as estimates.
  • The check still shows which paths and which caps are missing. That is often the more useful finding.

Where a product or a setting is named above, check your provider's docs. Names change.

Why this exists

Every layer did the sensible thing. The system still went down.

At Google, Sunil saw an outage pattern that began with every layer doing the sensible thing: rate limits refused calls, callers retried, and the retries brought the system down. Agents add more ways to send work back — judges, validators, tool errors and human reviewers — and unlike a rate-limit refusal, those are billed in full and happen on every run.

Worked example

A scheduling agent with a judge on the draft

A scheduling agent drafts meeting invitations. Nothing in it is broken, nobody has misconfigured anything, and the numbers below are what it costs on an ordinary Tuesday.

  1. A clean run is 43,300 tokens

    Read the request, check calendars, draft the invitation, send it. If nothing sends the work back, that is the whole bill. This is the number most cost models stop at.

  2. A judge reads every draft: 11,500 tokens, every time

    Before anything is sent, a model checks the draft against a rubric. It reads the whole output, so its input is large. It runs whether the draft passes or fails, which makes it an always-on check rather than a send-back. Running total: 54,800.

  3. The judge sends about two drafts back per task

    Each of those rounds costs three things: the drafting step runs again (8,900), the judge re-reads the new draft (11,500), and the critique is carried into the redo so the input is bigger than last time (1,500). One round is 21,900 tokens, and there are two of them.

  4. So a normal task costs 98,600 tokens, not 43,300

    54,800 + 2 × 21,900. That is 2.28× the clean run, and it is the everyday number — the one that decides whether the product makes money. At 200 tasks a minute it asks for 19,720,000 tokens a minute, which is 197.2% of a 10,000,000 quota shared with every other workflow on the same key. The workflow is over its quota before anything has gone wrong.

  5. The rate limit is capped and quiet, until it is not

    Transport refusals average zero rounds on a normal day: that is a measurement, not a blank. They bite at peak. When they do, this code restarts the task, so 16,700 tokens of completed work are paid for again on every attempt.

  6. At the cap, one task costs 170,600 tokens

    54,800, plus three judge rounds at 21,900, plus three transport rounds at 16,700. That is 3.94× a clean run, and across 200 tasks a minute it is 34,120,000 tokens a minute, or 341.2% of the quota.

ReadingFigureWhat it means
One judge round21,900tokens · 8,900 repeated + 11,500 review + 1,500 growth
Typical cost per task98,600tokens · 2.28× a clean run, on a normal day
Typical demand at peak19,720,000tokens a minute · 197.2% of the quota
Worst case per task170,600tokens · 3.94× a clean run, every loop at its cap
Worst case at peak34,120,000tokens a minute · 341.2% of the quota

The row that matters is the second one, not the last. A check that only counted rate limits would have reported 8,660,000 a minute, or 87% of quota, and called it survivable. The judge loop is what puts it at 197%, and it does that every single day.

  • Pass 1 Every path that sends work back is listed
  • Pass 2 Each round is costed as repeated, review and growth
  • Pass 3 Every loop has a limit, a decided outcome, and an owner
  • Fail 4 Both numbers fit: what you pay daily, and your busiest minute
  • Pass 5 One token budget for the whole task, not one per call

Check 4 fails because normal load is already over the quota. Everything else passes: the paths are listed and costed, both are capped with a decided outcome and a named owner, and the 120,000 budget sits between the typical cost and the worst case, so it binds.

Closing

Stopping has a cost you can list. Not stopping has no ceiling.

Download the Excel version  All resources 

Published 28 September 2026 · Content CC BY 4.0 · Code MIT · © The Living Craft