Rework Cost Check
What does one task really cost once limits, judges, validators and reviewers send work back?
Prefer a spreadsheet? Download the Excel version
Built by Sunil Mathew, co-authored with Claude
- What to bring
- One agent workflow. Token counts from a handful of traces, or estimates. How often each check or limit sends work back. Your peak tasks per minute and your tokens-per-minute quota.
- What to do
- List the paths that send work back, cost one round of each, set a cap and a final outcome for each, and add your per-task budget.
- What you leave with
- Two numbers — typical cost per task and worst case at peak — plus a Pass, Attention or Fail on five checks, and the next fix to make.
Your numbers stay in your browser. Nothing is sent or stored, and there is nothing to sign up for.
Four numbers about a normal task
Cost one round of each
Whatever sent the work back, a round costs the same three things: the steps that run again, the check that re-reads the new attempt, and the extra context the redo carries. Add only the paths your workflow actually has. Zero is a real answer; blank is not.
-
Transport refusal
429 rate limit, 503 unavailable, or a timeout
The model or tool API did not answer. This is a failure to reach the service, not a problem with the output.
Bites at peak only Context grows no Attempt billed usually not
1 How often does it send work back?
2 What does one round cost?
=One round —tokens
3 When does it stop, and who owns that?
No cap means the worst case has no ceiling, and the tool will say Unbounded.
-
Judge / critic rejection
an LLM-as-judge fails the draft
A model or a rubric reads the finished output, decides it is not good enough, and sends it back with feedback.
Bites every run Context grows yes Attempt billed yes
1 How often does it send work back?
2 What does one round cost?
=One round —tokens
3 When does it stop, and who owns that?
No cap means the worst case has no ceiling, and the tool will say Unbounded.
-
Validation failure
bad JSON, schema mismatch, guardrail trip
The output did not parse, did not match the schema, or tripped a guardrail, so the model has to produce it again.
Bites every run Context grows usually Attempt billed yes
1 How often does it send work back?
2 What does one round cost?
=One round —tokens
3 When does it stop, and who owns that?
No cap means the worst case has no ceiling, and the tool will say Unbounded.
-
Tool error
not found, permission denied, bad arguments
A tool the agent called returned an error: not found, permission denied, bad arguments. The agent has to re-plan or try again.
Bites every run Context grows yes Attempt billed yes
1 How often does it send work back?
2 What does one round cost?
=One round —tokens
3 When does it stop, and who owns that?
No cap means the worst case has no ceiling, and the tool will say Unbounded.
-
Human rejection
a maker-checker sends it back
A person reviewing the work sends it back. Their reading costs no tokens; the redo does.
Bites every run Context grows yes Attempt billed yes
1 How often does it send work back?
2 What does one round cost?
=One round —tokens
3 When does it stop, and who owns that?
No cap means the worst case has no ceiling, and the tool will say Unbounded.
-
name it yourself
Anything else in your system that makes model work happen a second time.
Bites you decide Context grows you decide Attempt billed
1 How often does it send work back?
2 What does one round cost?
=One round —tokens
3 When does it stop, and who owns that?
No cap means the worst case has no ceiling, and the tool will say Unbounded.
-
name it yourself
Anything else in your system that makes model work happen a second time.
Bites you decide Context grows you decide Attempt billed
1 How often does it send work back?
2 What does one round cost?
=One round —tokens
3 When does it stop, and who owns that?
No cap means the worst case has no ceiling, and the tool will say Unbounded.
Add a path only if it makes the model do work a second time. That is what this tool counts. A send-back that spends no further tokens — a person editing the output themselves, say — belongs in your process notes, not here.
All seven paths are in your workflow.
One budget for the whole task
No price entered yet, so this stays in tokens. The scheduling example leaves the price blank on purpose: it is the one figure nobody else can supply, because it depends on your provider, your tier and your own mix of input and output tokens. Put your number in Your price per million tokens above and the cost per task appears here. At that rate, a typical task costs —, and a thousand of them cost —. That is the everyday figure, not the worst case.
For scale: the same task at frontier rates
Your typical cost per task, priced at what three frontier models charge. It is a range because input and output are priced differently: the low figure is all input, the high figure is all output, and real work sits between them, usually nearer the low end. Checked 28 September 2026.
| Model | Input per 1M | Output per 1M | Your typical task | Worth knowing |
|---|---|---|---|---|
| Claude Opus 5.5 Anthropic | 4.00 | 20.00 | — | US-only inference costs 1.1×. Fast mode costs 2×. |
| GPT-6 Astra OpenAI | 10.00 | 50.00 | — | Long context costs $20 in and $75 out, so a growing redo can double this. |
| Gemini 3.1 Pro Google | 2.00 | 12.00 | — | Prompts over 200k cost $4 in and $18 out. |
Figures are US dollars, which is how all three publish. Rates change, and two of these carry a condition that doubles them, so check the provider before you quote one: Anthropic · OpenAI · Google .
Pass, attention or fail
Each check reads your figures and says where you stand, then what to do about it. Work down the list; the one to fix first is called out at the bottom.
- Pass
- Nothing to do here.
- Attention
- It works today, and something about it will surprise you. Worth fixing.
- Fail
- This is the one that stops you scaling. Fix it first.
- Not answered
- The check needs a figure you have not entered yet.
- Not answered
1. Every path that sends work back is listed
A path you have not listed is spend you cannot see. Most teams count their retries and miss the judge entirely, because a judge rejection looks like normal operation rather than a failure.
No path is marked as present yet.
Do this
- Not answered
2. Each round is costed as repeated, review and growth
A round is three things and teams usually count only the first. The review call and the growing context are what make a judge loop expensive, and neither appears in a retry counter.
No path is marked as present yet.
Do this
- Not answered
3. Every loop has a limit, a decided outcome, and an owner
A loop with no cap has no worst case. This is the check that turns "it depends" into a number, and an uncapped path is the single most common finding here.
No path is marked as present yet.
Do this
- Not answered
4. Both numbers fit: what you pay daily, and your busiest minute
One number is not enough, because the two ways this hurts you are different. The typical cost is what you pay every day, and it decides whether the product makes money. The worst case at peak decides whether one bad minute takes the whole system down. A workflow can be fine on one and fatal on the other.
Peak tasks per minute and the quota are needed for this one.
Do this
- Not answered
5. One token budget for the whole task, not one per call
A task that goes round three times spends three times. A budget checked on each call sees the attempts one at a time and never sees the total, so it cannot stop a task that is running away. To do that, every round has to count against one limit that belongs to the task.
Choose how your budget is enforced.
Do this
Your next step
Where each one lives, in your own system
Tokens for one clean run
- Open your tracing tool, for example Langfuse, LangSmith or Arize Phoenix, and filter to tasks that finished with nothing sent back.
- Take 20 to 50 of them and use the median of total input plus output tokens.
- No tracing? Every model API response carries a usage field with input and output token counts. Log it with a task id for a day, then add up the calls per task.
Always-on checks
- In the same traces, find the judge, validator or guardrail calls that ran on the first attempt, and take their median tokens.
- If the check is code rather than a model, a JSON schema check for instance, its token cost is 0.
Average rounds per task, per path
- Take a normal week and count send-backs divided by tasks.
- Judge: failed verdicts in your judge or eval logs.
- Validation: parse errors, schema failures and guardrail blocks that triggered a redo.
- Tool error: tool calls that returned an error and made the agent re-plan or retry.
- Human: items sent back divided by items reviewed, from your review queue.
- Transport refusal: 429, 503 and timeout counts from your client or gateway logs. Use a normal hour, which is usually close to 0. The peak matters for the worst case, not for the typical cost.
Work repeated
- Find out what your code does after a send-back: does it restart the task, redo the step, or retry the call only? The answer is in your orchestrator or retry code.
- Then add up, from traces, the tokens of the steps that run again.
- Hidden repeat — model SDK retries. Many SDKs retry some errors automatically, often twice by default. Those stack on top of any retries your own code makes.
- Hidden repeat — graph re-entry. A graph framework looping back to an earlier node repeats every node in between.
Review call
- The judge's or checker's token count on a redo, from the same trace.
- A judge usually reads the whole output plus a rubric, so its input is large.
Context growth
- In one trace with a send-back, compare the input tokens of the first attempt with those of the redo. The difference is the growth: the note, the error or the rejected draft carried forward.
- If each round keeps all earlier notes, growth rises every round. Enter the average growth, or the last round’s growth for a cautious number.
Max rounds per task
- Your retry settings, the loop limit in your orchestrator (graph frameworks usually have a recursion or step limit), and your SDK’s retry setting.
- If a path has no explicit limit in your code, leave the cap blank. The tool will show the worst case as unbounded, and that is the finding.
Peak tasks per minute
- The busiest minute in the last 30 days from your app metrics. Use the 99th-percentile minute, not the average.
- Pre-launch: use your launch estimate and mark it as an estimate.
Tokens-per-minute quota
- The limits page in your model provider’s console, for your account tier and model.
- Most providers also return rate-limit headers on each response showing the limit and what remains.
- Note who else draws on the same key or organisation. The quota is usually shared.
Per-task budget
- Your own code or gateway config.
- If nothing stops a task at a token total, the answer is "Not enforced".
Price per million tokens
- Your provider’s pricing page. Input and output are priced differently.
- Blend them by your own ratio of input to output tokens, taken from traces.
No data yet?
- Use estimates and label them as estimates.
- The check still shows which paths and which caps are missing. That is often the more useful finding.
Where a product or a setting is named above, check your provider's docs. Names change.
Every layer did the sensible thing. The system still went down.
At Google, Sunil saw an outage pattern that began with every layer doing the sensible thing: rate limits refused calls, callers retried, and the retries brought the system down. Agents add more ways to send work back — judges, validators, tool errors and human reviewers — and unlike a rate-limit refusal, those are billed in full and happen on every run.
A scheduling agent with a judge on the draft
A scheduling agent drafts meeting invitations. Nothing in it is broken, nobody has misconfigured anything, and the numbers below are what it costs on an ordinary Tuesday.
-
A clean run is 43,300 tokens
Read the request, check calendars, draft the invitation, send it. If nothing sends the work back, that is the whole bill. This is the number most cost models stop at.
-
A judge reads every draft: 11,500 tokens, every time
Before anything is sent, a model checks the draft against a rubric. It reads the whole output, so its input is large. It runs whether the draft passes or fails, which makes it an always-on check rather than a send-back. Running total: 54,800.
-
The judge sends about two drafts back per task
Each of those rounds costs three things: the drafting step runs again (8,900), the judge re-reads the new draft (11,500), and the critique is carried into the redo so the input is bigger than last time (1,500). One round is 21,900 tokens, and there are two of them.
-
So a normal task costs 98,600 tokens, not 43,300
54,800 + 2 × 21,900. That is 2.28× the clean run, and it is the everyday number — the one that decides whether the product makes money. At 200 tasks a minute it asks for 19,720,000 tokens a minute, which is 197.2% of a 10,000,000 quota shared with every other workflow on the same key. The workflow is over its quota before anything has gone wrong.
-
The rate limit is capped and quiet, until it is not
Transport refusals average zero rounds on a normal day: that is a measurement, not a blank. They bite at peak. When they do, this code restarts the task, so 16,700 tokens of completed work are paid for again on every attempt.
-
At the cap, one task costs 170,600 tokens
54,800, plus three judge rounds at 21,900, plus three transport rounds at 16,700. That is 3.94× a clean run, and across 200 tasks a minute it is 34,120,000 tokens a minute, or 341.2% of the quota.
| Reading | Figure | What it means |
|---|---|---|
| One judge round | 21,900 | tokens · 8,900 repeated + 11,500 review + 1,500 growth |
| Typical cost per task | 98,600 | tokens · 2.28× a clean run, on a normal day |
| Typical demand at peak | 19,720,000 | tokens a minute · 197.2% of the quota |
| Worst case per task | 170,600 | tokens · 3.94× a clean run, every loop at its cap |
| Worst case at peak | 34,120,000 | tokens a minute · 341.2% of the quota |
The row that matters is the second one, not the last. A check that only counted rate limits would have reported 8,660,000 a minute, or 87% of quota, and called it survivable. The judge loop is what puts it at 197%, and it does that every single day.
- Pass 1 Every path that sends work back is listed
- Pass 2 Each round is costed as repeated, review and growth
- Pass 3 Every loop has a limit, a decided outcome, and an owner
- Fail 4 Both numbers fit: what you pay daily, and your busiest minute
- Pass 5 One token budget for the whole task, not one per call
Check 4 fails because normal load is already over the quota. Everything else passes: the paths are listed and costed, both are capped with a decided outcome and a named owner, and the 120,000 budget sits between the typical cost and the worst case, so it binds.
Stopping has a cost you can list. Not stopping has no ceiling.
Download the Excel version All resources
Published 28 September 2026 · Content CC BY 4.0 · Code MIT · © The Living Craft