Resources/ agentic system design 07 · The Agent Failure Triage Quiz
Tell a failed action from an unanswered one
One illustrative incident, worked one piece of evidence at a time: a purchase order that timed out, and a supplier that received it twice.
10 concept questions · 2 kit questions · 4 options each ·
answers stay in this browser
Built by Sunil Mathew, co-authored with Claude
01 What to bring
The list of tools your agent can call that change something outside it. At least one of them moves money or places an order.
02 What to do
Answer each question before you read the feedback. Every answer explains itself, including the wrong ones, and says what that choice would actually cause.
03 What you leave with
A recovery step you can name for each of your own write tools, and a list of the takeaways you missed, in your own order.
04 Read this first
Two things sit behind these questions. Episode 4 is where the procurement incident comes from. The Agent Failure Triage Kit is where the three failure classes and the order of the triage questions are written down.
Agentic system design, Episode 4 — The episode is on LinkedIn and is not linked here yet. The kit covers the same ground in writing.
The Agent Failure Triage Kit — the three failure
classes, the four triage questions and the response playbook, in writing.
Illustrative case · not a real client, employer or event
One order, two purchase orders, no error anywhere
A procurement agent at a mid-size manufacturer raises purchase orders when stock runs low. A buyer approves the order; the agent places it.
It places every order through one tool, create_purchase_order. That tool sits on an MCP server the platform team runs. MCP (Model Context Protocol) is the standard way a client offers tools to a model. The MCP server calls the supplier’s ordering API.
On Tuesday at 09:14 the agent orders 1,200 steel brackets at ₹8,40,000. By 09:20 the supplier’s portal shows that same order twice, PO-44812 and PO-44813, ₹16,80,000 committed.
The agent’s own logs say no order was placed. Nothing raised an error. You are on call.
This incident is illustrative, and is not a real client, employer or event. You are the engineer on call, and each question hands you one new piece of evidence.
0/10 concept
Not startedAnswer question one. Nothing is stored and nothing is sent.
0/12
05 The triage
Commit to an answer, then read the evidence back.
Your answers stay in this browser. Nothing is stored, nothing is sent, and closing the
tab clears them.
01Diagnose
The agent’s client log
09:14:02.114 call create_purchase_order sku=BRK-19 qty=1200 value=840000
09:14:32.115 call create_purchase_order no response after 30000 ms
The call to create_purchase_order timed out after 30 seconds. No error came back. What do you know for certain?
Hint
Read the second log line again. Which side of the call does it describe?
A Wrong. A timeout tells you the response did not arrive. It says nothing about what the receiver did with the request.
B Wrong. A rejection is an answer, and it would have come back as an error. Silence is not a rejection.
C Correct. A timeout is missing information, not a failure. Every step after it has to be correct whether the order exists or not.
D Wrong. That is one cause of silence, not something the silence proves. A reply that is merely slow looks identical at 30 seconds.
A timeout on a write is an unknown outcome, not a failed one. Takeaway 1.
02Wrong lever
The incident channel, 09:41
"Raise max attempts from 3 to 5 and add exponential backoff: 2s, 4s, 8s, 16s. That is why we got two orders."
The team proposes five attempts with exponential backoff instead of three. What does that fail to fix?
Hint
Ask what the extra attempts know that the first one did not.
A Correct. Retry policy and recovery design are two separate decisions. Attempts and delays answer how often to try; nothing in them establishes what the first try did.
B Wrong. Spacing changes when the second call arrives, not whether it is a second order. The supplier accepted both calls as they came.
C Wrong, and it is the same lever again. A longer gap in front of an unknown outcome still leaves the outcome unknown.
D Wrong. The expensive case is a supplier that is available and slow. Against a slow success, more attempts means more orders.
A retry policy answers how often to try. It never answers whether the first attempt landed. Takeaway 2.
03Diagnose
The MCP server’s log, from the platform team
09:14:02.140 tools/call create_purchase_order forwarded POST /v2/orders
09:14:34.210 tools/call create_purchase_order attempt 2 forwarded
09:14:37.902 supplier responded 201 Created order=PO-44812
09:14:41.377 supplier responded 201 Created order=PO-44813
The agent stopped waiting at 30 seconds. The supplier created PO-44812 at 35 seconds. What did the agent’s silence tell you about the hops behind it?
Hint
Compare two clocks: when the client stopped waiting, and when the supplier answered.
A Wrong, and the log says so. PO-44812 came from attempt one, five seconds after the agent had given up on it.
B Correct. MCP adds hops. Silence at the client says nothing about what happened at the MCP server or at the supplier.
C Wrong about the diagnosis. The server had no error to report, because it was still waiting too. An invented error would have made a real order look like a failure.
D Wrong, and it gives up one step early. The agent cannot learn it from silence. It can ask, which is question six.
A client timeout bounds your own patience. It does not stop the work downstream. Takeaway 3.
04Spot the claim
The tool definition, from the server’s tools/list response
{ "name": "create_purchase_order",
"annotations": { "title": "Create purchase order", "idempotentHint": true } }
The supplier’s API documentation has no idempotency header and no duplicate suppression.
The tool is annotated idempotentHint: true. Which part of this setup is declared but not enforced?
Hint
Ask who reads the annotation, and who places the order.
A Correct. idempotentHint declares; it does not enforce. The specification also tells clients to treat tool annotations as untrusted unless the server is trusted, and the hint defaults to false when it is absent.
B Wrong, and this is the belief that produced PO-44813. The hint travels with the tool definition. The supplier never sees it.
C Wrong target. The timeout is a client choice and it was honoured exactly. The unenforced claim here is about what a repeat call does.
D Wrong. A 201 with an order number is the confirmation you were missing all morning. The unsupported claim is the annotation above it.
An annotation is metadata about a tool, not behaviour at the system that acts. Takeaway 4.
05Next check
Four proposals on the board, 11:20
The team agrees to add an idempotency key: a value that tells the receiver two requests are the same action, so it applies the action once. Four versions are proposed.
Which of these keys protects the agent on a retry?
Hint
For each one, ask when the key is created and who keeps it.
A Wrong. A new value on each call describes the call, so two calls are two actions. This is the reversal waiting to happen, and it is question eight.
B Correct, and all four properties are doing work. Created before the first attempt, held outside the model, reused rather than regenerated, and passed downstream to the system that actually applies it.
C Closer, and it fails two ways. Two orders the team means to place twice collapse into one, and any re-planned argument produces a new identity for the same intended order.
D Wrong, and it needs the one thing you do not have. The retry exists because no response arrived, so there is no order number to send.
A key names the action you intended, and it exists before the first attempt. Takeaway 5.
06Next check
The supplier’s support team, in writing
Their API accepts duplicate orders and has no idempotency header. Our own reference field, our_ref, is stored on every order they hold and is searchable.
The supplier cannot deduplicate. What does the MCP server do before it repeats the call?
Hint
You have a field of your own on every order they hold.
A Correct. If the downstream system cannot deduplicate, look before acting again. The check belongs on the server, where the reference is known and the answer is a fact rather than a guess.
B Wrong. A key the receiver ignores is a comment. Nothing about the second call changes.
C Wrong, and it accepts the ₹8,40,000 commitment overnight. A repair after the fact is not a control, and the supplier may have picked and shipped by then.
D Wrong. An alert reports that it happened. The question is what the server does instead of placing the second order.
Where the receiver cannot deduplicate, read state first and act second. Takeaway 6.
07Wrong lever
The proposal after the incident review
"Turn retries off for every write tool. On a timeout, stop the task and page on-call."
What does turning retries off move rather than remove?
Hint
Ask what the paged engineer has to do first.
A Correct. Turning retries off moves the problem instead of removing it. Somebody still has to find out whether PO-44812 exists, by hand, at whatever hour the task stopped.
B Wrong about which problem. It does stop this duplicate, and it leaves every timed-out order in an unknown state with nothing checking automatically.
C Wrong. Nothing here turns on token cost. What moves is an unresolved question about state.
D Wrong. Retry behaviour is orchestration, not a model choice, and asking the buyer does not establish whether the order exists.
A person paged with an unknown outcome is the same problem, with fewer facts. Takeaway 7.
08Diagnose
Thursday. The key is in place, and it happened again.
09:02:10 call create_purchase_order idem_key=po-9f2c4e11 no response after 30000 ms
09:02:44 call create_purchase_order idem_key=po-41ab7c93 201 Created order=PO-45190
The supplier holds PO-45189 and PO-45190 for the same 1,200 brackets.
The team added an idempotency key. The supplier still received two orders. What is the most likely cause?
Hint
Read the two key values in the trace, then ask who produced them.
A Possible in general, and not what this trace shows. Check your own chain first: two different keys went out, so the supplier was asked for two different actions.
B Wrong. One retry produced this. A second call carrying a new key is a second order whatever the limit is.
C Correct. A key the model mints on each call gives no protection. The key has to be created once for the intended order and held where re-planning cannot reach it.
D Wrong. The hint neither creates nor breaks deduplication. It is metadata a client may read when it decides what to retry.
If the model can mint the key, the key belongs to the call and protects nothing. Takeaway 5.
09Predict
Friday. The key is created once and stored outside the model.
The key is generated when the task decides to order, written to the task record, and reused on every attempt. The next call times out at 30 seconds.
What has changed about what you know after this timeout?
Hint
Separate two questions: what happened, and what is safe to do next.
A Wrong. Nothing about a timeout became informative. The retry is safe to make; it is not a report on what the first attempt did.
B Correct. A timeout is still missing information. The key does not answer the question, it removes the cost of asking it again.
C Wrong. No response arrived, so nothing was learned about storage. The key was sent, not acknowledged.
D Wrong. The server was waiting on the same call. A hop that is also waiting has nothing extra to tell you.
A key makes the retry safe. It does not make the timeout informative. Takeaway 1.
10Next check
Monday. Eleven write tools, one week.
The agent can call eleven tools that change something outside it. You have a week before the next order run.
Which question do you ask of each of those tools first?
Hint
The answer names two things a tool has to let you do.
A Wrong first question. You would end the week with a tuned policy and the same unknown outcomes. That is the retry lever again.
B Correct. Recovery design is the separate decision, and those two answers are what it consists of. Each tool then ends the week with a named step rather than a tuned number.
C Wrong. That reads a declaration. It tells you what a client may believe, not what the receiving system does with a second call.
D Wrong, and it arrives after the money moves. An alert reports. The question is what the system does instead.
For every write tool: can I tell whether it landed, and can I repeat it safely. Takeaway 2.
KitNext checkThe Agent Failure Triage Kit · the triage tree
The kit asks four triage questions in a fixed order. Which one comes first?
Hint
One of the three failure classes is the expensive one to get wrong.
A Correct. Side effects come first, because misclassifying an uncertain action as a temporary failure is the expensive mistake.
B It is in the tree, second. Asked first, a timed-out payment is filed as missing evidence and a retry goes through behind it.
C Third in the tree. Asked first, an uncertain write looks retryable, which is how one customer is refunded three times.
D Fourth, and it is the last resort. Asked first, everything becomes an escalation and the tree does no work at all.
Ask about side effects before anything else.
KitNext checkThe Agent Failure Triage Kit · the three failure classes
The kit gives uncertain_action its own response. What is it?
Hint
Each of the three classes has a different first move. This one is about establishing what happened.
A That is temporary_failure, where the dependency could not be reached and nothing outside the system changed.
B Correct, and the kit says a timeout on a write is always this class. The first move is to establish what actually happened, not to try again.
C That is missing_evidence, where the read completed and the answer is that there is nothing there.
D The kit escalates when state cannot be read, and sends the key with it. Escalating first hands over a question the system could have answered.
Reconcile before you re-issue.
06 Where you stand
The score is the small half. The missed takeaways are the list.
0/10Not started
Answer the concept questions to see where you stand.
Read these again
Takeaway 1. A timeout is missing information, not a failure. A timeout tells you the response did not arrive. The action may already have taken effect. Everything you do next has to be correct whether it did or it did not. Read it again
Takeaway 2. Retry policy and recovery design are separate decisions. Tuning retry counts or backoff answers how often to try. It never answers the question that matters after a timeout: did my first attempt land? Read it again
Takeaway 3. MCP adds hops, and silence at the client says nothing about them. Between the model and the system that acts there is a client, an MCP server and a downstream API. A client that stops waiting has learned nothing about what the other two did. Read it again
Takeaway 4. idempotentHint declares; it does not enforce. It is metadata on a tool definition that a client may use when deciding what to retry. It changes what the client believes, not what the downstream system does, and it defaults to false. Read it again
Takeaway 5. An agent-safe idempotency key belongs to the intended action. It is created before the first attempt, stored outside the model, reused on every retry and forwarded downstream. A key the model mints fresh on each call gives no protection. Read it again
Takeaway 6. If the downstream system cannot deduplicate, look before acting again. The server checks whether the action already exists before it repeats it. That check runs where your own reference is known, not in the model. Read it again
Takeaway 7. Turning retries off moves the problem instead of removing it. The same unanswered question goes to a person, with less information than the server had and at a worse hour. Read it again
A timeout tells you the response did not arrive. The action may already have taken effect. Everything you do next has to be correct whether it did or it did not.
02
Retry policy and recovery design are separate decisions.
Tuning retry counts or backoff answers how often to try. It never answers the question that matters after a timeout: did my first attempt land?
03
MCP adds hops, and silence at the client says nothing about them.
Between the model and the system that acts there is a client, an MCP server and a downstream API. A client that stops waiting has learned nothing about what the other two did.
04
idempotentHint declares; it does not enforce.
It is metadata on a tool definition that a client may use when deciding what to retry. It changes what the client believes, not what the downstream system does, and it defaults to false.
05
An agent-safe idempotency key belongs to the intended action.
It is created before the first attempt, stored outside the model, reused on every retry and forwarded downstream. A key the model mints fresh on each call gives no protection.
06
If the downstream system cannot deduplicate, look before acting again.
The server checks whether the action already exists before it repeats it. That check runs where your own reference is known, not in the model.
07
Turning retries off moves the problem instead of removing it.
The same unanswered question goes to a person, with less information than the server had and at a worse hour.
→ Where this comes from
The fix was never a retry setting
Every wrong answer in this quiz is something a competent team has proposed in a real incident review. Raise the retry count. Add backoff. Mark the tool idempotent. Turn retries off and page somebody. Each one is a reasonable instinct, and none of them answers whether the first attempt landed.
The two questions that do are the ones in question ten, and they are asked of a tool rather than of a system. That is why the audit is a list of tools and not a design review.
Source The annotation claim
Takeaway four and question four describe the Model Context Protocol specification,
version 2026-07-28, checked 26 September 2026. Hint defaults and client
behaviour can change between versions, so re-read it before you rely on this page.
idempotentHint: "If true, calling the tool repeatedly with the same arguments will have no additional effect on its environment." Default: false. schema/2026-07-28/schema.ts, ToolAnnotations