Resources / agentic system design 07 · The Agent Failure Triage Quiz

Tell a failed action from an unanswered one

One illustrative incident, worked one piece of evidence at a time: a purchase order that timed out, and a supplier that received it twice.

10 concept questions · 2 kit questions · 4 options each · answers stay in this browser

Built by Sunil Mathew, co-authored with Claude

01 What to bring

The list of tools your agent can call that change something outside it. At least one of them moves money or places an order.

02 What to do

Answer each question before you read the feedback. Every answer explains itself, including the wrong ones, and says what that choice would actually cause.

03 What you leave with

A recovery step you can name for each of your own write tools, and a list of the takeaways you missed, in your own order.

04 Read this first

Two things sit behind these questions. Episode 4 is where the procurement incident comes from. The Agent Failure Triage Kit is where the three failure classes and the order of the triage questions are written down.

  • Agentic system design, Episode 4 — The episode is on LinkedIn and is not linked here yet. The kit covers the same ground in writing.
  • The Agent Failure Triage Kit — the three failure classes, the four triage questions and the response playbook, in writing.

Start the quiz  Read the kit first

Illustrative case · not a real client, employer or event

One order, two purchase orders, no error anywhere

A procurement agent at a mid-size manufacturer raises purchase orders when stock runs low. A buyer approves the order; the agent places it.

It places every order through one tool, create_purchase_order. That tool sits on an MCP server the platform team runs. MCP (Model Context Protocol) is the standard way a client offers tools to a model. The MCP server calls the supplier’s ordering API.

On Tuesday at 09:14 the agent orders 1,200 steel brackets at ₹8,40,000. By 09:20 the supplier’s portal shows that same order twice, PO-44812 and PO-44813, ₹16,80,000 committed.

The agent’s own logs say no order was placed. Nothing raised an error. You are on call.

This incident is illustrative, and is not a real client, employer or event. You are the engineer on call, and each question hands you one new piece of evidence.

0/10 concept
Not started Answer question one. Nothing is stored and nothing is sent.
0/12
05 The triage

Commit to an answer, then read the evidence back.

Your answers stay in this browser. Nothing is stored, nothing is sent, and closing the tab clears them.

01 Diagnose
The agent’s client log
09:14:02.114  call create_purchase_order  sku=BRK-19 qty=1200 value=840000
09:14:32.115  call create_purchase_order  no response after 30000 ms

The call to create_purchase_order timed out after 30 seconds. No error came back. What do you know for certain?

Hint

Read the second log line again. Which side of the call does it describe?

02 Wrong lever
The incident channel, 09:41

"Raise max attempts from 3 to 5 and add exponential backoff: 2s, 4s, 8s, 16s. That is why we got two orders."

The team proposes five attempts with exponential backoff instead of three. What does that fail to fix?

Hint

Ask what the extra attempts know that the first one did not.

03 Diagnose
The MCP server’s log, from the platform team
09:14:02.140  tools/call create_purchase_order  forwarded  POST /v2/orders
09:14:34.210  tools/call create_purchase_order  attempt 2  forwarded
09:14:37.902  supplier responded 201 Created  order=PO-44812
09:14:41.377  supplier responded 201 Created  order=PO-44813

The agent stopped waiting at 30 seconds. The supplier created PO-44812 at 35 seconds. What did the agent’s silence tell you about the hops behind it?

Hint

Compare two clocks: when the client stopped waiting, and when the supplier answered.

04 Spot the claim
The tool definition, from the server’s tools/list response
{ "name": "create_purchase_order",
  "annotations": { "title": "Create purchase order", "idempotentHint": true } }

The supplier’s API documentation has no idempotency header and no duplicate suppression.

The tool is annotated idempotentHint: true. Which part of this setup is declared but not enforced?

Hint

Ask who reads the annotation, and who places the order.

05 Next check
Four proposals on the board, 11:20

The team agrees to add an idempotency key: a value that tells the receiver two requests are the same action, so it applies the action once. Four versions are proposed.

Which of these keys protects the agent on a retry?

Hint

For each one, ask when the key is created and who keeps it.

06 Next check
The supplier’s support team, in writing

Their API accepts duplicate orders and has no idempotency header. Our own reference field, our_ref, is stored on every order they hold and is searchable.

The supplier cannot deduplicate. What does the MCP server do before it repeats the call?

Hint

You have a field of your own on every order they hold.

07 Wrong lever
The proposal after the incident review

"Turn retries off for every write tool. On a timeout, stop the task and page on-call."

What does turning retries off move rather than remove?

Hint

Ask what the paged engineer has to do first.

08 Diagnose
Thursday. The key is in place, and it happened again.
09:02:10  call create_purchase_order  idem_key=po-9f2c4e11  no response after 30000 ms
09:02:44  call create_purchase_order  idem_key=po-41ab7c93  201 Created  order=PO-45190

The supplier holds PO-45189 and PO-45190 for the same 1,200 brackets.

The team added an idempotency key. The supplier still received two orders. What is the most likely cause?

Hint

Read the two key values in the trace, then ask who produced them.

09 Predict
Friday. The key is created once and stored outside the model.

The key is generated when the task decides to order, written to the task record, and reused on every attempt. The next call times out at 30 seconds.

What has changed about what you know after this timeout?

Hint

Separate two questions: what happened, and what is safe to do next.

10 Next check
Monday. Eleven write tools, one week.

The agent can call eleven tools that change something outside it. You have a week before the next order run.

Which question do you ask of each of those tools first?

Hint

The answer names two things a tool has to let you do.

Kit Next check The Agent Failure Triage Kit · the triage tree

The kit asks four triage questions in a fixed order. Which one comes first?

Hint

One of the three failure classes is the expensive one to get wrong.

Kit Next check The Agent Failure Triage Kit · the three failure classes

The kit gives uncertain_action its own response. What is it?

Hint

Each of the three classes has a different first move. This one is about establishing what happened.

06 Where you stand

The score is the small half. The missed takeaways are the list.

0/10 Not started

Answer the concept questions to see where you stand.

Open the kit 
07 The seven takeaways

What every question is actually asking.

  1. 01

    A timeout is missing information, not a failure.

    A timeout tells you the response did not arrive. The action may already have taken effect. Everything you do next has to be correct whether it did or it did not.

  2. 02

    Retry policy and recovery design are separate decisions.

    Tuning retry counts or backoff answers how often to try. It never answers the question that matters after a timeout: did my first attempt land?

  3. 03

    MCP adds hops, and silence at the client says nothing about them.

    Between the model and the system that acts there is a client, an MCP server and a downstream API. A client that stops waiting has learned nothing about what the other two did.

  4. 04

    idempotentHint declares; it does not enforce.

    It is metadata on a tool definition that a client may use when deciding what to retry. It changes what the client believes, not what the downstream system does, and it defaults to false.

  5. 05

    An agent-safe idempotency key belongs to the intended action.

    It is created before the first attempt, stored outside the model, reused on every retry and forwarded downstream. A key the model mints fresh on each call gives no protection.

  6. 06

    If the downstream system cannot deduplicate, look before acting again.

    The server checks whether the action already exists before it repeats it. That check runs where your own reference is known, not in the model.

  7. 07

    Turning retries off moves the problem instead of removing it.

    The same unanswered question goes to a person, with less information than the server had and at a worse hour.

→ Where this comes from

The fix was never a retry setting

Every wrong answer in this quiz is something a competent team has proposed in a real incident review. Raise the retry count. Add backoff. Mark the tool idempotent. Turn retries off and page somebody. Each one is a reasonable instinct, and none of them answers whether the first attempt landed.

The two questions that do are the ones in question ten, and they are asked of a tool rather than of a system. That is why the audit is a list of tools and not a design review.

Source The annotation claim

Takeaway four and question four describe the Model Context Protocol specification, version 2026-07-28, checked 26 September 2026. Hint defaults and client behaviour can change between versions, so re-read it before you rely on this page.

The Agent Failure Triage Kit  All resources

Published 26 September 2026 · free to use and to pass on