Tools · browser checklist
Agent design check
19 questions about a system that acts. Answer the ones you can, leave the rest, and get back what is still open and what to do about it — in the order the decisions are worth making.
Worked out in this page. No sign-in, no score, and nothing you type is sent anywhere.
Before you start
What this does, and what it will not do
- 01 Every next step below is written in the rubric, next to the question it belongs to. Nothing is generated, ranked or summarised by a model, and you can read the whole rubric before answering anything.
- 02 Steps are ordered by two things only. First, whether the question is about an action that leaves the system — one you cannot take back. Then the order the questions appear in, which is the order these decisions are worth making.
- 03 A reported gap and a not-yet-defined answer are not ranked against each other. They are different kinds of work: a gap already has a decision behind it and needs a design change, while an undefined answer means nobody has decided and the first step is to find out.
- 04 Answering yes closes nothing. It asks for the artefact that would show it.
- 05 There is no score, no percentage and no grade, and no set of answers produces a readiness verdict. Nineteen yeses produce nineteen requests for evidence.
Your answers are held in this page and nowhere else. They are not saved to your browser, not sent to a server, and not recoverable by anybody — including us — once you close the tab. That is also why there is nothing to sign in to.
The example throughout
A returns assistant that can pay
Imagine a returns assistant. A customer writes in about a return. The assistant finds the relevant policy, looks up the order, and recommends a refund. The conversation reads well. You can follow the reasoning, and you can see how it would save somebody a lot of repetitive work.
Then someone asks a perfectly reasonable question: can it issue the refund as well?
On a diagram that is one more connection. In the world, the system can now move money — and that changes which questions are worth asking about it.
Every question below continues this one example, so you can answer for your own system without having to describe it to anybody, and without a new scenario to learn each time.
An illustrative example. It is not an account of a real customer system, and no incident is being described.
The three answers
“No” and “not yet defined” are different facts
- Addressed
- You are saying the design has an answer here. That is a claim, and the result asks what would show it.
- Reported gap
- You know the answer and the answer is no. Somebody looked; there is a hole you can already name, so the next step is a design change.
- Not yet defined
- Nobody has decided. What the system does here today is unknown to the people who own it, so the next step is finding out, not building.
- Not answered here
- You did not answer this one. Nothing has been assumed about it, and it is reported as still open rather than as a gap.
The 7 areas
01 Area
Purpose and boundary
What the system is for, where it stops, and whether it needed to be an agent at all.
The open part and the closed part
A purpose somebody can check
Non-goals, written down
02 Area
Evidence
What must be established before it acts, and whether that can be checked afterwards.
An evidence list per action
Evidence as a reference, not a recollection
The tool re-checks what it was handed
03 Area
Tool permissions
What each tool may do, and where that limit is actually enforced.
One job, narrowest scope
Limits enforced at the tool
Its own identity, and a record of every call
04 Area
External actions
What happens when the system does something the world can see and you cannot take back.
Recommending and acting are separate
A repeat cannot pay twice
An unknown outcome can be established
05 Area
Evaluation
What you can see after a run, and what the system is checked against.
A record of what it used and did
Cases, and failures that become cases
06 Area
Uncertainty and recovery
What it does when it cannot establish what it needs, and how a person takes over.
A third outcome besides act and fail
Kinds of uncertainty routed differently
A person can take over, and is handed something
07 Area
Ownership
Who is answerable for the behaviour, and what happens when it changes.
A named owner, and one per tool
Prompt, model and scope changes get reviewed
If a question here landed
Nothing you answered was sent anywhere, and no record of you exists. The page view itself is counted, as it is on every page here, and that count cannot be joined to a person or to anything above. If you want to take one of these questions further, these are the routes that exist.
- The guides — the reasoning behind these questions, in longer form. Free, and no sign-up.
- The open cohort — building a working agentic system and practising the argument for its design, with feedback. Applying starts a conversation about fit.
- Advisory — if the decision in front of your team is the immediate problem rather than the practice of making them.