ARPFramework

Eight questions, three of which decide it

A definition is worth as much as the test that applies it. Each of these eight is answerable about a system in production, and each has an answer that sounds like a pass and is not.

Eight questions, 3 load bearingAnswered per resource classReviewed September 2026

The eight questions

Answer them about one resource class at a time, about a system in production rather than in a demonstration, and in the order below, which runs from when the work starts to what happens when the system reaches the edge of what it was given.

  1. 01Does the system start the work itself, or does a person start it every time?
  2. 02Can you print the list of actions the system may take on its own, with the limits on each?Load bearing
  3. 03Does the action land in the system of record, or in a queue somebody rekeys?Load bearing
  4. 04Does the system notice its own failures before a person does?
  5. 05For one action taken last quarter, can you recover the inputs, the policy, the alternatives and the evidence?Load bearing
  6. 06When the system is wrong, is there a recorded reversal, or a person cleaning up by hand?
  7. 07When conditions move, does the plan move with them without being asked again?
  8. 08Does the system stop when it is outside its bounds or below its confidence, and hand over with its evidence?

Mark each one passed or failed and keep the marks per class, because a business whose cash application answers all eight and whose capacity planning answers two has learned something a single verdict would have hidden. The five unmarked questions grade how completely a system does this. The three marked ones grade whether it does it at all.

Question by question

Question 1

Does the system start the work itself, or does a person start it every time?

Autonomy is first a question of who begins. A system that does excellent work once somebody opens it, selects a scope and presses run is a tool being operated, and the judgement about when the work was needed stayed with the operator. That judgement is most of the value, because a decision taken three days late is usually the wrong decision however well it was computed.

A passing answer

A change in the world starts the work: a receipt, a bank line, a forecast revision, a missed delivery date. The system decides that something now needs deciding, and no person asked it to look.

An answer that sounds like a pass

It runs on a schedule, or it responds to a prompt, and the vendor calls the schedule proactive because the person did not have to remember.

Question 2

Can you print the list of actions the system may take on its own, with the limits on each?

Load bearing

Bounded authority is what separates autonomy from an unowned risk. If the grants cannot be enumerated then nobody can say what the system is permitted to do, which means nobody can say what it did was permitted. This is the question an auditor, an insurer and a regulator each arrive at first, and a system that cannot answer it will not be allowed near anything that matters.

A passing answer

A named person can produce the current grants as a document: which action types, up to what value, under which conditions, granted by whom, on what date, revocable and dated on revocation.

An answer that sounds like a pass

Authority is described as configurable, or it lives in a model prompt, a role name, or a set of permissions written for human users and inherited by the software.

Question 3

Does the action land in the system of record, or in a queue somebody rekeys?

Load bearing

An action that a person has to transcribe has not been taken by the system. The work, the timing and the error rate all still belong to the human, and the autonomy is a demonstration rather than a fact.

A passing answer

The posting, order or payment exists in the ledger with the system named as its author, and the reversal path is recorded against it.

An answer that sounds like a pass

It drafts the entry for review, or it writes to a staging table that an integration picks up on a schedule.

Question 4

Does the system notice its own failures before a person does?

A system that acts without a person in the path has also removed the person who used to spot the mistake. If nothing replaces that, the error rate has not changed but the time to discovery has grown, which is a worse position than the one before. Detection is the cost of the authority, not an enhancement on top of it.

A passing answer

The system compares what it intended against what the world shows, on a defined interval, and raises its own discrepancies: an order it placed that no receipt matched, a payment it scheduled that no bank line confirmed.

An answer that sounds like a pass

Failures surface through a monitoring dashboard somebody reads, through a downstream reconciliation run weeks later, or through a customer.

Question 5

For one action taken last quarter, can you recover the inputs, the policy, the alternatives and the evidence?

Load bearing

Every claim about a system's judgement rests on being able to inspect one instance of it after the fact. Reconstruction is what makes an autonomous decision reviewable, contestable and learnable from, and it has to survive the model changing, the policy changing and the person who granted the authority leaving. A decision that cannot be reconstructed can only be trusted or distrusted wholesale.

A passing answer

A single reference retrieves the state the system read, the rule or objective it applied, the options it weighed and rejected, the authority it acted under, and the record the action produced, all as of the moment it acted.

An answer that sounds like a pass

There is a log of what happened, or a chat transcript, and the reasoning is reproduced by asking the system today what it would do now.

Question 6

When the system is wrong, is there a recorded reversal, or a person cleaning up by hand?

Autonomy at any useful scale assumes error, because the alternative assumption is that the system is never wrong. The design question is not how to avoid the wrong action but what the wrong action costs to undo. If undoing is manual, the recovery work scales with the volume of autonomous work, and the benefit is spent on the cleanup.

A passing answer

Each action type has a defined reversal that the system can execute: a compensating entry, a cancellation, a credit, a release of a commitment, recorded against the original with both visible in the ledger.

An answer that sounds like a pass

Reversal is a support procedure, a data fix, or a documented runbook for a person to follow when something goes wrong.

Question 7

When conditions move, does the plan move with them without being asked again?

A plan is a statement about a position, and positions change. A system that plans once and holds the answer until somebody requests a rerun has produced a report, and the gap between the plan and the world is filled by people with spreadsheets. Continuity is what makes the difference between a system that knows the current position and one that knew a position.

A passing answer

A late shipment, a cancelled order or a changed forecast causes the affected plans to be recomputed and the dependent actions to be revised or withdrawn, within the window the business actually moves in.

An answer that sounds like a pass

Replanning happens on the next nightly or weekly run, or when a planner opens the screen and refreshes it.

Question 8

Does the system stop when it is outside its bounds or below its confidence, and hand over with its evidence?

The value of bounded authority depends entirely on what happens at the boundary. A system that proceeds anyway has no bounds; a system that stops and says nothing useful has moved the work to a person without moving the context. What makes escalation work is that the handover carries what the system knew, so the human decides rather than starts again.

A passing answer

Out-of-bounds and low-confidence cases stop before the action, reach a named role with the state, the options and the reason for stopping attached, and are tracked until somebody decides.

An answer that sounds like a pass

Everything is routed to a person for approval, which is review rather than escalation, or low confidence is handled by proceeding and flagging the result afterwards.

What failing a load-bearing question settles

Three of the eight are marked, and they are marked because each one removes a different necessary condition rather than a degree of quality. Without enumerated authority nobody can say what the system was permitted to do. Without closure nothing it decided has actually happened. Without reconstruction no single decision it took can be reviewed after the fact.

Stop scoring and say so

A system that fails one of the three is a system of record with assistance on top of it, whatever the assistance is called and however good it is. That is a description rather than an insult: assistance is worth paying for, most of the software in this market is assistance, and a well built assistant saves real money. It is simply not autonomous resource planning, and the remaining five questions cannot make it so.

The practical instruction is to stop at that point and record the failure plainly, because the alternative is the arithmetic that ruins every scored assessment: five passes out of eight, a majority, a rounded verdict, and a conclusion that survives the missing condition it was supposed to test. A score is a summary of a judgement already made. If one of the three has failed, the judgement is made.

  • Question 2Authority is described as configurable, or it lives in a model prompt, a role name, or a set of permissions written for human users and inherited by the software.
  • Question 3It drafts the entry for review, or it writes to a staging table that an integration picks up on a schedule.
  • Question 5There is a log of what happened, or a chat transcript, and the reasoning is reproduced by asking the system today what it would do now.

Each line above is the answer that most often stands in for a pass on that question.

The other five are a grade, not a verdict

A system can miss any of the unmarked five and still be doing this, less completely. A system that acts on bounded authority into its own ledger and can reconstruct any decision, but only replans nightly, is an autonomous system with a stale clock, and the remedy is an engineering one. Naming which of the five failed is more useful than a total, because each failure has a different repair and they are not interchangeable. The maturity model is where a passing system is then placed, class by class.

How to use this in a vendor meeting

The eight are written to be answerable from published documentation, not from a conversation, which is what makes them usable by somebody who is not an engineer and cannot be shown the source. Each one asks for an artefact that either exists or does not.

Ask for the artefact, not for the answer

Every question above has a document behind it. The authority question asks for the current grants as a printable list. The closure question asks for one posting in the ledger with the system named as its author. The reconstruction question asks for a single reference that retrieves one decision taken last quarter, with the state, the policy, the alternatives and the evidence attached. Ask for the artefact and the question answers itself, usually within the meeting.

Ask about a system already running somewhere, not about a roadmap, name the resource class before you start, and ask what happens when the answer is wrong. Those three moves convert most of the eight from a discussion into an observation, and an observation is what the test is for.

A vendor who cannot answer the authority question has answered it

Authority is the one question whose answer cannot be assembled during the meeting. A system either has the grants written down, dated, attributed to whoever issued them and revocable, or it does not, and if it does then somebody can produce them, because the grants exist precisely so that they can be produced. There is no third state in which a system has explicit bounds that nobody can show you.

So the substitutes are worth recognising by ear. The bounds are configurable. The limits are set per customer during implementation. It operates under the permissions of the user it runs as. The model is instructed not to exceed the policy. Each of those is an answer to a different question, and the reason they arrive in place of the grants is that the grants have not been written. Recording the substitute verbatim is enough. The question does not need to be pressed.

None of this is a reason to be hostile in the room. The test is about an architecture, the person answering did not design it, and most of these answers are given honestly by people who have never been asked for the grants before. The definition is the reason the questions are these eight, and it is the more useful thing to send afterwards.

A system that passes the eight still has a level, and it has one per resource class.

The maturity model