Twelve objections put to the category by people who run finance and operations functions, answered at the length the question deserves, with the ones nobody has settled marked as open.
12 questions·2 unresolved·Reviewed September 2026
The questions that arrive in the first hour
A category that cannot survive its own hard questions is marketing. These twelve arrive in the first hour of any serious conversation about letting software act, and they arrive from the people whose signature is on the control environment it would act inside: the controller, the head of internal audit, the person who will be asked what happened. A definition, a maturity model and a test, published without them, is the easy half.
No answer here appeals to what the technology will be able to do later, and none asks the reader to grant a premise they came in doubting. Where the honest answer is a control, the answer names the control.
2 of the 12 are marked unresolved. Those questions have no settled answer anywhere, and a confident paragraph written over that gap would be the failure the mark exists to prevent. The mark travels with the answer, so a reader who arrives at one section from a link elsewhere sees it without having read anything above.
The order runs from the question that decides whether any of this is worth attempting to the question of where it stops. Several of the answers name a control instead of an argument, and those controls are set out at working depth on the governance page.
Objection 01
Why let a system that can be wrong act at all?
Because the alternative is not a system that cannot be wrong. It is a person who can be wrong, working from a plan produced overnight, holding less of the position than the system can read and leaving no record of what they looked at before deciding. The argument for letting software act is never that it is right more often in general. It is that for a named class of decision the error rate has been measured, the cost of an error is capped by a ceiling somebody granted, and the action carries a reversal procedure. Measured means sampled against the same decision taken by a person on the same evidence, and the sampling continues after go live and does not end with the pilot. A decision that fails any of those three tests should not be autonomous, which is a much narrower permission than most descriptions of automation ask for.
Objection 02
Unresolved
Who is accountable when an autonomous action causes a loss?
This is unsettled, and a page that answers it cleanly is describing a preference, not a practice. The settled half is the grant. A named person sets the action types, the ceilings and the conditions, signs them, and answers for the total exposure they authorised, in the same way a delegation of authority has always worked. The unsettled half is a single action inside a valid grant that produced a loss nobody would have chosen. Software vendor contracts push that to the buyer through limitation of liability, professional standards do not yet address a decision with no human author, and there is no body of tested cases to read. Until that changes, the safe reading is that the ceiling on a grant is the amount the company has decided it can lose with no recourse to anybody else, because in practice that is what it is. A programme that has not had that conversation with the person signing the grant has not finished designing the grant.
Why this one carries the mark
The answer above states where the question stands and does not settle it. Nothing here should be read as a resolution, and a system described as having solved this one is describing something narrower than the question.
Objection 03
What becomes the system of record?
The ledger remains the system of record for balances, and the subledgers remain the systems of record for their own transactions. Nothing here asks for that to move. Those records are authoritative because they are controlled, reconciled and signed, and none of those properties come from the software that happens to hold them. What is new is a second record that did not exist before: the decision record, holding what was observed, what was decided, under which grant, and what was posted as a result. The two are joined by reference, never merged, so the ledger keeps its own integrity and the decision record can be retained on a different schedule. A system that asks to become the system of record is asking for a migration and should be priced as one. A system that asks to write to the ledger and record why can be audited against the ledger it wrote to.
Objection 04
How does an audit trail work when the actor is software?
It works as it always has, with the identity of the actor replaced by a pair instead of a single name: the system instance that acted, and the grant it acted under, with the grantor named on that grant. Each action has to carry the evidence as it stood at decision time, the version of the policy then in force, the alternatives weighed and rejected, the grant identifier, and the artefact the action produced. The part that is genuinely harder than a human trail is freezing the inputs. A person's evidence can be reassembled from the documents they filed, while a system's evidence is a query result that will not return the same values tomorrow. Evidence therefore has to be captured at decision time and retained, never recomputed for the auditor afterwards, because the position it described has moved.
Objection 05
How does segregation of duties survive when one system holds several roles?
It does not survive unchanged, and saying that it does is the usual failure. The two person rule was never about two people. It was about making collusion necessary for a loss, so that no single actor could both create an obligation and settle it. When one system can raise an order, receive against it and release the payment, that property is gone at the level of the actor and has to be rebuilt at the level of the grant. The authority to create an obligation and the authority to settle one become separate grants signed by different people, and no single grant carries both. Two controls carry the rest: the system cannot widen its own grant, and its actions are checked after the fact against a record it cannot write to. Where that second check does not exist, segregation has been lost, not redesigned, and the honest description is a control accepted as a risk, not a control replaced.
Objection 06
Unresolved
Should the system's confidence decide how much authority it has?
In principle yes, in practice nobody has a defensible method, so this one is unsettled. The appeal is obvious: act alone when sure, escalate when not. The difficulty is that a confidence score is only meaningful if it is calibrated, meaning that decisions taken at ninety percent confidence turn out correct about ninety percent of the time, and calibration has to be measured per action type against outcomes that arrive weeks later. Very few deployments carry that measurement, and a score that has never been checked against outcomes is a number with a decimal point, not a probability. What can be done now is narrower and still worth doing. Hold the ceiling where a person set it, let low confidence route work to a person earlier than the ceiling would, and never let high confidence raise the ceiling. Confidence can spend caution. It should not be allowed to spend money.
Why this one carries the mark
The answer above states where the question stands and does not settle it. Nothing here should be read as a resolution, and a system described as having solved this one is describing something narrower than the question.
Objection 07
What happens when two agents disagree?
A disagreement settled at run time by whichever component writes last is a design defect, not a disagreement. The resolution is a stated precedence: for each class of decision one objective is named as dominant, and where two grants would authorise conflicting actions, neither acts and the conflict escalates with both positions recorded. The failure to design against is not deadlock, which is at least visible. It is the quiet case where a demand plan and a treasury plan each act correctly inside their own grant and the combination is a position nobody chose. That is why the conflict test has to run against the combined effect on the record, not against the two instructions, and why the escalation has to name what each side was optimising, not merely that a clash occurred.
Objection 08
How is an autonomous action reversed?
Through a defined procedure in the system that took the action, producing its own record, and never through a correction made elsewhere by whoever noticed. The mechanics decide what can be granted. An action still inside the system can sometimes be withdrawn before it lands. An action that has left, such as a released payment or an order sent to a supplier, is not deleted but offset by a second action that compensates it, and both remain visible, because the first one is now part of what somebody else did. Reversal cost follows a gradient that has nothing to do with the size of the number. A journal is reversible for the cost of an entry. A payment is reversible only through a recall that may fail. A message that reached a customer is not reversible at all. Grants should be written against that gradient, not against value alone.
Objection 09
What stops one error from being repeated a thousand times before anybody notices?
Nothing inside the decision itself, which is why the control sits outside it. Three mechanisms carry this. A rate limit per action type, so a defect produces a number of actions somebody can still review, not a day of them. A standing reconciliation that compares what the system intended against what the record shows, and suspends the action type on divergence instead of reporting it. And a ceiling on aggregate value per period that is separate from the ceiling per action. A system with per action limits and no aggregate limit is the one that produces the incident, because every individual action was inside policy and the total was never tested against anything.
Objection 10
Can this coexist with the systems already in place?
Yes, and in a business of any size that is the only way it arrives. The systems of record stay where they are. The new part reads from them, decides, and writes back through the same interfaces a person's transaction uses, so every action meets the same validation and lands in the same population an auditor samples. The real constraint is not integration, which is ordinary work. It is that a decision is only as good as the account it reads, so the first scope should be a domain whose data is already reconciled and whose exceptions are already understood. Choosing the domain with the worst data because it shows the largest theoretical gain is the most reliable way to stall the programme, and it stalls it at the point where nothing has shipped and the cost is already visible.
Objection 11
Does it replace those systems immediately?
No, and for most businesses not for years. The ledger, the subledgers and the transactional systems keep doing what they do, and there is no version of this that starts by switching them off. What moves first is where planning and decision making sit, and that can move without the record moving at all. Replacement becomes a question only once the decision layer has held authority long enough that the older planning modules are genuinely unused, at which point retiring them is an ordinary decommissioning exercise, not the premise of the programme. A plan that requires a replatform before anything autonomous runs has taken both risks at once, and it will be judged on the migration, not on whether the loop closed.
Objection 12
What should never be fully autonomous?
Three categories, stated as categories so they survive cases nobody has met yet. Decisions that bind the company legally, because signing is an act of a legal person and delegating it is a decision about the company, not a decision inside it. Decisions that affect a person's employment, pay or record, because the subject of such a decision is owed an answer from somebody who can be asked why, and a system cannot hold that obligation. And decisions whose reversal costs more than the decision saves, which covers disclosure, anything that reaches a customer or a regulator, and any action with a physical consequence. Everything outside those three is a question of grant size and evidence, not one of principle. Drawing the line by transaction value instead is the common error, because it puts small irreversible actions on the permitted side of it.
Send a better answer
Every answer above is the best one available at the date in the header, which is a different claim from the best one there is. If you have run one of these decisions in production and the answer here is wrong, incomplete, or true only of a narrower case than it implies, the correction is worth more than the paragraph it replaces.
Corrections go through the route described on the about page, which also records how a change to a published answer is marked when it is made. Corrections to the two unresolved questions are the ones most wanted. A worked account of who carried a loss on a bounded autonomous action, or of a confidence measure calibrated against outcomes for a named action type, would move either one further than an argument.
An answer that names a control is worth what the control is worth.