The Off Switch: Who Can Stop the Agent at 2 a.m.?
Every vendor says "AI agent governance." Almost none of them can tell you who could stop the agent at 2 a.m., or how you'd know by nine.
On the afternoon of 11 August 2020, three people at Citibank looked at the same screen and approved the same payment. A loan-operations contractor entered it. A colleague checked it. A manager signed it off. The bank's manual required all three — a "six-eyes" control, built so that no single person could send money by mistake.
The payment was an interest instalment of about $7.8 million to the lenders on a Revlon loan. The screen belonged to Flexcube, the loan-servicing software Citi used. To send interest only, the operator had to tick two boxes — labelled FRONT and FUND — as well as the one marked PRINCIPAL. That was how the principal got routed to an internal holding account instead of out of the door. None of the three noticed the two boxes weren't ticked. The screen didn't show what would happen if they weren't. By six o'clock that evening, roughly $893 million had gone to the lenders: the interest, plus the entire outstanding principal of a loan not due until 2023.
They found out the next morning. About $500 million wasn't returned, and it took Citi a lawsuit, a loss at trial — the judge called it "a banking error of perhaps unprecedented nature and magnitude" — an appeal and two years to get it back.
Citi had governance. It had a policy, a three-person approval, a manual, and a control with a name. What it didn't have is the subject of this piece. Governance is not a document. It's the four things you can point to at 2 a.m.
Read the middle one carefully. Deloitte's definition of a "mature governance model" rests on three capabilities: clear boundaries for what an agent may decide on its own; monitoring of what it's doing in real time; and an audit trail of what it did. Roughly four in five organisations surveyed report no mature model covering them. This piece keeps the first, folds the second into the third — a trail nobody reads before nine is a post-mortem, so the trail has to be the thing that wakes someone — and adds two the surveys don't ask about: approval before the irreversible, and the stop.
First, understand what "governance" is standing in for(and why the word does no work at 2 a.m.)
Every vendor deck has a governance slide — a policy, a committee, a review cadence, a responsible-AI statement. All worth having. None of them is running at two in the morning when a workflow with credentials to your ERP does the wrong thing a hundred times before anyone is awake.
What's running at two in the morning is the software. So the only governance that counts in that hour is built into how the software behaves — what it can reach, what it must wait for, whether it can be stopped, and what it writes down. Everything else is a description of intent.
The gap between the two is the gap Citi fell through. The policy said three people must approve. The software let three people approve a screen that didn't show the consequence, and then sent an irreversible payment with no way back. The policy was satisfied. The money was gone.
And that was a system with three humans in it. In Deloitte's agentic-transformation survey, 61% of leaders expect that within four years most of their agents will be generally autonomous, with humans acting as oversight rather than as a step. When the human isn't in the loop, the controls built into the software aren't the strongest layer of governance. They're the only one running.
So let's name them: the Four Controls
Think of the automation as a contractor you've given the keys to. There are four questions you'd want answered before you left the building — and they're the same four for software.
If it went wrong at 2 a.m., who could stop it — and how would you know by nine?If either half of that takes longer than a sentence to answer, you don't have governance. You have a document.
The $893 million, one control at a time
Citi isn't an AI story. That's the point. It's a story about where controls live, and it fails all four questions.
The workflow that paid interest was able to pay principal. The limit lived in a manual and two checkboxes, not in the system. A boundary that exists only in the instructions is one the software doesn't know about.
Three people approved. But the screen showed the payment without its consequence — that the untouched boxes meant the full principal would go. An approval that doesn't show what's being approved is a signature, not a check. This is the one most vendors think they've solved by "adding a human in the loop."
Once the wires went, there was no stopping them. A wire is final by design; the only stop available was never to send it. For anything that acts irreversibly, the stop has to come beforethe action, or it doesn't exist.
Citi had a complete record — that's how a colleague reviewing the previous day's transactions found it, the next morning. The trail was accurate and useless, because nobody read it before nine. A trail that isn't read before the damage is done is a post-mortem — which is why control 04 has to notice, not just record.
Two months after the wire, US regulators fined Citi $400 million over its risk management and internal controls broadly — not this payment, though the timing was lost on no one. Ask which of the four controls the fine would have bought. A framework doesn't buy any of them. Software does.
The stop, specifically(the AI agent kill switch, and 45 minutes at Knight Capital)
Control three is the one the surveys don't ask about and the one that decides how bad a bad morning gets. The public case is Knight Capital, 1 August 2012: a deployment error left old code active on one server, and when the market opened it began sending orders it shouldn't have. It took the firm about 45 minutes to stop it. The loss was about $440 million pre-tax by Knight's own account — the SEC put it above $460 million — and the firm was merged out of existence within a year.
The SEC's order found that Knight had received 97 automated warning emails before the open, not designed as alerts and not acted on — and that the firm lacked written procedures for responding to an incident. Read that as the four controls: a trail nobody watched, a boundary nobody set, and a stop nobody had rehearsed. Forty-five minutes is what "we'd just turn it off" costs when turning it off isn't a designed, tested action.
Where the four controls usually live(policy vs contract vs promise vs runtime)
Ask any vendor about the four and you'll get an answer. The useful question is whereit lives. Four places, in rising order of usefulness at 2 a.m. The last column isn't the whole of governance — policy, accountability and audit all matter — but it's the part that's awake.
Read across any row: the first three columns are the same control described three ways, and none of them is running at 2 a.m. Read down the last column: it's the only one where the answer is a behaviour rather than a sentence. Most of what is sold under the word "governance" lives in the first three columns. Citi's six-eyes control lived in the first. That's the twist, and it's the whole buying question: not whether a vendor has the four, but which column each one is in.
Six objections
"Human-in-the-loop solves this."
Citi had three humans in the loop. The approval step is only a control if the person can see the consequence of what they're approving — the amount, the action, what happens if they click. "A human approves" is the beginning of the design question, not the end of it.
"We'd just switch it off."
Switching off is easy when nothing is half-done. A workflow mid-way through a batch of payments, or a filing on a government portal, is not nothing. The stop has to define what happens to the step in flight — complete it, undo it where that's possible, or halt it in a known state — and someone has to have tried it before the night they need it. Knight's forty-five minutes is what an unrehearsed stop looks like.
"Isn't this the same argument as The Four Exits?"
It's the other half of it. Four Exits was about where data goes once the agent can read it — reach. This is about what the agent may do, and whether you can take it back — permission and reversibility. Reach and rights were one pair. This is the second.
"If a deterministic engine is what makes this enforceable, isn't it just workflow automation with an AI label?"
The answer is where the autonomy sits. The agent's autonomy is in designing the workflow and in reading and deciding within it — which supplier this is, what the total should be, whether this document is confident enough to pass or should stop. The engine's determinism is in executing what was decided: the same step, the same way, every time, logged. Judgement where judgement is needed; certainty where money moves. A system that's autonomous at the point of execution isn't more capable. It's just harder to stop.
"Isn't this just the EU AI Act's human-oversight requirement?"
Partly, and it's worth knowing. Article 14 of the EU AI Act requires that a high-risk system can be overseen by people — including the ability to "intervene in the operation of the high-risk AI system or interrupt the system through a 'stop' button or a similar procedure that allows the system to come to a halt in a safe state." NIST's AI Risk Management Framework puts the same idea under its Manage function; ISO/IEC 42001 sits alongside. So the stop is not a novelty. Regulation already mandates one. What regulation doesn't specify is what this piece is about: who can press it at 2 a.m., how fast it takes effect, what happens to the step in flight, and whether anyone has rehearsed it. A stop that exists to satisfy Article 14 is in the first column of the matrix. The Act requires that it exist. Only the runtime decides whether it works.
"A vendor would say this."
Yes. So take the four questions at the end to our demo before anyone else's, and ask for each one which column it's in. If any answer starts with "our policy is," you've placed it.
What 2 a.m. looks like at your scale
Strip the zeros off Citi and it's an ordinary night in an ordinary back office. What follows is an illustration — the timings are a scenario, not a measurement. A supplier-payment workflow is working through a batch of four hundred. Somewhere in the configuration, a tolerance has been set a decimal place too wide, and the eleventh invoice is about to be paid at ten times its value. Two versions of the next seven hours follow.
The difference is not the AI. It's which column the approval control lives in.
Where the four controls live in AutomatR
Framed as what you'd see, not what we'd claim. All four sit in the runtime column, for the reason that runs through this series: the AI designs the workflow, and a deterministic engine runs it. An engine that can't behave differently on Tuesday than it did on Monday is what makes a control enforceable rather than assured.
- Boundaries. Credentials are provisioned per workflow with least privilege and held by the engine; the designer never holds production credentials.
- Approval.Irreversible actions are declared at design time. At run time the engine suspends at that step, persists state, and renders the pending action — target, record, values — as a card; the resume carries the approver's identity and decision into the log.
- Stop. An operator with the role can pause or terminate a run; the engine completes the current atomic step, reverts it where the target system supports reversal, or halts it in a defined state, and records who stopped it and why.
- Trail.Every step, its inputs, decision and reason are written in order; healed steps record before and after; a suspended run or a failed check raises a notification to the queue owner's role; any run can be replayed step by step — the "Replay Test" finance teams ask for at audit.
The AI designs the workflow; a deterministic engine runs it.
Control 04 is the one we'd rather show than describe: pick any run from a pilot and we'll replay it in front of you, step by step, with the reasons. Send your first process through automatr.tech/contact and ask for the replay.
The Four Controls Card
Take these to any vendor demo, including ours — and for each answer, ask which column it's in.
- 01BoundariesWhat can this workflow reach that it shouldn’t need to — and what stops it? Show me the credentials, not the policy.
- 02ApprovalShow me the approval screen for a payment. Does it show the amount and the target, or a task ID?
- 03The stopWho can halt a run at 2 a.m., what happens to the step in flight, and when did you last rehearse it?
- 04The trailReplay yesterday’s run for me now, step by step, with the reason at each step — and show me what would have woken someone at 2:05.
A vendor who answers all four with behaviour you can watch has governance. One who answers with a slide has a framework. Both are worth having. Only one is running at 2 a.m.
If your automation went wrong at 2 a.m. tonight, who could stop it — and how would you know by nine?
Further reading: going deeper on control
- →The Four Exits — reach: where the data goes once the agent can read itautomatr.tech/blog/four-exits-agent-data-leaks
- →The 99% Lie — why the false-pass rate is the number that keeps approval honestautomatr.tech/blog/the-99-percent-lie
- →The Four Surfaces — the runtime the controls live inautomatr.tech/blog/four-surfaces
- →EU AI Act, Article 14 — Human oversight — the regulatory floor for the stopartificialintelligenceact.eu
- →NIST AI Risk Management Framework 1.0 — the Manage functionnist.gov
- →SEC, “SEC Charges Knight Capital With Violations of Market Access Rule” (16 October 2013)sec.gov
- →Al Jazeera, “$900m gaffe: US judge says Citigroup can’t recoup Revlon payout” (16 February 2021)aljazeera.com