AI Strategy

Should reversibility set the limit on agent autonomy?

← Back to Blog

Irreversible tasks “require stricter liability firebreaks and steeper authority gradients.”

Tomašev et al. (2026)

Picture a procurement team that has given its AI agent two permissions that look equally sensible on paper: approve invoices under $2,000 without review, and cancel duplicate purchase orders it judges redundant. Both score low-risk in the governance model. Then the agent cancels an order for a component that has already shipped, triggering an automated destruction protocol at the vendor's warehouse. The dollar amount is small. The decision cannot be undone.

Nobody asked whether the action was reversible. They asked whether it was risky, and risky is not the same question.

Two papers published in the first quarter of 2026 converge on a version of this problem, and both stop just short of naming what I think is the actual variable at stake. Delegating authority to an AI agent is a governance design question that should be gated by how reversible the decision is, not by the agent's demonstrated capability or a generic risk score.

The Research

Tomašev et al. (2026) model delegation as a trade-off across speed, cost, accuracy, and risk, and they treat reversibility as its own axis rather than folding it into a generic risk score (Emerging). Their framing is blunt: irreversible tasks "require stricter liability firebreaks and steeper authority gradients" regardless of how the task scores on other dimensions (Tomašev et al., 2026). That is a decision-science argument dressed in agent-governance language, and it has an older ancestor.

Dixit and Pindyck's (1994) real-options work established that irreversibility, not probability of loss, is what should trigger formal deliberation over fast execution (Established). A decision with a ten percent chance of a bad outcome that you can reverse tomorrow is a different animal from a decision with a one percent chance of an outcome you can never undo. Risk scoring collapses that distinction. Reversibility restores it.

The organizational half of the argument comes from Saini (2026), whose "Compliant Failure" vignette describes an enterprise with a full pre-deployment governance checklist that still produced repeated near-misses, because oversight lived on paper rather than in real-time supervision (Emerging). Saini's proposed fix is a governance layer where "each agent is associated with a clear business owner, a defined risk profile, and documented decision boundaries" (Saini, 2026). That is a Trust Architecture claim, whether or not the paper uses the term. Lee and See (2004) described the same architecture two decades earlier: appropriate reliance on automation depends on transparency about what a system can do, feedback when it acts, and a mechanism for recalibrating trust as evidence accumulates (Established). Neither Tomašev et al. nor Saini makes reversibility the primary sort. Both treat it as one input among several. I think that ordering is backward.

What This Means in Practice

Most organizations I have studied sort AI agent permissions by category first: customer-facing versus internal, financial versus operational, high-dollar versus low-dollar. Reversibility rarely appears as its own column. A support agent that issues a $30 refund and a support agent that deletes a customer's account both look low-risk by dollar value. Only one of them can be undone with an email.

The fix is not more approval steps. It is a different sort. Before an agent goes live, ask what happens if the action turns out to be wrong: can it be reversed within a day, a week, never? An agent approving marketing copy for a scheduled post is reversible right up until publication; an agent that deletes a production database record, migrates a customer off a legacy system, or cancels a shipped order is not. The dollar value tells you how much a mistake costs. Reversibility tells you whether you get a second chance to fix it.

This changes where the human gate sits. Instead of gating by dollar threshold or task category, gate by the irreversibility test itself, and write down how you scored it. That record is what lets you defend the boundary later and adjust it as the agent's track record accumulates.

Three Things to Take Away

Sort by reversibility before you sort by risk category

A low-risk irreversible action still needs a human gate, even when every other signal says the task is routine (Tomašev et al., 2026). The procurement cancellation above scored low on every axis except the one nobody was measuring.

Draw the delegation boundary before an incident forces it

Saini's (2026) Compliant Failure case shows that formal governance without embedded real-time supervision still produces near-misses; the checklist existed, the boundary did not. Weick and Sutcliffe's (2015) high-reliability-organization research makes the same point across industries: resilient organizations design for anticipated failure modes in advance rather than waiting for an incident to reveal where the line should have been (Established).

Reserve guardrail agents for the irreversible tier only

Tomašev et al. (2026) warn that overseers flooded with verification requests eventually default to heuristic approval, which defeats the purpose of the gate (Emerging). Saini's (2026) guardrail agents are designed to intervene only when uncertainty or impact crosses a threshold; put one on every action and you have rebuilt the approval bottleneck you were trying to remove.

My Two Cents

Risk scores are popular because they feel objective and produce a tidy number. Reversibility is popular with no one, because it forces a harder conversation about what the organization can actually live with if the agent is wrong. I suspect that is exactly why both papers mention it without centering it. It is easier to build a dashboard around risk than to sit in a room and decide, in advance, which mistakes you can afford to have made. The test I apply is simple: if we cannot undo this by Friday without real cost, a human signs off, no matter what the risk model says.

Read to Learn More

Fügener, Grahl, Gupta, and Ketter (2022) study the cognitive mechanics of human-AI delegation empirically, which grounds the trust-calibration argument in lab evidence rather than framework language.

Swayne (2026), writing for AI Insider, offers an accessible summary of the DeepMind delegation paper's trust-efficiency frontier for readers who want the concepts before the technical appendix.

References

Dixit, A. K., & Pindyck, R. S. (1994). Investment under uncertainty. Princeton University Press.

Fügener, A., Grahl, J., Gupta, A., & Ketter, W. (2022). Cognitive challenges in human-artificial intelligence collaboration: Investigating the path toward productive delegation. Information Systems Research, 33(2), 678-696. https://pubsonline.informs.org/doi/10.1287/isre.2021.1079

Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50-80. https://journals.sagepub.com/doi/10.1518/hfes.46.1.50_30392

Saini, S. (2026). Governing the agentic enterprise: A new operating model for autonomous AI at scale. California Management Review Insights. https://cmr.berkeley.edu/2026/03/governing-the-agentic-enterprise-a-new-operating-model-for-autonomous-ai-at-scale/

Swayne, M. (2026). DeepMind study proposes rules for how AI agents should delegate. AI Insider. https://theaiinsider.tech/2026/02/17/deepmind-study-proposes-rules-for-how-ai-agents-should-delegate/

Tomašev, N., Franklin, M., & Osindero, S. (2026). Intelligent AI delegation [Preprint]. arXiv. https://arxiv.org/abs/2602.11865

Weick, K. E., & Sutcliffe, K. M. (2015). Managing the unexpected: Sustained performance in a complex world (3rd ed.). Wiley.