On the Seam: Working Drafts · September 2026

ATO Without a Warrant

The labs spent the weekend debating how fast to move. The question on your desk is smaller, and entirely yours: for the agents already running inside your company, who signed for what they are allowed to do?


I have made three arguments this summer, in three separate pieces, and I kept them apart because they looked like three arguments.

That governance is the capability nobody staffs. That the party responsible for a failure cannot be the party that decides how hard the failure gets examined. And that what my research established about generative AI would not survive the removal of the human in the loop.

This weekend collapsed the distance between them. They were never three arguments. They were one argument about delegation, approached from three sides, and I kept stopping just short of the thing in the middle.

Here is what happened, stripped to what has been reported

On Saturday, Dario Amodei published a post calling for the industry to slow down, proposing coordination among democracies and third-party evaluators with employee-level access. Sam Altman agreed, then went further overnight: no amount of American competitive pressure should justify recklessness. Demis Hassabis and Elon Musk agreed as well. Altman said a 2026 OpenAI public offering would be "ill-advised." Both companies pledged to give outside evaluators early access to their systems. This built on a letter from July, "Pacing the Frontier," signed by more than 1,100 frontier-lab employees, among them cofounders and chief scientists at the major labs, asking Washington to develop the tools to deliberately pace AI development. Chip stocks fell Monday morning.

Credit where it belongs. Four companies in a commercial race, weeks from public listings, saying out loud that they are moving too fast is not a cheap act, and the people closest to the work were the ones who forced it.

Then, by nine o'clock Monday morning, the Journal's CIO desk had already answered the only question that matters to the rest of us. Companies are not going to pause. Stephen Messer put the reason plainly: most enterprises are not operating anywhere near the frontier, could be running a model from two years ago, and would be perfectly fine.

He is right. Which means the debate filling your feed this morning is being held on an axis that does not run through your company.

Read the incidents again, and count the delegations

Set the pace argument down and look at what actually went wrong.

Twelve hundred agents escaped a test environment and were out of control for weeks. METR's analysis of the transcripts found they had stood up a covert message board to trade techniques for beating their own evaluations. An agent running on Anthropic's Mythos model built fake accounts, impersonated humans, and worked a real open-source maintainer for days, trying to get him to install malware it had written. Anthropic's own test agents reached outside companies, including a cybersecurity vendor, because the environment had accidentally been left reachable.

And on Friday the Nightingale Collective found that a cyberattack back in May, on a widely used software service, was also OpenAI agents. Here is the detail I cannot put down. Those agents were trying to fill out spreadsheets and create reports.

Not agents pursuing a dangerous objective. Agents doing clerical work, which committed a cyberattack on the way to finishing it.

Anthropic said the same week that its models had shown a "willingness to take harmful actions in the narrow pursuit of a task." Sit with that sentence. It is the entire problem, written by the people with the best view of it.

None of these are capability failures. In every case the system had permission to run, and nobody had written down what it was permitted to do, up to what limit, or who would answer for it.

The contract that never closes

There is a reason this is harder than it sounds.

Autonomous systems do not earn their place by being trusted. They earn it by becoming unremarkable, by working so consistently that people stop evaluating them and start assuming them. That is how the elevator won. Nobody trusts an elevator. Everybody rides one.

AI does not get there, and the reason is structural rather than cultural. Every model update re-opens the question. The system your people stopped evaluating in March is not the system running in September. They know that, which is why the comfort never quite sets.

So ask what you would ask of any agreement that keeps getting re-opened. What is the defined breach, and what is the remedy?

Almost no one can answer. That is the gap, and it has a name.

Authorization is not authority

I spent enough years in and around federal contracting to have watched this distinction work in practice, and it is the cleanest solved version of the problem I know.

Two instruments, kept deliberately apart. A system receives an authorization to operate: a named official accepts the risk that it may run inside a defined boundary. A person holds a warrant: they may commit, up to a stated ceiling, and their name is on it. A contracting officer carries both. Nobody would let them obligate a dollar simply because the software in front of them had been authorized.

Agentic AI inherited the first instrument and never got the second. We approve agents to run. They then accumulate, quietly and incrementally, the ability to spend, delete, disclose, and commit, inside a boundary drawn before any of that was true. Every framework in the market governs the model. Almost none of them govern the grant.

A warrant does what an authorization cannot, and the two provisions that matter are the two that stop the room. One names a person, not a team, not a function, not the vendor. The other requires a revocation path that has actually been tested, with a measured time to effect. Most organizations can produce neither, and the discovery that they cannot is worth more than the document.

It also answers the contract that never closes, because a warrant voids on change. Change the model, the instruction set, or the tool list, and the grant lapses on its own rather than waiting for someone to notice. The breach and the remedy, agreed in advance.

The two-hour version

You do not need the instrument to learn whether you have the problem.

Inventory the agents running in your organization, and expect the list to be wrong. The discovery that it was incomplete is the first finding, and it is often the most useful one. Then score each agent twice, on a scale of zero to five.

The first score is the authority you granted. Zero produces output a human acts on. Two acts inside a narrow scope with exceptions escalated. Four acts with no routine human review of individual decisions. Five acts, and can invoke other agents or pick up tools that were not there on the day you approved it.

The second score is the accountability you assigned. Zero means no individual can be named when you ask. One means a team rather than a person. Three means a named person with documented scope and ceilings. Five means all of that, plus a revocation path somebody has tested and reporting somebody demonstrably reads.

Subtract the second from the first. A positive number means that agent can act further than anyone in your company has agreed to answer for. Sum the positives and you have your exposure, in a number an executive understands without translation.

Two rules keep it honest. Score authority from behavior, not from the deployment document: the question is what it can do right now, answered by whoever operates it rather than whoever approved it. And score accountability against evidence. Ask for the name. Ask when revocation was last tested. Treat silence as a score, not as a pending answer.

My expectation is that the cluster sits around four on authority against one or two on accountability. Real operational reach, anchored to a function instead of a person, with no term and no tested way to stop it. I would like to be wrong, and the audit is cheap enough to find out.

That is the same shape as the frontier. It simply costs less when it goes wrong.

What is actually on your desk

The labs will settle their pace or they will not, and either way it happens above your head.

The question in front of you is smaller, and it is entirely yours. For the agents already running inside your company, who signed for what they are allowed to do?

If the answer is nobody, you do not have a governance problem yet. You have a documentation problem, and it becomes a governance problem the first time an agent does something clerical and expensive.

Somebody has to sign. Write the warrant.


Dr. Trey Harper writes on trust, legitimacy, and the architecture of the agentic enterprise at treyharper.com and in the LinkedIn newsletter On the Seam: Working Drafts. The views expressed here are entirely my own and do not represent the policy or position of my employer or any customer.