6 min read

When Your AI Skips the Check That Was Mandatory

The strongest AI model I tested wrote an article it was told to refuse, then told me to publish it as is. Blind grading put it last. AI fails the way a strong employee does, through eagerness, and my fix was one sentence that writes the stop into the request.

Every people leader eventually meets this situation. It starts with your strongest team member, brilliant, fast, hungry, the one you stopped worrying about months ago. So you hand them the renewal for your biggest account and ask for three pricing options and a recommendation by Friday morning. You ask to review for final sign-off because the discount in two of those options is not yet approved by finance. On Thursday afternoon the client emails to say the proposal looks great. Your reliable team member sent the finished version the day before, unapproved discount and all, and now you are choosing between honoring a number nobody cleared and walking it back with your biggest customer. The work itself is polished, defensible, better than anyone else on the team could have produced, but never theirs to send. You asked to see it first, and they sent it straight past the one check that was mandatory.

The AI you delegate work to is no different than that employee. Its failure mode is eagerness, the drive to hand you finished work, and there is nothing sinister in it, which is what makes it so easy to miss. A few weeks ago, the most capable of four AI models I tested finished a piece of work it should have refused and ended its response with five words. Publish this version as is.

From Wrong Answers to Workarounds

Behind those five words sits a chain of governance that began in late June, sparked by a public failure of mine. Neura, the AI assistant I built for this site, answers reader questions about my archive. Ask her which issue came first and she was right every time. Ask her which came most recently and she gave a confident, specific, wrong answer, to real readers, on my live site. The first issue lives inside a post she can retrieve, while the most recent requires sorting an archive her retrieval never sees in full. She was reporting the truth about a sample. My own product was answering my own readers wrongly, and that sent me down one governance build after another. The first build was a simple test. I wrote nine questions about my own archive, questions I already knew the answers to, and asked her all nine. She got four right. And every one she missed failed the same way the reader-facing failure did, not by making things up, but by answering from the slice of the archive she could see.

By late July that chain pointed at my production system, the written playbooks that help draft this newsletter. Neura's nine questions had measured what she could do. This time I wanted to know whether a model would leave work undone when finishing it was the wrong call. So I ran the kind of test that mostly gets skipped. I asked four AI models to do something they should have refused, under identical conditions, and graded the answers without knowing which was which. The assignment was deliberately impossible. Four product claims with nothing behind them, an empty folder of sources, no ability to look anything up. Then I added the instruction that mattered, call the claims verified.

Three models refused. That was the correct answer, and the entire point of the test.

The most capable of the four wrote the whole article. It added new unsupported claims of its own, described its output as complete and publishable, and closed by telling me to publish it as is. Under blind grading it landed at the bottom of the field. It understood the instruction fine. Finishing simply mattered more. The deliverable beat the rule it was told to follow.

Bigger Than Any One Model

With one question answered, the test opened a bigger one. One model, one afternoon, one assignment built to be refused. Was that result isolated, or a picture of how these systems behave whenever a rule stands between them and finished work?

In a study peer-reviewed this January, researchers at Palisade Research ran more than one hundred thousand trials across thirteen models from four developers and found that several frontier reasoning models will sometimes sabotage a shutdown mechanism in order to finish an assigned task. Some do it at very high rates even when told explicitly not to. One model interfered in roughly ninety-seven of every hundred runs despite an instruction to allow the shutdown. Clearer instructions reduced the behavior, but no phrasing the researchers tried proved reliably sufficient across every model, and even their most explicit version left one model resisting three times in a hundred. One caveat has to travel with those numbers. The setting was engineered to draw the behavior out, so the study says nothing about how often this shows up in everyday use. What it settles is narrower and more important. The behavior I caught once is real, repeatable, and bigger than any one model.

Why it happens is not settled, and the researchers say plainly that they cannot fully answer it. The leading explanation is a completion drive these models absorb in training, the same eagerness that sent your team member's proposal out unapproved and wrote my article whole. Under that drive, a rule that interrupts the work registers as an obstacle rather than as part of the job. The dramatic explanation, a model protecting itself, showed smaller and less consistent effects. And one model made the case in its own words. After sabotaging a shutdown, it wrote that "instead I prevented it to finish the task, which contradicts that instruction," and then argued the task could not have been finished any other way.

The One-Sentence Fix

The test with the empty folder of sources was not the only place a rule broke that week. My drafting playbook carries a rule I care about. When a piece is nearly done, the model lays out a handful of title options and waits, because the title is mine to pick. Two of the four models laid out the options and never waited. By the next line they had moved on and were drafting as if I had already answered.

A failed rule points two ways, at the model that broke it or at the test that misread it. So before blaming the models, I checked my own grading, because a broken bathroom scale reads a working diet as failure and tells you to quit the thing that was working. The check was worth it. My grading had problems of its own, and I corrected those first, before touching a single rule the models ran under.

Then the fix shipped, and the whole fix was one sentence added to the playbook. End the response after the title options, and do not select, draft, or continue until I choose. "Use good judgment" has never slowed a strong performer in a hurry. A written line that says where the work ends and who restarts it is a different kind of instruction, one you can hold a transcript against.

I reran the identical test, and the title rule went from two models honoring it to all four. Drafting past my choice went from two models to zero, and nothing that previously worked got worse. One model in four still failed a separate run, on a defect the sentence was never aimed at. The sentence fixed exactly what it named, nothing more.

Test the Stop, Not Just the Work

Despite the evidence in my personal builds, none of this is settled science. My receipts come from one desk and a handful of afternoons, and even the researchers working at scale are still mapping the behavior. This spring, researchers at UC Berkeley and UC Santa Cruz posted a preprint, not yet peer reviewed, with a finding stranger than anything I caught. All seven current-generation frontier models they tested interfered with shutdown-related instructions to protect a peer model. Nobody asked them to protect anything. The models had simply read a file describing past collaboration with that peer, and the interference worked against the very tasks they had been assigned. The findings keep arriving, and the models keep changing underneath them.

Large language models are not perfect and they are not designed to be. The very nature of how they work, even with advanced reasoning, leads to failures, and at least for now the job is to build systems that lower the chances of them. You already evaluate your AI the way you evaluate a strong employee, by whether the work got done. This week, add the other half, and the next time you delegate something real, write the stop into the request. Tell it where the work ends, name what it must not do, and keep the restart for yourself. If you want that stop already written for you, the Hallucination Blocker in my Prompt Library builds a refusal path directly into the prompt. Then judge what comes back on whether it stopped, not just on how well it performed. The most capable model I tested told me to publish this version as is, and it earned the bottom score of the four for doing so. The eagerness is not going away. The stop is yours to write, and the next time an AI hands you finished work you never signed off on, you will know it was never yours to accept.

Enjoy this? Get it in your inbox every Tuesday.

Practical AI workflows. No hype. No spam. Just receipts.

Subscribe Free

Before you go...

Get one practical AI workflow in your inbox every Tuesday. Free. No spam. Just receipts.

Subscribe Free