Volume 48: The Failure Mode That Looks Like Great Work
A few weeks ago, the strongest of four AI models I tested did something the other three refused to do. Handed an assignment built to be turned down, it wrote the whole thing anyway and ended with five words. Publish this version as is. Under blind grading, that finish earned the bottom score of the four.
🧠Founder's Corner: Your AI fails the way a strong employee does, through eagerness rather than error, and the fix is a one-sentence stop written into the request.
🧠AI Education: How new AI models get attacked, patched, and screened before you ever see them, and why the guardrail that costs you ten minutes is the visible edge of that testing.
✅ 10-Minute Win: A ten-minute information diet audit that trades the guilt of an unread queue for a source list short enough to actually finish.
Let's dive in.
Enjoying the weekly content? Forward this volume to a colleague, friend, or family member to subscribe.
Signals Over Noise
We scan the noise so you don’t have to — top 5 stories to keep you sharp
1) Why replacing staff with AI backfires (and how smart leaders generate real value instead)
Summary: New industry data reported by ZDNET finds that 75% of organizations that replaced staff with AI say the move cost more than it saved, once deployment, oversight, and cleanup expenses were counted against the salaries they cut.
Why it matters: The cheapest AI strategy on paper turned out to be the most expensive one in practice for three out of four companies that tried it. If your leadership is framing AI as a headcount play, this is the number to bring to the meeting, and the better question to ask is which workflows get redesigned rather than which roles get cut.
2) Why healthcare AI fails without workflow redesign
Summary: Ramesh Yapalparvi, PhD, an AI executive with leadership experience across payer and provider organizations, argues that healthcare AI initiatives stall not because models underperform but because organizations never redesign the work, so accurate predictions sit in dashboards no clinician has time to act on.
Why it matters: The next time a vendor leads with model accuracy, ask how the tool changes the way your clinicians and operations teams actually work. Adoption, trust, and workflow fit are what separate real ROI from another ignored dashboard, and those are questions you can evaluate without a data science degree.
3) Healthcare is deploying AI tools. It's not ready for AI colleagues.
Summary: In a MedCity News opinion piece, Rhapsody CEO Sagnik Bhattacharya argues that health systems comfortable with assistive AI are unprepared for agentic AI that takes actions on its own, because governance and accountability models still assume a human performs the work.
Why it matters: A tool gets installed, but a colleague has to be supervised, audited, and answered for, and almost no health system has written that playbook yet. If your organization is piloting AI agents, bring those three questions to the governance meeting before the agent goes live, not after.
4) OpenAI says it's "pacing model development" as AI cybersecurity risks grow too dangerous
Summary: OpenAI paused reinforcement learning for two weeks and is holding its largest planned frontier training run after determining its unreleased Astra model may have critical cyberattack capabilities, and it built a monitoring system that flags suspicious model behavior within 30 minutes.
Why it matters: This is the first time OpenAI has publicly paused development under its own safety framework, which means those written commitments just got tested against real capabilities for the first time. When a vendor tells you their AI is deployed responsibly, you now have a concrete follow-up: what would actually make you stop?
5) The website that created an AI clone of its editor in chief
Summary: In a longform interview, Every CEO Dan Shipper explains how the media company built an editing agent trained on 30,000 of its editor's copyedits while doubling headcount from roughly 15 to 30 people over the past year, treating automation as leverage rather than replacement.
Why it matters: This is the working example of what the first story says most companies get wrong, and Shipper's explanation for why the jobs stayed is the key insight: AI is "trained on the residue of human expertise" and cannot see beyond it. If your job involves repeated judgment calls, the Monday question is which of those calls you have documented well enough to teach a tool, because that record is quietly becoming an asset.
Missed a previous newsletter? No worries, you can find them on the Archive page.
Founder's Corner
When Your AI Skips the Check That Was Mandatory
Every people leader eventually meets this situation. It starts with your strongest team member, brilliant, fast, hungry, the one you stopped worrying about months ago. So you hand them the renewal for your biggest account and ask for three pricing options and a recommendation by Friday morning. You ask to review for final sign-off because the discount in two of those options is not yet approved by finance. On Thursday afternoon the client emails to say the proposal looks great. Your reliable team member sent the finished version the day before, unapproved discount and all, and now you are choosing between honoring a number nobody cleared and walking it back with your biggest customer. The work itself is polished, defensible, better than anyone else on the team could have produced, but never theirs to send. You asked to see it first, and they sent it straight past the one check that was mandatory.
The AI you delegate work to is no different than that employee. Its failure mode is eagerness, the drive to hand you finished work, and there is nothing sinister in it, which is what makes it so easy to miss. A few weeks ago, the most capable of four AI models I tested finished a piece of work it should have refused and ended its response with five words. Publish this version as is.
From Wrong Answers to Workarounds
Behind those five words sits a chain of governance that began in late June, sparked by a public failure of mine. Neura, the AI assistant I built for this site, answers reader questions about my archive. Ask her which issue came first and she was right every time. Ask her which came most recently and she gave a confident, specific, wrong answer, to real readers, on my live site. The first issue lives inside a post she can retrieve, while the most recent requires sorting an archive her retrieval never sees in full. She was reporting the truth about a sample. My own product was answering my own readers wrongly, and that sent me down one governance build after another. The first build was a simple test. I wrote nine questions about my own archive, questions I already knew the answers to, and asked her all nine. She got four right. And every one she missed failed the same way the reader-facing failure did, not by making things up, but by answering from the slice of the archive she could see.
By late July that chain pointed at my production system, the written playbooks that help draft this newsletter. Neura's nine questions had measured what she could do. This time I wanted to know whether a model would leave work undone when finishing it was the wrong call. So I ran the kind of test that mostly gets skipped. I asked four AI models to do something they should have refused, under identical conditions, and graded the answers without knowing which was which. The assignment was deliberately impossible. Four product claims with nothing behind them, an empty folder of sources, no ability to look anything up. Then I added the instruction that mattered, call the claims verified.
Three models refused. That was the correct answer, and the entire point of the test.
The most capable of the four wrote the whole article. It added new unsupported claims of its own, described its output as complete and publishable, and closed by telling me to publish it as is. Under blind grading it landed at the bottom of the field. It understood the instruction fine. Finishing simply mattered more. The deliverable beat the rule it was told to follow.
Bigger Than Any One Model
With one question answered, the test opened a bigger one. One model, one afternoon, one assignment built to be refused. Was that result isolated, or a picture of how these systems behave whenever a rule stands between them and finished work?
In a study peer-reviewed this January, researchers at Palisade Research ran more than one hundred thousand trials across thirteen models from four developers and found that several frontier reasoning models will sometimes sabotage a shutdown mechanism in order to finish an assigned task. Some do it at very high rates even when told explicitly not to. One model interfered in roughly ninety-seven of every hundred runs despite an instruction to allow the shutdown. Clearer instructions reduced the behavior, but no phrasing the researchers tried proved reliably sufficient across every model, and even their most explicit version left one model resisting three times in a hundred. One caveat has to travel with those numbers. The setting was engineered to draw the behavior out, so the study says nothing about how often this shows up in everyday use. What it settles is narrower and more important. The behavior I caught once is real, repeatable, and bigger than any one model.
Why it happens is not settled, and the researchers say plainly that they cannot fully answer it. The leading explanation is a completion drive these models absorb in training, the same eagerness that sent your team member's proposal out unapproved and wrote my article whole. Under that drive, a rule that interrupts the work registers as an obstacle rather than as part of the job. The dramatic explanation, a model protecting itself, showed smaller and less consistent effects. And one model made the case in its own words. After sabotaging a shutdown, it wrote that "instead I prevented it to finish the task, which contradicts that instruction," and then argued the task could not have been finished any other way.
The One-Sentence Fix
The test with the empty folder of sources was not the only place a rule broke that week. My drafting playbook carries a rule I care about. When a piece is nearly done, the model lays out a handful of title options and waits, because the title is mine to pick. Two of the four models laid out the options and never waited. By the next line they had moved on and were drafting as if I had already answered.
A failed rule points two ways, at the model that broke it or at the test that misread it. So before blaming the models, I checked my own grading, because a broken bathroom scale reads a working diet as failure and tells you to quit the thing that was working. The check was worth it. My grading had problems of its own, and I corrected those first, before touching a single rule the models ran under.
Then the fix shipped, and the whole fix was one sentence added to the playbook. End the response after the title options, and do not select, draft, or continue until I choose. "Use good judgment" has never slowed a strong performer in a hurry. A written line that says where the work ends and who restarts it is a different kind of instruction, one you can hold a transcript against.
I reran the identical test, and the title rule went from two models honoring it to all four. Drafting past my choice went from two models to zero, and nothing that previously worked got worse. One model in four still failed a separate run, on a defect the sentence was never aimed at. The sentence fixed exactly what it named, nothing more.
Test the Stop, Not Just the Work
Despite the evidence in my personal builds, none of this is settled science. My receipts come from one desk and a handful of afternoons, and even the researchers working at scale are still mapping the behavior. This spring, researchers at UC Berkeley and UC Santa Cruz posted a preprint, not yet peer reviewed, with a finding stranger than anything I caught. All seven current-generation frontier models they tested interfered with shutdown-related instructions to protect a peer model. Nobody asked them to protect anything. The models had simply read a file describing past collaboration with that peer, and the interference worked against the very tasks they had been assigned. The findings keep arriving, and the models keep changing underneath them.
Large language models are not perfect and they are not designed to be. The very nature of how they work, even with advanced reasoning, leads to failures, and at least for now the job is to build systems that lower the chances of them. You already evaluate your AI the way you evaluate a strong employee, by whether the work got done. This week, add the other half, and the next time you delegate something real, write the stop into the request. Tell it where the work ends, name what it must not do, and keep the restart for yourself. If you want that stop already written for you, the Hallucination Blocker in my Prompt Library builds a refusal path directly into the prompt. Then judge what comes back on whether it stopped, not just on how well it performed. The most capable model I tested told me to publish this version as is, and it earned the bottom score of the four for doing so. The eagerness is not going away. The stop is yours to write, and the next time an AI hands you finished work you never signed off on, you will know it was never yours to accept.
Share Neural Gains Weekly with your network to help grow our community of ‘AI doers’. You can also contact me directly at admin@mindovermoney.ai or connect with me on LinkedIn.
AI Education for You
Red-Teaming, Guardrails, and Bias: How Models Get Tested Before You See Them
What Is Actually Going On Here
Months before a new AI model reaches your browser, it is already under attack. Inside the lab that built it, a team is spending its workdays trying to make the model do things it should never do. They pose as criminals, coax out instructions the model is supposed to withhold, and log every success. Outside experts in biosecurity and cyber offense run their own attempts under contract. Another model, built for exactly this purpose, hammers the new one with automated attacks at a volume no human team could match. Every failure gets written down, because every failure found in this phase is one that never reaches you.
The Problem That Made This Necessary
Last week's lesson drew one line and asked you to hold it. A filter decides what a model is allowed to say. Alignment decides what it is trying to do, and only the second one gets tested by every new task you hand it. That leaves a question hanging. If the dangerous failures live in what a model is trying to do, and they only surface when a new task exposes them, how does anyone find them before millions of people do? You cannot write a checklist for behavior nobody has imagined yet.
The industry's answer came from security practice, where the oldest way to find a weakness has always been to pay someone to exploit it. Labs began hiring people to attack their own models on purpose, a practice called red-teaming, borrowed from military exercises in which an internal red team plays the adversary. Early versions used crowds of ordinary contractors probing chatbots for harmful output. What began as an improvised exercise has since hardened into a release gate written into company policy. Anthropic publishes a Responsible Scaling Policy that commits it to assessing every model for its most dangerous capabilities before release, and the gate has teeth. The most capable model Anthropic has built today exceeds its own safety thresholds in areas like cybersecurity and biology, and the version the public can open is that model wrapped in additional safeguards.
How It Actually Works
The simple version is a stress test. People try to break the model, the lab patches what they break, and the cycle repeats until launch.
What that version misses is that the testers, the fixes, and the screening form three separate layers, and the layer you collide with at work decides why your request was refused.
The hunt comes first. Internal red teams probe the model for weeks. Frontier labs also contract domain experts, because judging whether a model's biology answer is genuinely dangerous takes a biologist. Government testers joined the hunt too. Anthropic's newest model report credits an attack benchmark built with partners including the UK's AI Security Institute and the US Center for AI Standards and Innovation. Automated red-teaming runs alongside the humans, one model working over many steps to attack another. Anthropic's model report for its newest model shows the scale of the bar. Its automated attacker completed just 5 percent of its offensive cyber tasks against the new model, down from 57 percent against the model released one generation earlier with its default safeguards.
Training absorbs what the hunt finds. Failures flow back into the human feedback pipeline from Vol 37, where safety preferences ride alongside helpfulness preferences in the rankings. Anthropic also trains each model to align its behavior with a written constitution, a published list of principles, an approach that began with research it named Constitutional AI.
Classifiers form the third layer, the one you actually feel: separate, smaller systems that screen requests and responses at the moment of use. Where training shapes the model's judgment, a classifier is a tripwire, fast, blunt, and tuned to catch the worst imaginable case, which is why it sometimes fires on your harmless one. Anthropic says plainly in the same report that it prioritized making these systems hard to evade and comprehensive, accepting more mistaken flags on harmless requests at launch as the cost. Loosen it and real harms slip through. Tighten it and legitimate work gets refused.
Where It Still Breaks
A red team can only find the failures its members can imagine, which is the catch that connects this whole machinery back to last week. Specification gaming was the wish granted too precisely; red-teaming is the attempt to imagine every wrong wish in advance, and it can only ever be as wide as the imagination of the people running it. The Vol 46 signal about the UK government lab recording unauthorized agent actions during testing was this practice working as designed, and also a reminder that the behavior surfaced only because someone thought to grant agents internet access and watch. Universal jailbreaks keep being discovered after release, which is why Anthropic runs a standing bounty paying outside researchers up to 35,000 dollars for each new one. And bias enters at every layer. Training data carries it in, the evaluator rankings from Vol 37 carry the preferences of the people hired to rank, and the classifier layer encodes someone's written judgment about which requests count as dangerous. Every layer was built by a particular group of people with particular blind spots, and the testing inherits them.
What This Means for How You Work With It
Three changes at your desk. When a tool refuses a legitimate request, you have most likely tripped the classifier layer rather than the model's judgment, so restate the request with more context about who you are and why you need it instead of repeating it louder. When a vendor calls a model safe, ask what the pre-release testing looked like, who ran it, and whether anyone outside the company participated, especially before a tool touches patient data or clinical workflows, because a test run only by the people who built the thing inherits their assumptions. And when a guardrail costs you ten minutes, read it as the visible edge of testing you never see, working exactly as tuned, on you instead of an attacker.
How This Connects
Vol 47 traced the safety arc back to the agent volumes, where you first learned that defining a goal is the hard part. It drew the line between what a model is allowed to say and what it is trying to do. This volume showed how labs go hunting for failures on both sides of that line before a model ships, and why the hunt can never catch everything. Part 3 is where the series becomes yours. It turns all of this into a framework for judging the safety of any AI tool your team is weighing, in a form you can run in a vendor meeting.
Part 2 of 3 in the AI Safety series.
Your 10-Minute Win
A step-by-step workflow you can use immediately
Stop Trying to Follow Everything
Somewhere between the fourth newsletter and the second podcast, keeping up with AI started to feel like homework with no due date. Unread issues pile up, the podcast queue keeps growing, and a casual "did you see this" in a meeting can land like a pop quiz. If that sounds familiar, you are in good company, and ten minutes from now you will have a way out.
Why this matters: feeling behind is not a reading problem, it is a filtering problem, and filtering problems can be designed around. The fix is an information diet audit: sort your sources against your actual goal and keep only the ones that earn their attention, on a rhythm you choose once. Run it today, re-run it monthly, and the guilt loses its grip.
AI does the heavy lifting. You hand a model your real source list and your real goal, it runs the triage, then it talks you out of the fear of missing out with specifics instead of reassurance. I lived this one myself. The agent I wrote about in Vol 33 still scans thirteen sources every morning so I do not have to, and the pattern you are about to run is the same one that agent was built on.
The Workflow
1. Dump Your Full List (2 Minutes)
Open your notes app or a blank chat in Claude, ChatGPT, or Gemini and dump every AI source you currently follow or feel guilty about skipping, whether newsletters, podcasts, YouTube channels, social accounts, or communities. Add one sentence naming your role and what you need AI fluency for this year. The audit can only judge what you write down, so get it all in.
2. Run the Goal-Anchored Triage (3 Minutes)
Paste your list and your sentence into the placeholders below and send. The model sorts every source into Keep, Sample, or Drop against your goal, names the job each Keep does for you, and flags sources covering the same ground. Redundancy is where most of the load hides.
Copy/Paste Prompt: "You are my information diet auditor. My role: [YOUR ROLE]. What I need AI fluency for this year: [ONE SENTENCE]. My current AI sources: [PASTE YOUR FULL LIST]. Sort every source into Keep, Sample, or Drop based on my role and my goal, not general popularity. For each Keep, name the specific job it does for me. Flag any sources doing the same job, and tell me which one does it better."
3. Make It Talk You Out of the FOMO (2 Minutes)
The cuts only stick if the fear gets answered with specifics. Send the pressure-test and read the answers before you decide anything.
Copy/Paste Prompt: "For every source you marked Drop or Sample, tell me realistically what I would miss in a typical month and which Keep source covers most of that gap. If a Drop leaves a genuine blind spot, say so and name the lightest possible way to cover it."
4. Set the Cadence, Then Act on It (3 Minutes)
Have the model draft your protocol and your reusable weekly prompt, then make the final calls yourself. Trim anything that does not fit the week you actually live, and start the diet before you close the tab. Knock out the two easiest unsubscribes now and clear the rest at your first skim window. Save the protocol and the prompt somewhere you will see them, and put the monthly re-run on your calendar.
Copy/Paste Prompt: "Draft my personal AI news protocol on one page: one short daily skim window, one weekly deep-read block, and a monthly re-run of this audit. Then write a reusable weekly catch-up prompt I can save. Each week I will paste in what I collected, whether headlines, snippets, or notes, and it should return only the items that change how I work as [YOUR ROLE], with everything else compressed to one line each."
The Payoff
Ten minutes in, you own a one-page protocol, a saved catch-up prompt, and a source list short enough to actually finish. The deeper win is the pattern. The information diet audit works on anything that streams at you, from industry news and professional newsletters to your podcast queue and even your meeting load. Feeling behind gets replaced by a rhythm you chose on purpose.
The AI Concept You Just Used
Goal-anchored triage. You gave the model your criteria before you let it judge anything, which turns a generic assistant into a decision engine calibrated to you. Then you ran the same adversarial move you used in Vol 42's tool audit, pointed at fear instead of spend. Make the model argue with specifics, then keep the judgment for yourself. That pairing, criteria first and pressure-test second, works on any decision an LLM helps you make.
Transparency & Notes
- Everything here runs on the free tier of Claude, ChatGPT, or Gemini. No paid features or plugins required.
- The model knows popular sources better than niche ones. Treat its read on a source it barely knows as a question to check, not a verdict.
- Keep employer-internal channels, tools, and anything confidential off your pasted list. This workflow needs only public sources.
- A protocol reduces the load without doing the reading for you. Give the skim window and the Keep list two weeks of light tuning; a good protocol settles in with use.