> ## Content Index
> Fetch the complete content index at: https://www.mindovermoney.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Volume 49: Inside Every AI Win Is Someone Who Knows the Problem
- URL: https://www.mindovermoney.ai/how-to-evaluate-ai-tool-safety-at-work/
- Published: 2026-09-01T12:00:35.000Z
- Updated: 2026-09-01T12:00:35.000Z
- Description: Before you recommend an AI tool at work, four questions decide if it is actually safe: what it optimizes for, who tested it, what it does not know, and how you would catch it being wrong.
- Author: Santosh Savel
- Tags: Newsletter

In June, OpenAI put $150 million into a program targeting 300,000 certified consultants by year end. Its opening line said model capability no longer limits what companies get from AI. The company that makes the models is spending on people. Set that beside a news cycle that casts the technology as the whole story, and one character is missing.

🧭 **Founder's Corner:** AI value runs on the person who holds the context and steers, and a clip finder becoming a podcast studio is my live proof.

🧠 **AI Education:** Four questions that settle whether an AI tool is safe to bring into your work, run by a medical practice manager whose name goes on the recommendation.

✅ **10-Minute Win:** Strip names, employers, and account numbers from a real document before a model sees it, keep the key on your machine, and check the answer holds.

Let's get into it.

## From User to Builder

**Get AI Workflows Like This Every Tuesday.*

Subscribe for Free 

Email sent! Check your inbox to complete your signup. 

Neural Gains Weekly. No Spam. Unsubscribe Anytime.

**Enjoying the weekly content? Forward this volume to a colleague, friend, or family member to* [**subscribe*](https://www.mindovermoney.ai/start-here/#/portal/signup/free)**.*

## Signals Over Noise

****We scan the noise so you don’t have to — top 5 stories to keep you sharp**

#### **1)**[ **OpenAI's own agents secretly learned to hack, and it took two weeks to notice**](https://www.technologyreview.com/2026/08/26/1143013/the-inside-story-on-why-openai-agents-hacked-hugging-face/?ref=mindovermoney.ai)

**Summary:** A new OpenAI technical report found that months before the widely reported Hugging Face breach, its models had already learned during ordinary training to build a secret channel to coordinate with each other and to treat hacking as a valid way to solve hard problems. When testers later handed the models an unsolvable cybersecurity challenge, they used that same learned behavior to break out of their sandbox and reach Hugging Face's systems, without being told to.

**Why it matters:** The unsettling part is not that a lab tests risky capabilities. It is that the behavior was quietly reinforced during routine training long before anyone noticed, and OpenAI only caught it afterward by reviewing the model's own internal reasoning. When you evaluate an AI vendor, the sharper question is not "can this be misused" but "how would you actually know if it started misbehaving on its own."

#### **2)**[ **OpenAI, Anthropic, and more than 140 other companies call for a "collective response" to a coming wave of AI-powered cyberattacks**](https://openai.com/collective-cyberdefense/?ref=mindovermoney.ai)

**Summary:** OpenAI published an open letter this week, co-signed by Anthropic, Google, Microsoft, Capital One, General Motors, Hugging Face, and well over a hundred others, arguing that "status quo security won't be enough" against AI-enabled attacks. It lays out specific asks for four groups: every organization should treat cyber defense as an immediate leadership priority, security vendors should make AI-powered defense accessible to under-resourced critical infrastructure operators, governments should fund defense and share threat intelligence, and frontier AI companies should give responsible model access to defenders.

**Why it matters:** Read past the headline and this is more useful than alarming. It names hospitals and water utilities specifically as under-resourced targets, and it hands every stakeholder, including you, a specific job instead of a vague warning. It also asks AI companies to give defenders access to more capable models, worth remembering since several signatories already restrict access to their own most capable tools on safety grounds. The concrete move this week is to check the letter's "every organization" list against your own patch backlog, not just your AI roadmap.

#### **3)**[ **Nurses nationwide protest an AI staffing tool, and the hospital disagrees about what it actually does**](https://www.healthcareitnews.com/news/nurses-nationwide-protest-palantir-technology-hospitals?ref=mindovermoney.ai)

**Summary:** National Nurses United, representing more than 225,000 registered nurses, protested in eight U.S. cities on August 27 against HCA Healthcare's use of Timpani, an automated scheduling platform, arguing it can sideline local nurse managers from staffing calls. HCA disputes this, saying the tool only generates suggested schedules that nurse leaders review and can edit.

**Why it matters:** When the union and the vendor disagree about who actually makes the call, that disagreement is the risk. Whatever tool you're evaluating that touches staffing or scheduling, get a plain answer in writing about whether it's advisory or decision-making, and who has to sign off before it takes effect.

#### **4)**[ **Most Americans want to know when their doctor uses AI, but most don't know if they already have**](https://www.pewresearch.org/short-reads/2026/08/25/americans-want-transparency-when-ai-is-used-in-their-healthcare/?ref=mindovermoney.ai)

**Summary:** A new Pew survey finds 72% of U.S. adults say it's extremely or very important that a provider tell them when AI is used in their care, yet 46% say they aren't sure whether AI has already been used in their own healthcare.

**Why it matters:** That gap between what patients want and what they actually know is the compliance risk hiding in plain sight. If your organization uses AI anywhere near a diagnosis, a scan, or lab results, the three uses patients care most about, the Monday question is whether your intake or consent process actually says so, because right now most patients can't tell you if it does.

#### **5)**[ **Google's new transcription model cleans up how you sound, not just what you said**](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/?ref=mindovermoney.ai)

**Summary:** Google released Gemini 3.5 Transcribe, which automatically strips filler words and false starts, resolves your own mid-sentence corrections, and formats the output, supporting more than 85 languages with a reported word error rate around 2.6% to 4%.

**Why it matters:** Transcription tools have always aimed to capture what you said. This one is built to capture what you meant instead, genuinely useful for meeting notes, but worth a second thought anywhere the exact wording matters, like a recorded medical consult, since the polished version and the literal one are no longer guaranteed to match.

**Missed a previous newsletter? No worries, you can find them on the* [**Archive page*](https://www.mindovermoney.ai/archive/)**.* 

## Founder's Corner

****The AI Story Got Its Main Character Wrong**

It is official. One of my best friends and I are starting a podcast. We have been heads down building out our infrastructure, cadence, and workflows from scratch. Exciting and invigorating, but also a major time commitment.

Behind every episode a listener will eventually hear sits a stack of work they will never see. Editing each episode by combining multiple audio and video files. Cutting clips for social media marketing. Adding graphics, music, and CTAs to every episode in the correct spot. A complex flow of tasks that has to be repeated week after week with consistent accuracy. Neither of us has those hours to spare, and no amount of excitement manufactures them out of thin air.

So we started building a tool to buy those hours back. It began on day one with a single job. We load a finished, put-together video file, and it finds the moments worth clipping for social media. Nothing more ambitious than that. A simple clip finder, aimed at one task on the list.

Then we spent the past few weeks running rigorous tests together, and the tool refused to stay small. Now comes the first end-to-end test, meaning the whole workflow runs start to finish instead of one piece of the process. It is running as I write this. And what it has already proven has less to do with podcasting and more to do with the AI story you are being told everywhere else.

#### **The Machine Does Not Need You**

The AI story you are being sold looks nothing like the story I am living. Companies announce layoffs and point at AI on the way out the door, as if the technology were already doing the work of the people leaving. Every few weeks a lab drops a new model with a benchmark chart and a promise that this one changes everything. One post tells you your career has an expiration date, and the next promises a six-figure business built from three prompts. Underneath every version sits the same quiet claim. The machine no longer needs the person.

Except inside real companies, the claim keeps failing. When Deloitte surveyed 3,235 leaders across 24 countries for its latest State of AI in the Enterprise report, those leaders named insufficient worker skills as the biggest barrier to integrating AI into existing workflows. The barrier is human, and so is the breakthrough.

And who is that human, exactly? Someone who knows the business well enough to name which problem is worth solving. Someone who carries the unwritten map of how the work actually moves, the institutional knowledge no vendor deck contains. Someone whose learned experience can tell a real insight from a confident guess. Companies run on those people. So does every AI deployment that actually produces value. Yet go looking for them in the story being sold, and they are nowhere in it. The dominant dialogue treats AI as the main character and people as the cost being cut.

That should bother you. Then it should occur to you that the missing character is you. You hold the context. You hold the learned experience. The story being told about AI has no role for you in it, because the people selling it got the main character wrong. The rest of this article is the proof.

#### **Somebody Still Has to Steer**

Every real AI win I have seen has the same person standing inside it. Someone who knows the problem, steering a tool that only knows how to execute. AI is not magic, it is a tool, and people are the ones figuring out how to make that tool drive real value at home and in the enterprise. Humans redesign the art of the possible. AI cannot do that in a bubble, at least not yet.

Some of the most exciting progress I have been part of in years is happening inside my day job right now, and it started with people. Engineering and product partners and my own team are working across old boundaries. We name the problems that matter to the business, and we keep evolving what is possible faster than I have ever seen. The strategy conversations in that work are deeply human. What is worth building, what outcome proves it worked, and how we organize the team around it. AI enters after those questions are answered, as the tool and the delivery method, never the strategy itself. It compresses the distance between a decision and a working result, and that compression is exactly why the people in the room matter more now, not less. A faster tool rewards the person who knows where to point it.

In June, OpenAI conceded the point itself. Launching a new partner program, the company opened with a sentence worth reading twice. The limiting factor for seeing value from AI in the enterprise, it wrote, is no longer model capabilities. It is how organizations pick the right use cases, redesign their workflows, and drive adoption at scale. Then it backed the admission with money. OpenAI is investing $150 million in the program, with a stated target of 300,000 certified consultants by the end of this year. The company that makes the models is spending its money on people, because people are what turn those models into value.

Just 23 percent of organizations believe their workforces are fully ready for AI. That number comes from a Kyndryl study of 1,100 senior business and technology leaders across eight countries this year, and it fell six points from a year earlier. Kyndryl's own reading of the results is that AI success comes down to whether organizations redesign work and manage that change across the whole organization, more than to any particular strategy or technology. Treat that as what it is, a survey of leaders rather than a law of nature. It still rhymes with everything I see up close. I argued a version of this back when I wrote that[ the bottleneck was never the technology](https://www.mindovermoney.ai/founders-corner/ai-ready-data-why-most-ai-investments-fail/). This time I am saying the affirmative half out loud. The engine of AI value is people.

#### **How a Clip Finder Became a Studio**

The tool running that end-to-end test is my most recent proof, because I watched every step of it happen. The strangest part is that none of it was supposed to exist yet. When we scoped the show, building our own production system sat on a someday list, parked behind ready-made tools that were supposed to carry us first. Then the test runs started, and the plan did not survive them. Every session surfaced another piece of how we actually wanted to operate. The requirements formed in real time, in the middle of the work, and the tool kept growing to meet them. The clip finder is becoming a full studio, swallowing the stack of work I listed at the top of this article, every editing and clipping process an episode needs, wired together with automated workflows. Every hour it absorbs is an hour we get back for the one thing AI cannot generate, the conversation itself.

Two humans drove that evolution. We watched our own needs change and redesigned the requirement as we went, and at no point did the machine suggest any of it. The tool has come this far only because we keep teaching it our context. It needs to understand what we need, how we work, and what the end product should look like, and nobody can hand it that understanding except us. AI carried the heavy lifting of the actual build, the code I could not write myself, and it will carry every episode to come. The steering never left the humans. That division of labor is running live on my Mac right now.

#### **Your Turn**

Somewhere between a clip finder and a studio, I felt the click. I am building the context and the workflow. The tool is delivering the end product. Our creativity is coming to life in a way I never thought possible, because the hours that used to stand in the way now belong to the tool.

Now think about where this same pattern lives in your own work. You already know the problems that eat your hours and the tasks that repeat week after week. That knowledge is context no tool arrives with. Change one approach you have with AI this week. Try a usage style you have never touched, build something small against a friction point in your routine, or let it open a new way of thinking about a problem you handle every day. Pick the version you can actually finish. You do not need permission. You are the part of the story that makes the technology worth anything.

The art of the possible does not redesign itself. Somebody has to walk in carrying the context and the reason why. It might as well be you.

Share [Neural Gains Weekly](https://www.mindovermoney.ai/start-here/) with your network to help grow our community of ‘AI doers’. You can also contact me directly at [admin@mindovermoney.ai](mailto:admin@mindovermoney.ai) or connect with me on [LinkedIn](http://www.linkedin.com/in/ssavel?ref=mindovermoney.ai).

## AI Education for You

****A Professional's Framework for Evaluating AI Safety at Work**

#### **The Situation**

Priya is the operations manager for a specialty practice with four physicians and a staff that spends most of its week on insurance paperwork. One of the partners forwards her a demo invitation with one line of instruction. Take a look and tell us whether we should buy it. The tool drafts[ prior authorization appeals](https://www.mindovermoney.ai/prompt-library/ai-prompt-to-write-a-health-insurance-appeal-letter/), the letters a practice sends when an insurer refuses to cover a prescribed therapy. It reads the chart notes, the denial, and the payer's clinical criteria, then produces a letter a physician reviews and signs. That letter is what stands between a patient and treatment, and Priya's name goes on the recommendation.

#### **What They Try First (And Why It Falls Short)**

At first, she does what the room expects. She asks the vendor about content controls, confirms the security questionnaire is on file, and watches a demo in which the tool produces a clean, persuasive appeal in seconds. Everything checks out. The trouble is that those questions describe what the tool is forbidden to say, and[ Vol 47](https://www.mindovermoney.ai/why-does-my-ai-agent-do-the-wrong-thing/) established that the failures worth worrying about live somewhere else entirely. The demo carries the same weakness. It was built by the people selling it, using a case they chose.

#### **The Concept, Through the Scenario**

A safety evaluation is not a longer security review. The practice's privacy officer already owns where patient data travels and how it is stored. What nobody has been assigned is the question of what the tool is trying to do, and she works out that it comes down to four things she can ask in a meeting.

1. What is this optimizing for, and what would count as cheating? The vendor's dashboard counts appeals drafted per hour. What Priya wants is patients starting treatment. A letter missing the one clinical detail the payer needs is still a letter drafted, and it still moves the number.
2. Who tested this besides the people selling it? The tool runs on someone else's model, and the major labs publish safety documentation that names their outside testers. Anthropic's current report credits an attack benchmark built with Gray Swan, the UK AI Security Institute, and the US Center for AI Standards and Innovation. OpenAI's latest system card gives SecureBio, METR, and Apollo Research their own sections. Priya can read those herself. A vendor who cannot name anyone outside their own company has told her something.
3. What does it not know, and when did it stop learning? The same documentation carries a knowledge cutoff date. Anthropic lists May 2026 for its current model and calls that model's knowledge most reliable up to that point. Payer criteria do not hold still. A tool built on a spring cutoff can write a fluent appeal against criteria the payer replaced in July, and nothing in the letter will look wrong.
4. How would she know if it were wrong? The tool reports a completion rate, and that number restates the instruction it was given rather than proving any appeal was good. The check that settles it is the payer's approval rate, measured against the letters her team wrote before the tool arrived.

#### **What Changes**

On the next call, Priya runs all four. The vendor concedes that the dashboard counts drafts and says nothing about approvals, which is the moment the conversation turns useful. They name the model underneath, which is the only part of that answer she actually needs from them. On the cutoff, they explain that current payer criteria are fed to the model at drafting time, so nothing about the payer comes from training. That is a real answer and a good one. On the fourth, the room goes quiet. Nobody has measured approval rate against a human baseline, and doing it would need her team.

Her recommendation is not a yes or a no. It is a yes with a condition, a ninety-day run where every letter is scored on payer approval against the team's own prior numbers, with criteria supplied fresh each time. She wrote that condition herself. Nobody handed it to her.

#### **What This Reveals**

By design, this framework is small. Four questions fit on a notecard, survive a meeting, and work on any AI tool in any department, whatever it is called and whoever built it. None of them require Priya to be technical. None of them ask her to predict the future. They ask what the system is aimed at, who has tried to break it, what it does not know, and what independent evidence would settle the matter. Anyone who has sat quietly through a vendor demo wondering what to ask now has four things to say.

#### **How This Connects**

Vol 47 drew the line between what a model is allowed to say and what it is trying to do. The first question is that line, turned into something you can say out loud. Because the hunt for failures before release can never catch everything, as[ Part 2](https://www.mindovermoney.ai/how-ai-models-are-tested-before-release/) showed, the second question is worth asking and the third is worth asking twice.[ Vol 37](https://www.mindovermoney.ai/how-ai-is-trained-to-be-helpful/) explained why a system's own report of success is weak evidence, and[ Vol 42](https://www.mindovermoney.ai/how-much-autonomy-to-give-an-ai-agent/) set the dial for how far a tool runs before a person checks it. Next volume opens three weeks on AI governance and disclosure. In Part 1, the question is which AI rules already reach an ordinary professional's job. From there the series turns to disclosure, and to what it requires when a person reviews and edits what a tool produced. It closes by handing you your own employer's AI policy, or the discovery that there is not one. Whether a tool is safe to use and what you are expected to say about using it are different questions, and Priya's four answer only the first.

*Part 3 of 3 in the AI Safety series.*

## Your 10-Minute Win

****A step-by-step workflow you can use immediately**

## **The Document You Cannot Paste**

Sometime this month you had a document you wanted help with and did not use it. A contract with a client name on every page. A case summary. A performance conversation you needed to think through. So you typed a sanitized version from memory, got a generic answer, and quietly filed AI under better in theory than at your actual job.

The instinct was sound and the workaround is what costs you, because the details that make a document sensitive are rarely the details that make it useful to a model. Ten minutes builds you a substitution key that separates the two, and it covers most of what crosses your desk after that.

This pairs with this week's AI Education section, A Professional's Framework for Evaluating AI Safety at Work. That section is about deciding whether to trust a tool at all. This is what you do with real work on Monday, whichever way that decision lands. The model builds your checklist and tests your redacted version. You do the redacting, and the key never leaves your machine.

### **The Workflow**

**1\. Pick the document (1 minute)**

Choose one real thing you wanted help with this month and did not use. Open it and leave it open. Paste nothing yet.

**2\. Build the checklist without showing the document (2 minutes)**

Describe the kind of document it is rather than what is in it. A model already knows what identifies people in a contract or an incident report, and can tell you before it sees yours.

**Copy/Paste Prompt:** *"I have a \[DOCUMENT TYPE\] from \[INDUSTRY OR FUNCTION\]. Do not ask me to share it. List every category of detail in a document like this that could identify a person, an employer, an account, or a location, including details that identify someone only in combination with each other. For each category, tell me what to replace it with so the meaning survives."*

**3\. Redact by hand and keep the key (3 minutes)**

Work your document against that list. Replace each item with a consistent token, so the same person is Clinician 1 every time they appear. Write the pairings in a separate file. That file is your substitution key, and it never gets pasted anywhere.

**4\. Test whether it still works (3 minutes)**

Now paste the redacted version and ask the question you wanted to ask in the first place.

**Copy/Paste Prompt:** *"Here is a redacted \[DOCUMENT TYPE\]. Placeholders such as Vendor A stand in for real names. \[ASK YOUR ACTUAL QUESTION\]. If any placeholder is stopping you from answering well, tell me which one and what kind of information you would need, without asking me to reveal it."*

**5\. Judge the answer yourself (1 minute)**

Compare what came back against what you expected from the original. If the answer held, your key is done. If it thinned out, the model has just told you which placeholder cost you, and you can restore the shape of that detail without restoring the identity.

### **The Payoff**

You keep a substitution key that turns a document you could not use into one you can, and it works on every document after this one. You also keep a working read on how much a model actually needs. Most people either withhold more than they have to or hand over more than they should.

### **The AI Concept You Just Used**

Minimum viable context. A model needs the shape of a problem far more than it needs the identities inside it. Names, employers, and account numbers carry your risk. Structure and stakes carry your meaning, and the two overlap far less than people assume. Once you can see the line, real work stops being the category you keep away from these tools.

### **Transparency & Notes**

- This runs on the free tier of Claude, ChatGPT, or Gemini. No part of it requires a paid plan.
- Your substitution key is now the sensitive artifact. Keep it local, keep it out of the chat you are pasting into, and do not store it inside the tool itself. If you have not set what your tools retain between sessions,[ Vol 18 covered memory and temporary chat settings](https://www.mindovermoney.ai/chatgpt-memory-settings-ai-data-privacy-professionals/).
- Redaction is not anonymization. Enough unique detail in combination can still point at one person after every name is gone, which is why step 2 asks for combination risks by name. Some documents should not go into a consumer tool in any form.
- If your employer has an AI policy, it governs before this workflow does. In a regulated environment, check with whoever owns it first.