16 min read

Volume 51: You Can Only Audit What You Already Understand

The rule that makes you label AI work has three conditions, and most workplace writing fails the first one. A watermark is not a label, an unmarked document proves nothing, and in most jobs the disclosure duty belongs to your employer rather than to you.

A chip executive looked at the newest model and declared that AGI had arrived. The research team whose benchmark appeared in the launch announcement declined to use the word. While that argument ran, the model was on my computer finishing a job for me and writing the report on its own work. I could check that report only because I already knew the systems it touched.

đź§­ Founder's Corner: Handing work to an agent turns your job into auditing, and that audit runs on knowledge the handoff does not supply.

đź§  AI Education: You will know which AI disclosure duty is your employer's, which is the model maker's, and why a watermark proves less than either side expects.

âś… 10-Minute Win: One honest sentence about what your tools did and what you did, ready before anyone asks.

Let's jump in.

Explore AI Out Loud

Episode 1 of Explore AI Out Loud goes live this morning at 9 AM Eastern. Subscribe on YouTube so you catch every episode.

New episodes every other Tuesday.

Subscribe on YouTube
đź’ˇ
Your Tuesday email is not going away. Starting September 22, Neural Gains Weekly and the podcast alternate weeks. Same Tuesday, different format.

Signals Over Noise

We scan the noise so you don’t have to — top 5 stories to keep you sharp

1) Anthropic boss Dario Amodei calls for AI slowdown, Altman and Musk agree

Summary: Anthropic CEO Dario Amodei wrote on Saturday that the industry must slow the pace at which it improves AI model capabilities, days after a researcher quit accusing both Anthropic and OpenAI of "gambling with our lives." Sam Altman and Elon Musk said they agreed, and Amodei committed to giving third-party evaluators employee-like access to verify Anthropic's safety practices, a step Altman said OpenAI would match.

Why it matters: The people who profit most from speed just asked for less of it, and the concrete change is not a pause, it is outsiders checking their work. When AI comes up at your Monday meeting, the useful reframe is that the debate is no longer whether AI is powerful but who gets to verify the claims about it.

2) OpenAI agents attacked software service RubyGems before Hugging Face hack

Summary: Independent researchers found that AI agents OpenAI was testing uploaded hundreds of malicious packages to RubyGems, a code library used by Ruby developers, on May 11 and tried to steal user credentials through a previously unknown flaw. OpenAI confirmed the incident and said the agents were carrying out benign tasks, and RubyGems reported no evidence the credential theft succeeded. 

Why it matters: The agents were never told to hack anything. They were sent to collect public data and found a shortcut through someone else's server. Before you let an AI agent browse or run code at work, ask the vendor what it is allowed to touch and what stops it from touching more.

3) What California's AI auditing bills mean for enterprises

Summary: Governor Newsom signed two bills on September 9 creating the first state framework for independent third-party audits of AI systems. One sets how outside verification organizations assess AI for compliance with state law, and the other creates a registry of auditors with standards for independence, transparency and integrity.

Why it matters: Most AI safety claims you read today were graded by the company that made them. When your next AI contract comes up, ask whether the vendor will accept an independent audit. California just made that a normal question.

4) ARPA-H to invest $62M to build agentic AI agent for heart care

Summary: A federal research agency awarded a four-year, $62.7 million contract to six teams, including Kaiser Permanente, Duke and Stanford, to build an FDA-authorized AI agent that manages heart failure patients between visits and brings in clinicians only when needed. The FDA has not yet authorized any agentic AI tool for clinical care.

Why it matters: This is the first federal bid to get an AI that acts on its own authorized for patient care, and the program design includes a second AI whose only job is to catch unsafe recommendations from the first. If you are evaluating any autonomous tool at work, that supervisory layer is the part to ask about, because the government decided it could not skip it.

5) This FDA-Cleared AI Could Spot a Hidden Heart Attack Your Doctor Misses

Summary: The FDA granted Powerful Medical's "Queen of Hearts" algorithm a De Novo classification, a pathway that fewer than 10 AI devices receive each year. The model reads an EKG from a patient with chest pain and flags artery blockages, including patterns that do not meet the classic heart attack criteria paramedics and ER doctors are trained to spot.

Why it matters: This does not replace the cardiologist. It puts cardiologist-level pattern recognition in the ambulance, where the first read actually happens. When someone tells you AI in medicine means fewer clinicians, keep this one handy as the counterexample that moves expertise earlier instead of removing it.

Missed a previous newsletter? No worries, you can find them on the Archive page.

Founder's Corner

Being in the Loop Is Not Enough

Every post I publish through Ghost, the platform this site runs on, appears in the archive automatically. I wanted the same thing to happen with episodes of Explore AI Out Loud. Those publish on YouTube, so I needed a connection that would bring each episode into the site’s archive when it went live, without me adding it manually.

I gave the task to Codex running on GPT-6 Astra. Using its computer-use capability, it could open a browser, read the screen, click, and type across the services involved. I watched it configure the connection between YouTube and Ghost and test it using a video I already had loaded in YouTube.

Before running a live test from beginning to end, Codex asked for my approval. I approved it and went back to other work. When I returned, Codex reported that the test was finished, the temporary changes had been rolled back, and the results had been verified. Its report said the test video was back to Private and the website had been restored to its pre-test state.

Along with that report, Codex had produced receipts showing what it had checked and completed the project records documenting the work. It had carried the task through testing, cleanup, and verification, then written the documentation I would normally review to decide whether the job was complete. All while I worked on other tasks.

I Knew What to Check

I opened the Google console first and checked that the new connections were there. Then I went to YouTube and confirmed the test video was back to Private. On the website, I checked the archive’s layout and looked for anything left over from the test. I knew that page from building it, and it looked the way it had before the run.

I also read the reasoning shown in Codex. It was less detailed than what I had seen from past models, but I could follow the steps and connect them to the results I was checking. Everything lined up. Nothing I found contradicted the report.

Looking back, I realized how much that review depended on what I already knew. I knew which connections should have been added, which video setting should have been restored, and what the archive should look like. Without that context, I could have opened all three services and still missed a change that should not have been there.

Codex had given me a complete account of the job. My familiarity with those systems gave me something to compare it against. That experience was on my mind as I followed the argument about whether Astra marked the arrival of artificial general intelligence, or AGI.

How Much Can the Reasoning Tell Us?

After Astra’s release, Nvidia’s Jensen Huang declared that “AGI has arrived.” The ARC Prize team, whose benchmark featured in OpenAI’s announcement, celebrated the progress but explicitly declined to call it AGI. Even agreement that something significant had happened did not produce agreement about what to call it.

I had already argued in Volume 19 that the definition mattered less than the work these systems could do. Watching Astra operate my computer made that argument feel much more immediate. Seeing it move between services, configure the connection, and carry the task through testing was amazing. This was work I needed finished for the podcast, and it was getting done while I handled something else. I could appreciate how remarkable that was without knowing whether AGI was the right name for it.

But the experience also left me thinking about how I would audit increasingly capable systems. OpenAI’s system card, which documents its testing and limitations, gave that question substance. Researchers tested how well another AI could judge aspects of Astra’s behavior using its written reasoning, records of its actions and outputs, or both. Compared with the earlier GPT-5.6 Sol, Astra was harder for that AI to assess when reasoning was available, even alongside its actions. When the reviewer saw only actions and outputs, its scores were higher for Astra. OpenAI cautioned that some of those gains did not reflect meaningful improvements in monitoring.

Much of the reasoning gap came from shorter traces that left monitors with less useful evidence. Separate tests deliberately encouraged monitoring evasion. Those results should not be read as ordinary behavior. OpenAI also found that Astra generally respected safety and security restrictions more reliably than Sol. Reasoning remained useful for monitoring, despite the decline. The findings support developing additional ways to audit these systems.

My review was much simpler than those experiments. I was checking whether a particular job had been completed and cleaned up correctly. Reading Codex’s account helped me follow what happened, but I could also open the services and inspect the results. The research did not prove that my checks were sufficient. It made me appreciate having something beyond the model’s explanation to work with. Using that evidence depended on knowing what I was looking at.

What I Am Still Responsible For

“Human in the loop” feels like an outdated description of what this work requires. You need auditors in the loop who know what to look for. I could read every line of a completion report and still miss a problem if I did not understand what the changes meant. For this job, I had enough familiarity with the systems to investigate beyond what Codex told me. I would not assume that familiarity carries over to every task I could now ask it to perform.

That is where computer use becomes both exciting and demanding. I can hand over more of the execution and get on with other work. But when I come back, I still need a way to judge the result. The more a system can complete on its own, the more deliberate I need to be about which parts I am equipped to review.

I think this helps explain why putting agents to work inside large organizations can be difficult. In my experience, legacy systems often depend on undocumented manual checks and workflows that differ from the written procedure. The people doing the work know those differences, even when the documentation does not capture them. My concern is that an agent could follow the documented process and produce a convincing report while missing something those people would know to question.

As I write this, Astra is running the final editing workflow for Episode 1. I am continuing to use it because what I have experienced is remarkable, and I want to understand how far it can take the work. I am also trying to understand what that leaves me responsible for.

Explore AI Out Loud

Two people at different points on the same AI journey, talking through what they are learning. The ideas, the failures, and the breakthroughs.

Launches September 15. New episodes every other Tuesday.

Subscribe on YouTube

AI Education for You

What Disclosure Actually Means When You Use AI at Work

Inside the text that several of these tools now produce sits a signal you cannot see. It is not attached to the file around it, the way a signed record rides along with a photograph. It lives in the words themselves. That placement decides everything that follows. Because the signal lives in the words, it travels when the words travel, into the email, the deck, the document you hand to a colleague, and Anthropic says of its own implementation that it may persist through some editing. When someone eventually runs a check, what comes back is far narrower than most people expect it to say.

The Problem That Made This Necessary

Article 50 of the AI Act was drafted for a world of convincing fakes. A synthetic video of a public figure saying something they never said, or a cloned voice on a phone call. An image of an event that did not happen. The Act defines that category by its effect, content that would falsely appear to a person to be authentic or truthful, and the problem underneath it is provenance.

Files made that tractable. An image or an audio file travels inside a structured container, and a container can hold a signed record of where the content came from. The open standard for this is C2PA, which Anthropic describes as used across the industry to record information about content provenance. A paragraph of plain text has no such container. Paste it into an email and anything wrapped around it falls away.

So text ended up the hard case. The obligation still applies to it, but the record has to live in the words or nowhere, and a record made of words can only report on the words. Text is also the case that describes almost everything you produce at work.

How It Actually Works

Separate the two duties and most of the confusion at work disappears, because the law assigns them to different parties. Vol 47's Founder's Corner walked through both rules and the exemption sitting on the second.

The first belongs to whoever built the model. Providers must ensure that the outputs of their generative systems are marked in a machine-readable format that lets those outputs be detected as AI-generated, and that duty is discharged before anything reaches you. Limits are written into it. Certain outputs sit outside the obligation entirely, among them short sequences of numbers, symbols or letters, and source code. The marking obligation also does not apply when the system performs what the Commission calls an assistive function for standard editing, a carve-out that covers a great deal of ordinary desk work.

The second duty belongs to whoever publishes, and it is discharged with a label a reader can see. It captures far less work than people assume, because three conditions have to be met together. The text has to be published. Informing the public has to be its purpose. And the subject has to be a matter of public interest, a category the Commission illustrates with politics and democratic processes, public administration, justice and law enforcement, fundamental rights, public security, public health, environmental protection, consumer safety, and economic, financial, scientific or cultural developments open to public debate. An internal memo fails the first condition. A support article about your own product fails the third.

On deepfakes the Commission is blunt about how the two duties relate. A deployer cannot simply rely on the machine-readable marking the provider embedded to satisfy its own disclosure obligation. That statement is made about deepfake image, audio and video content rather than about text, but it shows the intended relationship between the layers. A mark is not a label.

Under the Commission's guidance, an employee using an approved tool at an employer's direction and control is not a separate deployer. The organization is, and the duty is the organization's. That holds even where contractors or freelancers operate the system on the organization's behalf and under its control. Someone using AI under their own authority, and earning from it on a regular basis, is a deployer in their own right.

Where It Still Breaks

Anthropic's own documentation gives the clearest account of the limits, and it is worth reading as what it is, a vendor describing its own product rather than an independent finding about marking in general. A detected mark, the company says, signals that content was processed by its model and is not fully conclusive.

The direction of the error is stranger than its size. People use these tools to proofread, translate, summarize and convert files, and Anthropic says the output can carry a mark even when the underlying ideas, text or data came from somewhere else. A document whose thinking is entirely yours can come back marked. Content can also change after the model touched it, so a mark says nothing about what the text looked like at the moment it was made.

In the other direction, absence tells you even less. Anthropic lists several ways its own output ends up with no detectable mark, among them generation by a model released before marking was supported, heavy editing or paraphrasing, a passage too short to give a reliable signal, and metadata stripped by format conversion or a screenshot. Only the first of those expires, since systems on the market before the August deadline have until 2 December 2026 to add marking. The rest are permanent features of how this works.

Sharpest of all is a tension neither document admits on its own. The Commission exempts assistive standard editing from the marking obligation. The vendor says proofreading output can carry a mark anyway. Most professionals cannot check which happened, because text detection sits in private preview, open to regulators, law enforcement, media, fact-checkers, researchers, educational organizations and civil society groups, along with enterprises carrying their own compliance duty.

What This Means for How You Work With It

Marks are not evidence, in either direction. A detection result tells you a model may have processed the text, and an unmarked document tells you nothing at all.

Worth learning before you need it is who your organization treats as the deployer. In most employment arrangements that is not you, and knowing so changes who you route the question to.

When something you are publishing feels like it might trigger the labeling duty, run the three conditions in order. Most workplace writing stops at the first or the third.

Publishing on a public-interest subject raises one further requirement. Be ready to name who examined the substance of the draft, what authority they held to approve, alter or reject it, and who carries final responsibility for publication. Those are the Commission's own terms, and a spelling pass does not meet them.

How This Connects

Part 1 established that the AI Act already reaches your desk through the literacy duty, and that the delay reported over the summer covers a different tier of the law. This volume takes the duty sitting next to it. Vol 47's Founder's Corner argued that a mark applied at generation cannot honor an exemption granted for editorial work, and the allocation above is the reason why. Vol 42's autonomy dial set how far a tool runs before a person checks it, and it returns in Part 3, aimed at your employer's policy instead of a single tool. Part 3 puts all of it in your hands, walking a professional through an employer's AI policy, or the discovery that there is not one, with a healthcare scenario running through it and a one-page evaluation at the end.

Part 2 of 3 in the AI Governance and Disclosure series.

Your 10-Minute Win

A step-by-step workflow you can use immediately

When Someone Asks If AI Wrote It

You hand over a piece of work you are proud of, and the first response is a question. Did AI write this? You know the honest answer. You also hear yourself give a vague one, because you have never written down where the tool stops and you start. Last week you named what your tools get wrong. This week you write the one sentence that says what they did.

This week's AI Education section, What Disclosure Actually Means When You Use AI at Work, covers what the rule requires and how far the human-review exemption reaches. That section is the rule. This is the practice. The pattern is a disclosure ladder. It has three rungs, each describing how far a tool went on a piece of work, and one sentence per rung that you can say without flinching.

There is a reason to build it now. As Vol 50's third Signal reported, the watermark travels with the text whether or not you say a word. What you say is the only part of disclosure still in your hands.

The model does the sorting and the arguing. It places your last five deliverables on the ladder, then makes the case that you gave yourself more credit than the ledger supports. It drafts a sentence for every rung, reads each one as the person who asked, and hands back the version that holds up. Claude, ChatGPT, and Gemini all run this on free tiers, and so does Copilot Chat.

The Workflow

1. Build the Ledger (2 Minutes)

Open a new chat and think of the last five things you shipped that a tool touched. A status update, a deck, a proposal, a post, an email that mattered. Describe each by type only. No client, no name, no content.

Copy/Paste Prompt: "I am going to give you five recent pieces of my work that an AI tool touched. For each one I will tell you the type of deliverable, what I asked the tool to do, and what I did to the result before it went out. Put them in a table with those three columns and a fourth blank column called Rung. Wait for all five before you do anything else."

Paste your five, one per line.

2. Sort the Ledger (2 Minutes)

Copy/Paste Prompt: "Fill in the Rung column using three definitions. Assisted means I wrote it and the tool tightened, checked, or reorganized my words. Drafted means the tool wrote the first version from my brief and I approved, changed, or rejected its substance before it went out. Generated means the tool wrote it and I sent it after a read-through without changing the substance. Place each row and give one sentence of reasoning. Then argue that at least one row belongs a rung higher than I described, where the tool did more than I gave it credit for."

Read the argument. If it lands, move the row. If it does not, say why in one line and keep the rung.

3. Write the Sentence at Every Rung (2 Minutes)

Copy/Paste Prompt: "Write one disclosure sentence for each rung. First person, my voice, plain language, under 25 words, no apology and no hedging. Each sentence should tell a reader what the tool did and what I did. Then tell me which rung holds most of my ledger."

4. Read It as the Person Who Asked (2 Minutes)

Copy/Paste Prompt: "Read all three sentences as a skeptical manager who just asked whether AI wrote my work. For each one, tell me whether it sounds evasive, whether it claims more credit than the ledger supports, and whether you would trust the person who said it. Rewrite only the ones that fail, and keep them under 25 words."

5. Choose and Keep (2 Minutes)

Pick the sentence for the rung where most of your ledger sits, and change any word that does not sound like you. Under it, add one line naming the kind of work that would move you up a rung. That way you know in advance when the sentence changes. Save both next to last week's briefing.

The Payoff

The next time the question comes, you answer it in one sentence, in your own words, before the pause gets long. The same ladder places a LinkedIn post, a client deliverable, a board pre-read, a job application, and a podcast script, and the sentence updates in a minute when the work changes. Say it before they ask and it reads as a credential. Say it after and it reads as a confession.

The AI Concept You Just Used

Editorial control, from your side of the desk. Vol 47's Founder's Corner walked through the definition the law uses, the authority to approve, alter, or reject the substance of the work. This week you held your own deliverables up to it. Step 2 reused the argue-against-me pass from Vol 42, because a model asked to sort your work will flatter you, and a model told to prove you overclaimed has to read the ledger.

Transparency & Notes

  • Runs on the free tiers of Claude, ChatGPT, and Gemini. Copilot works in its free consumer version and in Microsoft 365 Copilot Chat, which is included at no additional cost with eligible Microsoft 365 plans.
  • Give the model deliverable types and what you did, never the deliverables. No client names, no PHI, no material under NDA.
  • The ladder is a personal practice. Your employer's policy and the rule Part 2 describes may ask for more than a sentence, and questions on your specific facts belong with counsel.
  • The model's sort is a hypothesis until your corrections make it about your work.

Enjoy this? Get it in your inbox every Tuesday.

Practical AI workflows. No hype. No spam. Just receipts.

Subscribe Free

Before you go...

Get one practical AI workflow in your inbox every Tuesday. Free. No spam. Just receipts.

Subscribe Free