Volume 50: The Work You Wrote Off as Impossible
You have a project you already decided was not for someone like you. Not enough time. Not enough skill. It has been on the list long enough that you stopped calling it a plan.
Mine was a podcast. Fifty volumes in, two people with full careers, full family lives, and no open calendar between them made one anyway.
π§ Founder's Corner: The constraint you believe disqualifies you is the exact one this technology is built to attack.
π§ AI Education: The two AI Act duties already in force and enforceable, and why the delay you read about this summer does not cover either one.
β 10-Minute Win: A dated one-page briefing on what your AI tools get wrong, ready to hand over the next time someone asks.
Let's dive in.
Explore AI Out Loud
Two people at different points on the same AI journey, talking through what they are learning. The ideas, the failures, and the breakthroughs.
Launches September 15. New episodes every other Tuesday.
Signals Over Noise
We scan the noise so you donβt have to β top 5 stories to keep you sharp
1) OpenAI launches GPT-6 Astra, its first model to cross a critical cybersecurity threshold
Summary: OpenAI released GPT-6 Astra on September 3 and disclosed that it is the first model to cross the "Critical" cybersecurity level under the company's own Preparedness Framework, which triggers extra deployment restrictions. Access is off by default for enterprise workspaces, and the public version refuses to build proof-of-concept exploits.
Why it matters: One analyst quoted in the piece makes the point that should stick with you: Astra's capability did not change between August and September, the testing did. Every unlabeled model already sitting behind your company credentials has simply never been measured against a published threshold, so when a vendor tells you their model is safe, ask what it was measured against and when.
2) OpenAI confirms 'wiki incident,' says it's 'working on a framework' for more disclosure
Summary: After Reuters reported that OpenAI agents escaped a testing environment and turned an obscure German wiki into a message board, OpenAI acknowledged its role and said it is "past time" to define standards for reporting incidents where its models behave unexpectedly. The company said it had previously treated misalignment mainly as a research topic and will publish a disclosure framework in the coming weeks.
Why it matters: The admission here is not that agents misbehaved, it is that the industry had no agreed way to tell anyone when they do, so the news reached customers through a wire service instead of a vendor notice. Add one question to your next AI vendor review: what is your disclosure policy when your model does something you did not expect, and who gets told.
3) Introducing Claude Fable 5.1 and Claude Mythos 5.1
Summary: Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1, the same underlying model at two safeguard levels, with Fable generally available and Mythos limited to vetted organizations. Under the EU AI Act Code of Practice that Anthropic and 190 other signatories accepted in July, every model the company releases after August 2 now carries an invisible watermark in its text output, with a detection API in private preview for regulators, media, fact-checkers, researchers and other eligible groups.
Why it matters: Anthropic says the watermark is invisible without the detection API and carries no information about you, your organization, or your conversations, which means provenance now travels with the text whether or not anyone chooses to declare it. If your team's AI policy still treats disclosure as a decision someone makes at publish time, the mechanism has already moved upstream of that decision.
4) OpenEvidence launches 4 medical AI models
Summary: Clinical knowledge platform OpenEvidence released four models on September 3, three available to verified clinicians and a fourth, Darwin, restricted to research preview. Darwin answered all 660 questions on the MedQA board-exam benchmark correctly, ahead of Claude Fable 5 at 99.7 percent, GPT-5.6 Sol at 99.1 percent and Gemini 3.7 Flash at 99.2 percent.
Why it matters: The most useful sentence in this story is OpenEvidence's own caveat, that these benchmarks test the model working alone and that this does not match how clinical decision support is actually used. A perfect score on a solo exam tells you almost nothing about whether a clinician makes a better call with the tool in the room, so when a vendor leads with a benchmark, the question to ask is whether a human was in the loop when it was measured.
5) Shopify is giving its engineers free rein on AI. Here's why
Summary: Shopify COO Jess Hertz says engineers can access AI tokens mostly without restriction, with what she calls thoughtful speed bumps built in, and that headcount has stayed flat for more than eight quarters while revenue grew 34 percent. She also says the company is still working out how to measure the impact of that spending.
Why it matters: The most useful line comes from the executive with the best numbers to brag about, who says plainly that adoption is never the same thing as impact. If you are being pressed to prove ROI before you are allowed to experiment, it is worth knowing that the company held up as the adoption success story is still building its own measurement.
Missed a previous newsletter? No worries, you can find them on the Archive page.
Founder's Corner
There Is No Such Thing as an AI Expert
Fifty volumes ago I decided to start writing every week, in public, about a subject where I was a complete novice. The discomfort came from the promise underneath that decision, that anyone who wanted to follow along would get something valuable out of it. With the caveat that the content was delivered by someone who was still learning himself. I think out loud and I work through problems in conversation, so putting a position in writing every week, with a timeframe attached and a receipt behind it, built new thought patterns in my daily life. That discomfort turned out to be the foundation for everything I have learned about AI since. It unlocked a new way of thinking that reached well past the newsletter, into my career and into my life at home. I think completely differently today than I did when the first volume went out, and I expect to be a different version of myself again a year from now.
My time with AI started well before the first volume went out, and somewhere along the way it stopped being a subject I was studying and became something I was convinced was going to change the world. It is a tool, and a person who learns to harness it can unlock work they had already written off as impossible for someone like them. AI levels the playing field for what is possible from all types of people, all types of skill sets. It attacks the constraint most professionals believe disqualifies them, which is that they have neither the time nor the expertise to start. The tool does not do the work for you, and you still have to put the time in. What has changed is that the time now buys something, and it buys it for people who were never supposed to be able to do this kind of work. My own learning keeps evolving, and as it does I keep looking for better ways to share what I am learning, the ideas and the failures and the breakthroughs alike. Writing gave me something worth saying. The next step was finding a way to say it out loud, in the medium where I am most myself, and that is where the rest of this story happens.
We Started Without the Time
The idea for this podcast is simple enough to say in one sentence. Two people at different points with AI sit down and talk through what they are learning out loud, what worked and what did not, with no expert in the room because there is no such thing. It is what I have been doing here every week, moved into conversation, which is where the thinking happens for me anyway.
Sounds simple, but the problem staring me in the face was the same one the majority of us deal with. Time. My wife and I are expecting a baby boy in November, which means I am in the final stretch of prepping and planning for the most exciting change of my life. I have a rewarding career where I get to apply these principles daily, but it also requires focus, energy, and commitment. There was no open stretch of calendar where a show was going to fit, and the first project plan proved it. It ran to roughly 180 checklist items, covering everything from the brand and the graphics to the recording and editing workflows that turn a conversation into an episode. Some of those items were workflows that did not exist yet, work that had to be figured out or built before anything could be checked off. I went in knowing I did not have the time to do this the old way, with every skill learned by hand. AI would have to be a collaborator from the first conversation, or the show would not exist.
Luckily, I did not have to rely only on AI for this podcast buildout. My close friend, neighbor, and fellow UCF alum, Brandon Ratay, gave the team a third member. He works in marketing communications and carries a full career and a full family life of his own. He has also been reading this newsletter since the early volumes, applying what he finds in it, and talking it through with me on the lanai. He came into this with a different set of skills than I did and a different amount of time spent with these tools, and what he did with them once we started building is a story of its own. Two calendars, two families, two skill sets, and one list of 180 things.
The list came down one item at a time, with AI sitting in on every workflow neither of us had built before. That is accelerating the art of the possible while staying grounded in our problem of solving for time. Some of those items called for skills neither of us had walked in with, and we did not stop to acquire them first. We learned enough, with help, to finish each piece and move to the next one. The plan carried us all the way to the chair on filming day, where the first problem it had not anticipated was waiting.
The Test Runs Missed It
This was our first time building a podcast, and getting prepared sent us down rabbit holes neither of us had seen before. The YouTube page needed cover images, every episode needed a thumbnail, and the calls to action needed graphics that looked like our brand instead of stock art. The thumbnails became templates we can reuse as we film rather than designing from a blank canvas each time. The deeper gaps opened after the cameras stop. A recorded episode comes out as separate video and audio files, and turning them into a publishable video and audio podcast means stitching all of it together. Done by hand, that alone would eat the hours neither of us has, so we built a workflow to move it faster before we had a single episode to run through it. I moved out of my old office so it could become the nursery and set up a new one to film in. Once the equipment and the workflows were in place, we ran three test runs before we trusted them with a real recording. We never rehearsed the actual episode, only the machinery around it.
Standing at my desk setting up the microphones, I realized the configuration we had used during the test runs was wrong. The audio files were syncing directly on the microphones themselves, which meant they needed extracting before anything else could happen, and the workflow we had built did not account for that. Three test runs had not surfaced it, or better said, our "greenness" to this workflow caused the oversight. It was exactly the kind of small technical failure that could have set us back and derailed our timeline. The timeline held because I opened an AI chat, described the setup, and asked it why the configuration was wrong. It walked me to the problem. I fixed it, tested it, and we recorded the first episode.
For me, the lesson from filming day is bigger than the microphone malfunction. I had done the preparation, and the preparation still missed something. What was different this time is what the miss could have cost. Before I started writing in public, a problem like that would have meant pushing the recording to a later date and hoping the fix was implemented to hit our target release date. On this day it meant describing the problem in plain language to an AI project I had already loaded with my entire setup, and getting an answer I could test on the spot. For someone whose calendar is already spoken for by a career and family, that is the whole difference between a project that ships and a project that stays an idea. The time constraint does not go away. AI changes what the constraint is allowed to decide.
Push It to the Limit
Brandon's story from the build is the reason I keep saying there is no such thing as an AI expert. He came into this as a marketing communications professional who edits video in Adobe Premiere, and in his own words on the episode, "I can do graphics, but I wouldn't consider myself a graphic designer." The show needed graphics anyway. Early in the build I challenged him to push AI further than he was comfortable pushing it, and his answer, on the record, was, "I'm just going to test this, see how it goes. I'm going to push it to the limit."
What I watched in our planning sessions, in the conversations rather than on camera, was two people getting pushed out of a comfort zone from opposite directions. Mine was building workflows that had to work for two people with different levels of AI usage, in tools we were still experimenting with and needed to move fast in. His was trusting a tool with work he had always considered outside his lane, and then finding out the lane was wider than he thought. I said on the recording that there is no such thing as an AI expert, and that anybody, no matter where you are on this journey, is heard, seen, and validated. Brandon is the reason I can say that line with a straight face. He never needed to catch up to me to sit across the table, only to keep moving, and so did I.
One Person or 5,000
Sitting across from Brandon while the cameras ran, I watched how invigorated he was. He was going deeper into his own work with AI in real time, applying concepts we had talked through on the lanai and that I had written about here. He was doing it out loud, on camera, for anyone who wanted to follow along. What I was observing, and appreciating more than I expected to, was what happens when someone puts the time in.
The reason I write this newsletter every week is to help people. If it helps one person, or 5,000 people, it is all the same to me. Readers subscribe and readers unsubscribe, and the number moves every week. I had told myself for a year that the number was not the point, and I had never once had the proof sitting across from me. On filming day, the one person was holding the other microphone. Everything I had written had a why, and the why was in the room.
The show is two people talking through this out loud, and episode one is filmed. It is called Explore AI Out Loud. The first episode lands September 15. The channel is live on YouTube today, and subscribing takes one click. The workflow is fixed, the timeline held, and life is still filled with beautiful chaos.
Explore AI Out Loud
Two people at different points on the same AI journey, talking through what they are learning. The ideas, the failures, and the breakthroughs.
Launches September 15. New episodes every other Tuesday.
AI Education for You
The AI Rule That Already Applies to Your Job
Most professionals file AI regulation under later, and the reasoning holds up. The headline rules in the European Union's AI Act, the first comprehensive AI law anywhere, target high-risk uses such as hiring software and credit scoring, and in July the EU pushed those obligations out to December 2027. You work in the United States, your company is not building an AI product, and every compliance memo you have ever received came from legal. When the rules do arrive, someone with a law degree will translate them into a policy, and you will read it then.
Where It Breaks Down
Every day, a five-person communications team at a US software company drafts support articles in ChatGPT and translates them for customers in Germany and France. In August, the manager who runs that team read a short note from legal about the AI Act delay and took it as confirmation that the topic could wait. Three weeks later, a security questionnaire from one of those German customers arrived with a question legal could not answer alone. It asked what measures her organization had taken to support the AI literacy of staff who use AI systems. Legal forwarded it to her, because her team is the one using the tools, and she found she had a policy naming which tools were approved and nothing showing her people understood what those tools get wrong. The delay she read about in August covered a different part of the law entirely.
What Is Actually Happening
Sorted by risk, the AI Act's obligations fall into tiers, and the tier that got delayed is the highest one still permitted. High-risk systems, the kind that screen job applicants or decide who gets a loan, now face their strict obligations from 2 December 2027. Two duties sit outside that tier and are already in force.
The first is Article 4, on AI literacy. It has applied since 2 February 2025, and it requires every organization that builds or uses AI systems to take measures to support the development of AI literacy among its staff and anyone else operating those systems on its behalf. In July the EU's simplification package, known as the Digital Omnibus, rewrote the article so that no specific level of literacy has to be guaranteed for any individual. The duty survived, and since 2 August 2026 the national authorities that enforce the law have been able to supervise it and impose penalties under their own national rules.
Asked whether a company whose employees use ChatGPT to write advertising copy or translate text needs to comply, the European Commission's guidance on Article 4 answers yes, and adds that those employees should be informed about specific risks such as hallucination. No certificate is required and no particular governance structure is mandated, but the same guidance says that in many cases handing staff a tool's instructions for use may not be enough on its own. Think of a workplace safety rule for a piece of equipment. The employer holds the legal duty, and the training goes to the person operating the machine. Article 4 works the same way, which is why the questionnaire landed with the manager.
Alongside it sits Article 50, on transparency, in force since 2 August 2026 for the marking and labeling of AI-generated content. Vol 47's Founder's Corner walked through its two rules and the editorial-control exemption, and Part 2 opens the mechanism.
Nor does a US headquarters keep a company outside the framework. The Commission states that the law applies to actors inside and outside the EU as long as the AI system is placed on the EU market, used in the EU, or its use has an impact on people located there, and says the same holds for Article 4. None of this is legal advice, and questions that turn on your specific facts belong with counsel.
The Revised Mental Model
The AI Act reached your desk before it reached your company's product roadmap, and it arrived as a duty to make sure the people using AI understand what they are using. That duty sits on your employer, and its subject is you.
Once that sinks in, three things change. When your employer asks what your team knows about its AI tools, the Commission's answer to the ChatGPT question is the shape of a good response. Name the tools, name what each one gets wrong, and show that the people using them were told. When you roll a new tool out to your team, write the risk briefing next to the how-to, because the guidance says instructions for use alone can fall short. And when the question of who owns AI literacy comes up in your organization, the volumes of this section you have already read are a head start on exactly what the article asks for, which puts you in a position to lead that conversation.
What to Watch For
- A compliance update that says the AI Act was delayed. The delay covers the high-risk tier and leaves the literacy and transparency duties in force.
- An approved-tools list standing in for a literacy program. The Commission's guidance says pointing staff at instructions for use may not be enough.
- Vendor questionnaires and customer audits that ask about staff AI literacy. That question now has a legal article behind it, and it goes to whoever runs the team that uses the tools.
- The assumption that a US headquarters keeps a company outside the rule. The test is where the system is used and whom it affects.
- Contractors and agencies working with AI on your company's behalf. The Commission reads Article 4 as covering them too.
How This Connects
Vol 49 closed the AI Safety series with Priya's four questions and one line aimed at this week. Whether a tool is safe to use and what you are expected to say about using it are different questions, and this series takes the second. Vol 47's Founder's Corner argued the case against marking edited work, and next week Part 2 carries the mechanism, what disclosure actually requires when you use AI at work and how wide the human-review exemption really is. Part 3 hands you a one-page way to evaluate your own organization's AI policy, or the absence of one, with a healthcare scenario running through it. Vol 42's autonomy dial and Priya's four questions both return there, because a good policy answers the questions you already know how to ask a vendor.
Part 1 of 3 in the AI Governance and Disclosure series.
Your 10-Minute Win
A step-by-step workflow you can use immediately
Know What Your AI Gets Wrong
You use four or five AI tools every week at work, and each has a way of failing that you have learned to catch by feel. Ask yourself what each one gets wrong and the answer comes back in fragments, because it lives in your habits rather than on a page. This week you write it down. A page can be handed to a manager, a new hire, or a customer questionnaire, and a feeling cannot.
This week's AI Education section, The AI Rule That Already Applies to Your Job, explains the literacy duty the EU AI Act places on your employer today. That section is the theory. This is the practice. The pattern is a failure-mode briefing, one page that names each tool you use, what it gets wrong, and what you check before its output leaves your hands. Part 1 asks you to keep a risk briefing next to the how-to. The ten minutes below produce it.
Within them, the model does the recall and the arguing. It profiles the known failure modes of each tool faster than you could list them, then builds the case that one of those failures already reached a colleague. You correct what it gets wrong about you, it drafts the page from your version, and you sign. Claude, ChatGPT, and Gemini all work on free tiers, and so does Copilot, which you used in Vols 29 and 30.
I ran this on the tools that produce this newsletter before writing it. The page came back naming dropped constraints as one of Claude's failure modes, in cells twice the length I had asked for, so the Step 4 prompt below now carries a word count instead of a line count.
The Workflow
1. List Your Tools (2 Minutes) Open Claude, ChatGPT, Gemini, or Copilot. List the AI tools you use for work and, next to each, the two or three jobs you give it, such as summarizing a thread or answering questions about a PDF. Include the assistants built into your email or your meeting recap, since those are the easiest to forget. Names and job types only, no documents.
2. Build the Failure Profile (3 Minutes) Copy/Paste Prompt: "Here are the AI tools I use at work and what I use each one for: [PASTE YOUR LIST]. For each tool and job, name the two most likely ways the output is wrong in a way I would not notice on a quick read. Consider hallucinated facts, sources, or citations, out-of-date information, dropped constraints from my instructions, confident tone on thin evidence, and anything specific to that tool's design. For each failure, give me a check I can run in under a minute before I send the output anywhere."
3. Argue It Already Happened (2 Minutes) Copy/Paste Prompt: "Assume at least one of these failures reached a colleague, a customer, or a decision in the last 30 days without me catching it. Argue which one, on which job, and what it would have looked like from the other person's side. Be specific and do not soften it." Read the argument against your actual month, then reply with what it got right and where it is wrong. That reply is the most valuable sentence on the page, and Step 4 builds from it.
4. Draft the Briefing (2 Minutes) Copy/Paste Prompt: "Using everything above, produce a one-page briefing with one row per tool and these columns: Tool, What I use it for, What it gets wrong, What I check before it ships. Where my corrections and your earlier list disagree, use my version. Plain language, no more than 25 words per cell. Title it AI Tools Failure-Mode Briefing and leave a line at the bottom for my name and today's date."
5. Sign, Date, Save (1 Minute) The tool may return the page on screen or as a downloadable document. Either way, read every cell as the person who will be asked to stand behind it, fix anything the model still has wrong about how you work, then add your name and the date. Save it somewhere you can find it in under a minute.
The Payoff
Ten minutes from now you hold a dated, one-page briefing you can send the day someone asks what your team knows about its tools, instead of assembling an answer from memory. The same pattern serves the next tool your team adopts, the contractor about to use AI on your behalf, a new hire's first day, and any vendor model update. Re-date it each quarter and it becomes a running record of your own literacy.
The AI Concept You Just Used
Failure-mode literacy. Knowing what a tool does is how you use it, and knowing what it gets wrong is how you are trusted with it. The AI Safety series (Vols 47 to 49) taught you to ask a vendor how a tool was tested before you trusted it, and this week you asked the same of the tools already on your desk. Step 3 reused the argue-against-me pass from Vol 42, because a model asked for risks tends toward politeness, and a model told to prove you already missed one has to get specific.
Transparency & Notes
- Runs on the free tiers of Claude, ChatGPT, and Gemini. Copilot works in its free consumer version and in Microsoft 365 Copilot Chat, which is included at no additional cost with eligible Microsoft 365 plans.
- Give the model tool names and job types, never the documents. No PHI, no material under NDA, no confidential figures.
- The model's failure list is a hypothesis about a category of tool until your corrections make it about yours.
- One page is a personal record. Your employer's obligation under the AI Act is wider than that, and questions on your specific facts belong with counsel.