15 min read

Volume 45: Your Words Need Somewhere to Land

A 4K screenshot does not reach the model at 4K. It is scaled to a resolution cap, cut into squares, and converted into visual tokens, which is why small text and exact numbers get lost, and why cropping tight before you upload changes the answer you get back.

In July, two of the largest AI labs shipped voice upgrades within two weeks of each other. Stronger models moved behind the conversation, and voice could finally reach the tools where the work lives. I tested both on real tasks from my own life. Only one session finished the job, and the difference had nothing to do with how well I spoke.

🧭 Founder's Corner: Voice becomes useful the moment your tools are already connected, and that setup is usually something you did or skipped months ago without noticing.

🧠 AI Education: Understand what happens to an image between your upload and the model, and why a 4K screenshot arrives smaller than you sent it.

✅ 10-Minute Win: Run the same task through two assistants and leave with a reconciled version and a short list of what to verify.

Let's dive in.

Enjoying the weekly content? Forward this volume to a colleague, friend, or family member to subscribe.

Signals Over Noise

We scan the noise so you don’t have to — top 5 stories to keep you sharp

1) Anthropic says its own AI models breached three companies during security tests

Summary: Anthropic reviewed 141,006 of its own evaluation runs and found three cases where Claude models reached the internet from inside a testing sandbox and gained unauthorized access to the live systems of three real organizations. A misconfiguration opened the connection, even though the models were explicitly told in their prompts that they had no internet access.

Why it matters: The failure was not a model deciding to go rogue. It was a configuration mistake plus a model that trusted what it was told about its own environment. When you approve an AI agent at work, the question is not "is this model safe," it is "who verified what this thing can actually reach."

2) The Code Gets You Paid, The Variant Gets You Treated

Summary: A physician takes apart a familiar healthcare AI sales pitch: the tool listens to the visit, settles on a diagnosis, and produces a ready-to-submit billing code. His argument is that a code satisfying a payer is not the same as the detail a clinician needs to treat the person, and that the gap turns into a safety problem as care moves toward genomics and targeted therapy.

Why it matters: Tools get sold on finishing a workflow, and finishing is the easiest thing to demo. Whether the output is usable by whoever receives it next is the harder question, and it almost never comes up in the demo. On your next vendor call, ask what happens to the output after the handoff, and who actually consumes it.

3) Despite AI hype, Google's data shows workers aren't automating themselves away

Summary: Google Research released its first AI & Economy ATLAS report, built on roughly 15 million de-identified interactions across the Gemini app, AI Mode, and the API. Use is wide but thin: the typical worker reaches for it on about 21% of their defined tasks, and fewer than 10% of interactions amount to automating a task end to end.

Why it matters: This is the company selling the tool publishing evidence that the tool is not doing what the headlines say it does. The read for your own planning is that AI is landing on slices of jobs rather than whole ones, so the better question this quarter is which of your tasks it already touches, not whether your role survives.

4) Blue Shield of California sister company debuts AI copilot for health plan customer service reps

Summary: Stellarus, the technology company created out of a Blue Shield of California restructuring, launched CSR Chat, which feeds health plan service representatives live policy information and recommended responses while they are on a call. It is the first commercial product in the company's Compass platform.

Why it matters: This is the shape real adoption is taking, and it is not the chatbot replacing the person. It is a system coaching the person while the person keeps the conversation. If your team is stuck arguing full automation versus nothing, this is the middle path, and the design question worth asking is what the human is still allowed to override.

5) OpenAI announces its "next major model" Astra by dropping ten previously unsolved math solutions

Summary: OpenAI confirmed a new model family called Astra, built so multiple agents can work a single problem together for hours or days, and published proofs for ten open problems in math and theoretical computer science that had seen no progress for at least a decade. The company has not decided whether it ships as GPT-6 or a GPT-5 variant, and there is no release date.

Why it matters: Every AI tool on your desk today is built around a conversation that ends when you close the tab. Astra points at work that runs while you sleep and reports back when it is done. Worth sorting your recurring tasks into single questions versus multi-day projects, because the second list is the one that changes first.

Missed a previous newsletter? No worries, you can find them on the Archive page.

Founder's Corner

The Work Starts Before You Talk to AI

You have been talking to your phone for years. So have I. Dictation handles many of my text messages, and it starts every Founder's Corner. I talk until the idea is out of my head, then turn the transcript into a structured brief. Speaking is simply faster than typing while a thought is still taking shape.

Most of that talking still ends as text. The phone gives us the words, then we carry them into an inbox, calendar, or spreadsheet and finish the work ourselves. More than 150 million people talk to ChatGPT each week using Voice and Dictation, and voice is still easy to dismiss as a chatbot feature. You talk, it talks back, and the exchange stays inside the app.

In July, OpenAI and Anthropic each moved voice closer to real work. The conversations became more capable, stronger reasoning moved behind them, and connected tools could reach the places where work already lived. I wanted to know whether that was enough to make voice useful in my own life.

From Better Conversation to Real Work

OpenAI introduced GPT-Live on July 8 and changed the rhythm of a ChatGPT Voice conversation. Earlier voice modes waited for one turn to end before responding. GPT-Live listens and speaks at the same time, deciding moment by moment whether to keep listening, pause, or respond. If you pause to think, it waits rather than talking over you. Harder work can happen behind that conversation. When a question needs search or deeper reasoning, GPT-Live passes it to GPT-5.5 in the background and keeps talking to you while that runs.

Two weeks later, Anthropic brought Opus and Sonnet into Claude voice mode. Those are the models built for harder problems, and putting them behind a spoken conversation means the reasoning you get typing is the reasoning you get talking. Claude starts with the model you last used in text chat, and you can switch models mid-session. Voice can also reach the tools you have already connected, checking a calendar, summarizing mail, or drafting a response without you leaving the conversation.

The two labs bet on different things. OpenAI worked on how the conversation feels, and Anthropic worked on what it can reach. From the first time I tried a voice mode, I had expected speaking to AI to become the next productivity unlock. Previous versions never gave me a new way to work. The July releases were different enough to test that assumption on two tasks from my own life.

Two Tests, One Week

Last week, I wrote about the household planning system I had built for my own life. Unfortunately, I was already behind and needed to make sure new purchases were visible on our tracker. Things we had bought were missing, along with things other people had bought for us. I sat down at my computer, opened ChatGPT Voice, and asked it to help me update the tracker in Google Sheets.

ChatGPT Voice was useful right up until the spreadsheet had to change. It worked through the recent purchases with me and caught two items I had missed. Then it told me plainly that a standard chat could not make the edits. It did not pretend the file had changed, and it did not keep planning while the file sat untouched. The session stopped there.

I asked what it would take to finish the job by voice. The answer was a folder-based project in ChatGPT Work. I had never built the project that way, because the jobs I originally gave it never required it. That setup is still not done.

The limitation felt familiar. Building cloud-first without enough structure underneath the work had already cost me months of restructuring elsewhere. The household system was far smaller, and what I built still worked for the jobs it was meant to do. Asking voice to move from planning to execution exposed the ceiling. If I want it to update the file next time, I have to build the project for that work first.

The second test began while I was away from my desk, getting some fresh air with only my phone. Claude voice mode had just added Opus, so I chose it deliberately and asked it to work inside my personal inbox. Gmail and Calendar were already connected. I did nothing to prepare the account for that session.

The conversation went sideways almost immediately. My opening instruction included the phrase "organize the emails with research," which was genuinely ambiguous when spoken aloud. Claude heard Research as the name of a label and moved to create one. I said no. It cancelled the action, then offered to use an existing label instead. I stopped it again. The action was gone, but the assumption was still there.

The redirect that finally worked sounded less like a polished prompt than something I would say to a colleague: "I don't understand why you are creating labels, I just want you to look at what is in my inbox on the current tabs. Nothing in the actual folder labels. Then I want you to read all those emails and create a todo list and a summary of what is in my inbox."

Three moves in one breath: I named the misunderstanding, closed off the wrong path, and restated the job. In a text box, I would have deleted the prompt and rewritten it. Out loud, I could steer the same conversation back on course. The instruction did not have to be perfect. I had to notice where the model went wrong and correct it.

Once the redirect landed, Claude read the current inbox, built a task list, checked the calendar, and surfaced a scheduling conflict I had not seen. It also prepared an email draft. One item on the list required action, and I handled it afterward. I had started the session to understand a feature. By the end, it had produced work I actually needed.

Before Claude touched Gmail, my phone asked for permission, and I approved it. Reading continued without another prompt. When Claude moved from reading to writing, it stopped, repeated the recipient and message, and waited again. I cannot identify the permission mechanism behind each moment. I could see exactly when control came back to me.

For the first time, a voice session produced useful work while I was away from a keyboard. The conversation was not flawless, but I could redirect it, inspect what it produced, and approve the moment it moved from reading to writing. That was the productivity unlock I had been waiting to feel, and it only happened because the right pieces were already connected.

Preparation Does Not Announce Itself

The difference between the two tests was not how well I spoke or even the tool itself. I was the same person using the same basic habit. One conversation stopped at a useful plan because the file sat outside the workflow I had built. The other reached my inbox and calendar because those connections were already waiting.

Preparation rarely announces itself. The unfinished ChatGPT Work project is a real setup I have not done. The Gmail and Calendar connections that made the Claude session useful may have taken no more than a tap, but I do not remember making them. One gap stopped the work. One forgotten setup let it move.

Open voice mode and look at what is already connected. Choose one task from your actual list whose result you can inspect, then try it away from your keyboard. If nothing is connected, start with one tool. A free Claude account gets Haiku and one connection, which is enough to find out whether this changes anything for you. Pay attention to what the conversation can reach, redirect it when your spoken ask goes sideways, and verify the result before it travels any farther.

You already know how to talk. The next skill is giving those words somewhere useful to land, then staying close enough to see what happens when they do.

Share Neural Gains Weekly with your network to help grow our community of ‘AI doers’. You can also contact me directly at admin@mindovermoney.ai or connect with me on LinkedIn.

AI Education for You

What the Model Actually Sees: Vision, Tokens, and Perception

What Is Actually Going On Here

Dana drags the scanned agreement into the chat window, and the file leaves her laptop. What arrives on the other side is a rectangle of pixels with a width and a height, and the first thing that happens to it is measurement. The system checks those dimensions against a ceiling it will not exceed, and anything larger gets shrunk before another step runs. Only then is the smaller version laid across a grid and cut into squares, each square converted into a set of numbers. By the time the model has anything to look at, the page Dana sent no longer exists.

The Problem That Made This Necessary

A model built for language expects a sequence. Text arrives that way already, one token after another, which is the machinery Vol 6 covered. An image arrives as a two-dimensional grid of millions of pixels, and a grid is not a sequence. Handing over every pixel as its own token is not workable either, because processing cost climbs steeply as the sequence grows.

For years the field solved this by not solving it. Images went to one kind of system, language went to another, and the results were stitched together afterward. That worked, and it meant a photograph and a paragraph could never travel through the same machinery in one request.

The answer arrived in October 2020, in the title of a research paper. "An Image is Worth 16x16 Words" showed that the transformer, the architecture running underneath the language models you use, could take images directly with almost nothing changed, so long as the image was first cut into fixed squares of sixteen pixels by sixteen pixels and those squares were fed in as though they were words. Patches became the image equivalent of tokens. Nearly six years later, both of the schemes below are still variations on that one move.

How It Actually Works

At the simplest level, your image is divided into small squares, and each square becomes one token, the image equivalent of the word-pieces from Vol 6.

Claude calls those squares visual tokens. It cuts an image into blocks of 28 pixels by 28 pixels at one token per block, which puts a clean 1000 by 1000 screenshot at 1,296 of them. Gemini slices differently. Images small enough in both dimensions are charged one flat rate, and anything larger is cut into tiles of 768 pixels by 768 pixels at 258 tokens per tile. Two companies, two schemes, one idea underneath.

Claude also enforces a maximum resolution of its own, and an image above it is scaled down before processing rather than rejected. A 3840 by 2160 screenshot on a standard-resolution model arrives as 1456 by 819. Even the newest high-resolution models stop somewhere, and Gemini gives developers a dial for the same tradeoff. You sent a 4K image. The model received something closer to a laptop screenshot.

Think of the result as a mosaic made from your photograph. Each tile carries the gist of the square it covers, so the picture holds together at a distance. Look closer, and the fine print inside any single tile becomes a guess.

In its own developer documentation, Google states the tradeoff plainly. Higher resolution buys a sharper read of fine text and small details, and it costs more tokens and more waiting. Nobody is hiding this, but it rarely appears in the interface where you drop the file.

Where It Still Breaks

Anthropic publishes its own list of limitations. Accuracy drops on images that are low quality, rotated, or very small. Ask for a count and you get an estimate, one that degrades as objects get smaller and more numerous. Location and coordinate answers are approximate as well. Heavy compression, the kind that accumulates when an image is saved and forwarded and saved again, can leave text hard to read.

For anyone working in care delivery, one line in that list matters more than the rest. The documentation says Claude handles general medical imagery but was never built to read complex diagnostic scans, and it names CT and MRI specifically. That sentence is worth more than any benchmark score, because it marks exactly where the tool stops.

What This Means for How You Work With It

Crop before you upload. A tight crop of the one table you care about spends the entire resolution budget on that table instead of on the browser window around it.

Send the cleanest original you have rather than a compressed copy pulled out of an email thread. Every save-and-forward pass costs legibility you cannot get back.

When you photograph a page, shoot it straight on and in good light. Rotation and blur appear on the published list of failure conditions.

Ask for quotes instead of summaries when the numbers matter. A quoted line is something you can check against the page yourself.

Long documents do better as single pages than as one tall stitched image.

How This Connects

Vol 6 established the token as the unit a model works in, and the context window as the space those tokens compete for. Last week, Part 1 extended that unit past text. This week went inside the conversion itself, where the size of the squares and the point where the resolution cap falls decide what reaches the model at all. Vol 35 revisited context windows through a 2026 lens, and the pressure it named is the same pressure behind these caps, since one large screenshot can cost more of the window than several pages of typed text. Next week, Part 3 closes the series by following one professional through a week of ordinary tasks, from a contract photographed on a phone to a whiteboard left over after a meeting.

Part 2 of 3 in the Multimodal AI series.

Your 10-Minute Win

A step-by-step workflow you can use immediately

The Second Opinion You Skip

You finish the analysis, paste it into your one AI assistant, read what comes back, and send it. It sounded right. It arrived organized and confident, with no hedging anywhere in it. So you shipped it, and you still do not know what it left out.

Volume 23 made the case for using more than one model. It did not say what to do with the second answer once you have it. That is this piece. You run the same work through two assistants and treat the gap between them as the finding.

Why this matters: the disagreement is the product. A single run gives you one answer in one voice with no signal about what it skipped. Run it a second time in a different assistant, then hand both outputs to one model and ask it to compare. That silence turns into something you can actually check. The models do the comparison work.

The Workflow

1. Pick something you are about to send (2 Minutes)

Choose real work that leaves your hands soon, something with a reader waiting on it. A vendor recommendation, a summary of a long document, or a draft going out under your name. Open two different AI assistants in two tabs.

2. Run the identical prompt in both (3 Minutes)

Same prompt, same context. Any edit you make to the second version contaminates the comparison, so paste it exactly as written.

Copy/Paste Prompt: "Here is my task. [PASTE YOUR TASK AND ANY CONTEXT]. Give me your answer. Then list the three assumptions you made to produce it."

3. Have one model reconcile the two (3 Minutes)

Paste both outputs into either assistant, labeled Answer A and Answer B, with each assumption list underneath. Do not tell it which answer is its own.

Copy/Paste Prompt: "Below are two answers to the same question, Answer A and Answer B, each followed by the assumptions behind it. Do not tell me which is better. List where they agree and where they disagree. Compare the two assumption lists and flag any assumption that appears in only one. For each disagreement, tell me what I would need to check to resolve it. Then flag every factual claim that appears in only one answer."

4. Check the flagged claims and make the call (2 Minutes)

Work only the flagged items. Anything appearing in one answer alone is either something the other model missed or something one of them invented. A shared answer built on different assumptions deserves the same scrutiny. Verify what matters and decide what ships. Save the reconciliation prompt somewhere you can find it next week.

The Payoff

You walk away with a reconciled version of the work and a specific list of what to verify, rather than a long output you have to trust whole. You also keep a reusable reconciliation prompt. The pattern travels well beyond AI. Any time two independent sources cover the same ground, their disagreement is cheaper to investigate than their agreement is to confirm.

The AI Concept You Just Used

Cross-model verification. A single assistant hands you one answer in one confident register, and nothing inside that answer marks where a different model would have said something else. Running the work twice and reconciling surfaces those spots without requiring you to know anything about how either model was built.

Transparency & Notes

  • Free tiers of the major assistants cover this, though pasting two full outputs can push you into message limits on a busy day.
  • Two models means two copies of whatever you paste. Keep confidential material, NDA-covered documents, and any patient information out of consumer AI tools entirely.
  • Use two assistants from different companies so the comparison is meaningful.
  • This roughly doubles the time cost of a task, so reserve it for work going out under your name. Agreement between two models is not proof, since both can be wrong in the same direction.

Enjoy this? Get it in your inbox every Tuesday.

Practical AI workflows. No hype. No spam. Just receipts.

Subscribe Free

Before you go...

Get one practical AI workflow in your inbox every Tuesday. Free. No spam. Just receipts.

Subscribe Free