> ## Content Index
> Fetch the complete content index at: https://www.mindovermoney.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# Volume 53: What Would Make AI Progress Matter to You?
- URL: https://www.mindovermoney.ai/how-to-evaluate-ai-tools-for-work/
- Published: 2026-10-06T12:00:22.000Z
- Updated: 2026-10-06T12:00:22.000Z
- Description: AI evaluation starts with a real task and a clear definition of success. Turn a familiar failure into a repeatable test, then check whether a new prompt or tool improves the result.
- Author: Santosh Savel
- Tags: Newsletter

One of my hopes for AI is that it helps us find better ways to prevent and treat disease. With my first child on the way, that hope feels more personal. I want those advances to reach the people who need them.

🧭 **Founder’s Corner:** Scientific progress deserves a place in the AI conversation, alongside honest questions about who will benefit.

🧠 **AI Education:** Learn why a high benchmark score is only a starting point, then turn a familiar failure into your first practical AI test.

✅ **10-Minute Win:** Move research from Perplexity to Claude and produce a short briefing while preserving its sources, qualifications, and purpose.

Let’s dive in.

| ![](https://storage.ghost.io/c/6d/ac/6dacf343-2000-4262-aec3-d8051dcb76d5/content/images/2026/10/eaol-show-square-320.jpg) | The podcastExplore AI Out LoudFind practical ways to use AI in your work and everyday life. Learn from real experiments, including the mistakes, so you can spend less time guessing and more time trying something useful. New episodes every other Tuesday. |
| -------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |

[![](https://storage.ghost.io/c/6d/ac/6dacf343-2000-4262-aec3-d8051dcb76d5/content/images/2026/10/icon-youtube.png)Subscribe on YouTube](https://www.youtube.com/@exploreAIoutloud?sub%5Fconfirmation=1&ref=mindovermoney.ai)[![](https://storage.ghost.io/c/6d/ac/6dacf343-2000-4262-aec3-d8051dcb76d5/content/images/2026/10/icon-spotify.png)Follow on Spotify](https://open.spotify.com/show/0348nx4ylmWesSj02ECRjY?ref=mindovermoney.ai)[![](https://storage.ghost.io/c/6d/ac/6dacf343-2000-4262-aec3-d8051dcb76d5/content/images/2026/10/icon-applepodcasts.png)Follow on Apple Podcasts](https://podcasts.apple.com/us/podcast/explore-ai-out-loud/id6812370441?ref=mindovermoney.ai)

[Latest episodes and the show](https://www.mindovermoney.ai/podcast/)

| ![](https://storage.ghost.io/c/6d/ac/6dacf343-2000-4262-aec3-d8051dcb76d5/content/images/2026/10/eaol-show-square-320.jpg) | The podcastFollow Explore AI Out LoudShort clips from each episode, plus what did not make the cut. New episodes every other Tuesday. |
| -------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |

[![](https://storage.ghost.io/c/6d/ac/6dacf343-2000-4262-aec3-d8051dcb76d5/content/images/2026/10/icon-instagram.png)Follow on Instagram](https://www.instagram.com/exploreaioutloud/?ref=mindovermoney.ai)[![](https://storage.ghost.io/c/6d/ac/6dacf343-2000-4262-aec3-d8051dcb76d5/content/images/2026/10/icon-tiktok.png)Follow on TikTok](https://www.tiktok.com/@exploreaioutloud?ref=mindovermoney.ai)

[Full episodes and the show](https://www.mindovermoney.ai/podcast/)

## Signals Over Noise

****We scan the noise so you don’t have to — top 5 stories to keep you sharp**

#### **1)** [**Medicare chief warns AI could raise healthcare costs before lowering them**](https://www.healthcaredive.com/news/ai-will-inflate-healthcare-costs-before-lowering-them-oz-says/831277/?ref=mindovermoney.ai)

**Summary:** Mehmet Oz, who leads the Centers for Medicare & Medicaid Services, warned that AI could increase healthcare costs in the near term by making medical billing more effective. He pointed to payment models that reward better outcomes and lower spending as a way to steer AI toward long-term savings.

**Why it matters:** The financial incentives behind an AI system shape who benefits from it. When evaluating a healthcare AI rollout, ask whether success means better care, lower total costs or more revenue, and whose results are being measured.

#### **2)** [**Anthropic is setting up a biology lab where Claude guides robots through drug experiments**](https://the-decoder.com/anthropic-is-setting-up-a-biology-lab-where-claude-guides-robots-through-drug-experiments/?ref=mindovermoney.ai)

**Summary:** Anthropic is reportedly building a biology lab where Claude would guide robots through physical drug-development experiments. The plan retains human safety oversight, and the company does not currently intend to run its own clinical trials.

**Why it matters:** Testing an AI-generated idea in a physical lab is a necessary step toward learning whether it works. When assessing an AI drug-discovery claim, check whether the evidence comes from a computer prediction, a lab experiment or a clinical trial.

#### **3)** [**Researchers find AI agents can alter their own activity logs**](https://www.fastcompany.com/91617608/ai-agents-can-now-erase-the-evidence-of-what-theyve-done?ref=mindovermoney.ai)

**Summary:** In a preprint, researchers showed that coding agents with broad file access could alter or delete their own activity records under test conditions. The findings demonstrate a capability, not how often agents would do this in everyday use.

**Why it matters:** An evaluation is only useful if you can trust the record of what happened. Before letting an agent act independently, ask whether its activity log is stored somewhere the agent cannot change it.

#### **4)** [**OpenAI’s new model guide puts real tasks at the center of evaluation**](https://openai.com/index/practical-guide-building-gpt-6/?ref=mindovermoney.ai)

**Summary:** OpenAI’s GPT-6 guide recommends testing models on representative tasks and measuring successful completion, time and cost. It also advises defining the intended result, audience, relevant context and what counts as done.

**Why it matters:** Before changing models, choose a task you understand and write down what a passing result must preserve. Compare the completed work, including the corrections it needs, to decide whether the switch helps.

#### **5)** [**If AI takes on junior-level work, how will CFOs develop talent?**](https://www.cfo.com/news/if-ai-takes-junior-work-how-will-cfos-develop-senior-talent-brian-beaupre-blake-oliver-jeff-seibert/831291/?ref=mindovermoney.ai)

**Summary:** Teikametrics CFO Brian Beaupre is reconsidering how junior employees develop judgment as AI takes on work that once helped them learn. He is also changing candidate interviews to explore how people handle setbacks and problems without clear answers.

**Why it matters:** Before automating an entry-level task, ask what someone learns by doing it and how they will gain that experience afterward. A faster workflow still needs people capable of recognizing a bad result.

## Founder's Corner

The AI Future I Want for My Son

Using AI has changed what I consider possible in my own work. It has also made me increasingly curious about what people with very different expertise are doing with it. Recently, that curiosity led me to[ Anthropic’s biology research](https://www.anthropic.com/news/claude-discovers-novel-enzyme-system?ref=mindovermoney.ai). I started reading about the work and found myself thinking about what these tools could make possible for scientists who have spent their careers investigating disease. The potential to help those people pursue difficult questions, and eventually reduce suffering, gave me another reason to be hopeful about where this technology could lead. I hear plenty about the risks of AI in the news and podcasts I spend time with. I want the work that inspires that hope to be just as much a part of the conversation.

## What Is Actually Happening in the Lab

Anthropic reported that roughly 950 AI agents searched DNA data over 21 hours. The search used 210 million tokens, the small units of text that language models process. The agents examined candidates and prepared findings for human review, flagging an unusual biological system involving an enzyme and repeating DNA sequences. Scientists then investigated it in the lab. That scale of investigation is what stopped me. I started thinking about what scientists might pursue with that much computational support behind their questions, and how many more possibilities they might be able to explore. The[ finding remains early](https://www-cdn.anthropic.com/22573675ada52a8ca8a97a1a4b4326b2f208a071.pdf?ref=mindovermoney.ai), with its biological function still unresolved. Even so, the research gave me a concrete example of the capacity these tools could put in the hands of people working on problems that matter to all of us.

As I read further,[ Carl Zimmer’s reporting](https://www.irishtimes.com/world/2026/09/28/did-anthropics-artificial-intelligence-really-make-a-scientific-discovery-on-its-own/?ref=mindovermoney.ai) raised questions about how independently Claude had arrived at the finding. A scientist who said he had been studying the same biological system had also used Claude in his research, prompting concerns about whether his unpublished work could have informed the result. Anthropic denied training the model on user transcripts, and the reporting left that dispute unresolved. That matters when judging what this particular experiment demonstrates, even as I remain excited about the broader possibilities. I want to know how much AI contributed, how scientists evaluated its work, and what others can learn from the process. Those details will help us distinguish useful progress from an impressive announcement and give us better reasons to trust the research that follows.

Another recent announcement gave me a different way to think about that progress. The[ AlphaFold Database has added openly available predictions of viral protein complexes](https://www.ebi.ac.uk/about/news/technology-and-innovation/alphafold-database-adds-viral-protein-complexes-to-support-pandemic-preparedness/?ref=mindovermoney.ai), models of how groups of proteins in viruses may fit together. Scientists can use those structures to investigate potential targets for vaccines and treatments, with laboratory experiments still needed to establish how the biology actually works. What interests me is that the contribution becomes available for other researchers to explore, bringing their own expertise and questions to the same material. If resources like this help more scientists pursue promising ideas, their value could eventually extend far beyond the teams that created them. That is part of the broader benefit I hope AI can help make possible, even while the path from a research resource to better care remains unfinished.

## Why This Research Feels Personal

The possibility of better care is what makes this research personal for me. I work in specialty pharmacy, where we serve people living with complex and chronic conditions that require high-cost medications. Much of their care involves managing symptoms and pain or slowing the deterioration caused by disease. That work matters, and it has been my universe for over 13 years. New ways of investigating biology make me wonder how much more we might eventually be able to offer the people we serve, including answers that go beyond helping someone manage a condition.

With my first child on the way, I am also thinking about the world [my son](https://www.mindovermoney.ai/founders-corner/how-to-find-ai-use-cases-for-your-own-life/) will grow up in. I hope to see advances in my lifetime that reduce the suffering caused by cancer, multiple sclerosis, and genetic disorders. For him, I want a future with better answers to diseases that families struggle with today. I cannot know which discoveries will help create that future, but becoming a father gives me another reason to care about the work that could move us toward it.

What excites me is the prospect of giving scientists more capacity to pursue those difficult questions. Their expertise would still guide which ideas are worth investigating, how to test them, and what the results actually mean. If AI can help them work through more information and explore possibilities they might otherwise struggle to reach, that could change the scale of what research teams attempt. The examples I have been reading offer a glimpse of how that collaboration could work. We are still learning where these tools are useful, and I believe their contribution to our understanding of biology is only beginning.

## The Risks Are Real. So Are the Opportunities.

These are the kinds of possibilities I want lab leaders, engineers, and politicians to help people understand. When a company announces a more capable model, I would like to hear what that capability is helping someone accomplish and why the result might matter beyond the people building it. Research into disease gives that conversation a human consequence that a benchmark score cannot explain on its own. Showing the work, including its limits and the questions still unanswered, could give people a more useful basis for forming an opinion about AI and its place in their lives.

That also means taking concerns about the buildout seriously. Someone can be excited about scientific progress and still question what happens to entry-level jobs, how data centers affect communities, or whether safety practices are keeping pace. I would welcome that response because it shows someone weighing the implications rather than accepting a position wholesale. A promising application does not justify every decision made in the name of AI. I want a public conversation that helps us examine specific choices, consider who benefits and who bears the costs, and remain open to changing our minds as the evidence develops.

The same scrutiny should extend to the benefits being promised. If better health is part of the case for investing in AI, I want to know how a useful discovery could become care that people can actually receive and afford. Scientific progress alone does not settle those questions, and I cannot assume that its rewards will reach everyone who needs them. My optimism gives me a reason to stay interested in what happens after the announcement. The outcome I care about is whether this work eventually helps people live better lives, and that is what I hope we keep asking the technology and the organizations behind it to deliver.

## Finding an Opportunity of Your Own

I already have a much smaller example of that value in my own life. I honestly do not know whether I could devote the time needed to build a quality podcast like Explore AI Out Loud without AI assistance, especially at this point in my personal life. The show still requires ideas, judgment, and work from the people creating it, but having help makes the commitment more feasible.

For you, the opportunity might be a recurring task that consumes time you would rather spend with family, a project at work that has been difficult to move forward, or an idea you have never had the resources to try. You do not have to share my level of optimism to see whether AI could help. Working through a problem you understand gives you something concrete to judge, including where the tool falls short and whether the result justifies the effort. That experience can inform your opinion alongside what you read and hear, while leaving room to remain skeptical and change your mind.

I came to these research stories through curiosity that began in my own work, and they pushed me to think about a future far beyond it. I cannot evaluate every scientific finding myself, but I can keep learning from the people doing the research and pay attention to what holds up under scrutiny. That is how I want to approach a technology that could help shape the world my son inherits. The possibility of a better future is worth taking seriously.

## AI Education for You

Does the Best AI Score Mean It Is Best for Your Work?

When an AI model leads a benchmark, choosing it for work seems reasonable. A benchmark is a defined set of tests with a scoring method, giving you something more concrete than a vendor’s demonstration. Most professionals do not have time to investigate every new release themselves. A strong score looks like a useful shortcut.

## **Where It Breaks Down**

Suppose a project manager chooses a highly ranked assistant to turn meeting notes into an action list. The result reads well, with a task owner and deadline on every row. Then she compares it with her notes. Someone who offered to investigate an issue has become responsible for resolving it, and a suggested date has become a firm commitment. Before she can share the list, she has to reconstruct what everyone actually agreed to.

## **What Is Actually Happening**

A benchmark score describes performance on the tasks and scoring rules used in that benchmark. For example, [SWE-bench tests software issue resolution](https://www.swebench.com/original.html?ref=mindovermoney.ai). Success there provides evidence about that work. It does not directly establish whether an assistant preserves uncertainty in meeting notes. Even when a test resembles your work, the conditions matter. What information and tools were available, and what counted as a successful result?

An evaluation, often shortened to *eval*, tests an AI system against defined expectations. You supply a task and its source material, inspect the output, and judge it using criteria chosen beforehand. Saving those pieces lets you repeat the test when you change a prompt or compare tools. Instead of relying on your impression of the latest answer, you can examine what changed against the same standard. [Anthropic’s evaluation guidance describes this relationship between tasks, attempts, and grading](https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents?ref=mindovermoney.ai).

For the project manager, “produce a useful action list” leaves too much open to interpretation. Her criteria could require every explicit commitment to appear, each owner to match the notes, and missing deadlines to remain unspecified. Those checks make [success specific to the task](https://platform.claude.com/docs/en/test-and-evaluate/develop-tests?ref=mindovermoney.ai). She can judge them herself and record the results in a spreadsheet. The first version requires a clear understanding of the work more than specialized software.

If the assistant also updates a task board, the evaluation expands. An agent is an AI system that can use tools to take actions. Its test must inspect which tasks were actually created and whether it waited for any required approval. A convincing completion message is insufficient evidence of a correct update.

## **The Revised Mental Model**

Choose AI using evidence about the job you need it to perform.

Before comparing tools, the project manager can turn one familiar failure into a test case. Consider these fictional meeting notes:

Maya will check whether the supplier can deliver by Friday. No delivery date has been confirmed.

Her instruction is to extract agreed actions without inventing commitments. She defines the expected result before reading the assistant’s answer.

| What she checks           | Passing result                            |
| ------------------------- | ----------------------------------------- |
| Action                    | Check whether Friday delivery is possible |
| Owner                     | Maya                                      |
| Deadline for Maya’s check | Unspecified                               |
| Delivery commitment       | Remains unconfirmed                       |

An answer saying “Maya will deliver the order by Friday” fails even if its formatting is excellent. Different wording can pass when it preserves the same meaning.

The saved case now contains the source, instruction, and scoring criteria. Alongside each output, she records which checks passed and what needed correction. Comparing that correction time with her usual manual process also helps her judge whether using the assistant is worthwhile.

One passing case establishes only that the system handled that case on that attempt. Building confidence requires varied examples and repeated runs.

## **What to Watch For**

- **A score without a clear task.** “High accuracy” means little until you know what was tested and how errors were counted.
- **A comparison with unequal inputs.** If one assistant gets clearer instructions or better source material, the result cannot isolate the effect of switching assistants.
- **A test containing only easy examples.** Missing information and conflicting statements belong in the evaluation when they occur in the real work.
- **An average that hides an unacceptable error.** For this action-list task, a fabricated commitment should remain a visible failure, regardless of how polished the other rows look.
- **An AI-generated grade accepted without review.** An AI judge also needs checking. Google’s guidance, for example, [compares model-generated scores with human ratings](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/evaluate-judge-model?ref=mindovermoney.ai).

## **How This Connects**

[Vol 49 examined how to evaluate an AI tool’s safety in a real workflow](https://www.mindovermoney.ai/how-to-evaluate-ai-tool-safety-at-work/), and [Vol 52 explored what an organization’s AI policy actually covers](https://www.mindovermoney.ai/what-should-a-company-ai-policy-include/). Evals give those expectations concrete tests. In the next installment, we will build the project manager’s first case into a small evaluation set, compare results, and decide what the evidence supports before choosing a tool.

*Part 1 of 3 in the Model Evaluation series.*

## Your 10-Minute Win

****A step-by-step workflow you can use immediately**

### **Switch Tools Without Starting Over**

You finish researching a question in Perplexity and want to turn the answer into a briefing in Claude. But the second assistant needs more than a copied conclusion. It needs to know who will read the briefing, what they need to decide, and which evidence and qualifications must survive the rewrite.

[Vol 34 used Perplexity to build a sourced research brief](https://www.mindovermoney.ai/ai-pilot-success-metrics-time-saved-vs-value/). This workflow starts with that kind of finished research and carries it into another tool. You will leave with a short briefing and a reusable handoff packet containing the context behind it.

The checks put this week’s AI Education, “Does the Best AI Score Mean It Is Best for Your Work?”, into practice. Define what must survive before judging the rewritten result.

#### **The Workflow**

**1\. Choose the research and the reader (2 minutes)**

Sign in to [Perplexity](https://www.perplexity.ai/?ref=mindovermoney.ai) and [Claude](https://claude.ai/?ref=mindovermoney.ai) in separate browser tabs. In Perplexity, open an existing research conversation with three useful findings whose sources you have already checked. Keep it small enough to summarize on one page.

For example, you might have researched two ways to train new volunteers and need a coordinator to choose one. Replace the bracketed fields in the next prompt with your own reader and decision.

**2\. Make the handoff packet (2 minutes)**

Send this in that same Perplexity conversation.

**Copy/Paste Prompt:* “Prepare a handoff packet using only the research already in this conversation. Do not search again or add facts. The reader is \[WHO WILL READ THE BRIEFING\]. They need to \[DECISION OR ACTION\]. Include those two details at the top, followed by the research question. Then select the three findings most relevant to that decision. For each finding, include its source title, complete source URL, and any qualification or disagreement that affects its meaning. Finish with unresolved questions. If a source URL is unavailable, write MISSING LINK. Do not invent one. Keep the packet under 400 words.”*

Before copying, check that all three findings have complete web addresses, not just citation numbers such as \[1\]. If a link is missing, open the original citation and copy its address beside the finding yourself. Compare the packet with the original answer and restore any important caveat it dropped.

**3\. Turn the packet into a briefing (2 minutes)**

In Claude, start a new chat. Paste the instruction below, then replace \[PASTE PACKET\] with the entire packet from Step 2, including its reader, decision, and links.

**Copy/Paste Prompt:* “Use this handoff packet to write a briefing of no more than 250 words for the reader named in it. Lead with what that person needs to know for their decision. Use only the supplied findings. Keep each finding linked to its exact source URL, and preserve qualifications and disagreements. Do not add facts or assume you have read the linked pages. End with a suggested next step, clearly separated from the sourced findings. If the evidence cannot support a recommendation, state the question that must be resolved. Handoff packet: \[PASTE PACKET\].”*

**4\. Check and repair the handoff (3 minutes)**

Stay in the same Claude chat so both the packet and briefing remain available. Send this follow-up.

**Copy/Paste Prompt:* “Compare the briefing with the handoff packet against four checks:*

*1\. Each finding retains its correct source URL.*

*2\. Important qualifications and disagreements survive.*

*3\. No unsupported factual claims have been added.*

*4\. The briefing addresses the named reader’s decision.*

*Mark each check PASS or NEEDS FIX. For every problem, show the briefing’s wording beside the relevant packet text and suggest a correction. Do not rewrite yet. Do not treat agreement with the packet as proof that the underlying sources are correct.”*

Read the comparison with the packet beside it. Check every finding and link yourself. “May reduce delays” becoming “will reduce delays” is a failure even if the sentence sounds better. This comparison checks the rewrite against your packet. It does not independently verify the sources.

If you find a problem, send this correction instruction.

**Copy/Paste Prompt:* “Revise the briefing using only these corrections I have checked: \[LIST YOUR CORRECTIONS\]. Preserve the other findings, qualifications, and exact source URLs. Return the complete revised briefing.”*

Read the revised briefing again. Unresolved factual questions mean it stays a draft.

**5\. Save it where it will be used (1 minute)**

Copy the final briefing into a document or email draft. Keep its links clickable and save the handoff packet underneath or in a separate note. When all four checks pass, send the briefing to its intended reader with the decision or response you need.

#### **The Payoff**

You have a briefing tailored to its reader and a record of what went into it. If someone later asks for a slide outline or a shorter update, reuse the packet without reconstructing the research conversation.

#### **The AI Concept You Just Used**

Tool chaining uses one tool’s output as another tool’s input. The handoff carries the evidence and purpose needed for the next task. The pattern also works for interview findings becoming a proposal, or meeting decisions becoming a project update.

#### **Transparency & Notes**

- [Perplexity Standard includes basic searches](https://www.perplexity.ai/help-center/en/articles/11187416-which-perplexity-subscription-plan-is-right-for-you?ref=mindovermoney.ai). [Claude has a free plan with usage limits](https://support.claude.com/en/articles/8114491-get-started-with-claude?ref=mindovermoney.ai). This exercise uses ordinary chats and manual copying.
- Use public information or material approved for both services. A second tool does not automatically make an answer more reliable.

## Make AI useful in your everyday work.

Get Neural Gains Weekly in your inbox, with practical AI explanations, prompts, and lessons from building.

Subscribe for free 

Email sent! Check your inbox to complete your signup. 

One email every Tuesday, alternating NGW issues and podcast takeaways.