How Specialized Does AI Have to Be to Actually Work?
Every few episodes, we stop interviewing and start synthesizing. This is our fifth Reflections episode, and it covers six conversations: Matt Truppo at Sanofi (twice), Ted Shortliffe, Barry Chaiken, David Hidalgo-Gato, and Danny van Leeuwen. Put side by side, they tell one story.
For four blocks we’ve argued that the LLM is becoming a commodity. Block five makes the claim stronger. The generalist AI agent is a commodity too. The agentic tooling has matured to the point where almost any technologist can stand one up. What is not a commodity is depth: in a workflow, in a knowledge base, in a patient’s own data.
Specialization is doing the work. Truppo set out to build, in his words, “one super agent to rule them all.” It failed. He pivoted to eleven specialized agents and clawed back about 30% of his calendar. David Hidalgo-Gato, building ambient AI for emergency medicine, went a mile deep on one specialty while roughly 120 competitors generalized across primary care. The bottleneck has moved. It used to be building the automation. Now it is listening: understanding a workflow well enough to know what to build.
The old methods are coming back. Ted Shortliffe, one of the founders of medical AI, offered the historical correction we keep returning to. Expert systems didn’t fail in the 1980s. They were early. The methodology was sound; the compute wasn’t ready. Today’s retrieval-augmented generation and knowledge graphs on top of LLMs are that symbolic tradition returning through the side door. As Truppo put it, not every tool needs to be, or even should be, AI.
Literacy is the bottleneck. Barry Chaiken pointed us to a striking finding about who can actually use these tools. In an Oxford study in Nature Medicine, AI tools identified the right condition about 95% of the time on their own, but when ordinary people used those same tools they landed under 35%, no better than not using AI at all. A separate JAMA Network Open trial, on a harder and admittedly different task, found physicians scored around 75% with or without AI, while the AI alone scored about 90%. Put the two together and a soft pattern emerges: AI on its own outperforms everyone; clinicians using AI likely outperform patients using AI; and in both studies, adding the human actually pulled results below what the AI managed alone. The model was identical. The user was the variable. His advice is practical: don’t hand the patient the AI. Use it yourself, and give them what you made with it.
Patients can be the pattern-makers. Danny van Leeuwen, a nurse for 50 years, went undiagnosed with MS for 25 years while the pattern sat in his records. When he finally requested those records, he got a four-pound box of paper, out of order. Access to data and access to usable data, he reminds us, are very different things.
The leaders worth trusting experiment on themselves. Truppo used his own calendar as the test rig and reported the failure honestly. As Leon put it, we used to experiment on our children with new technology. Now we all get to experiment on ourselves.
So the thesis we’ve built since episode two now has five layers: mundane wins matter, history rhymes, evidence tests and constrains, infrastructure is where the real work happens, and depth is the limiting reagent that decides whether any of it pays off.
After 37 episodes, the technology is no longer the question. The specificity of the work around it is.
Listen to the full conversation: https://practicalaiinhealthcare.com/episodes/#S1E38
You Might Also Enjoy
- S1E33 with Ted Shortliffe — One of the founders of medical AI on why the expert-systems era didn’t fail, it was early, and why that history is returning as knowledge graphs on top of LLMs.
- S1E35 with Barry Chaiken — A physician and two-time cancer survivor on the literacy gap and what changes when the clinician becomes the patient.
- S1E37 with Danny van Leeuwen — A nurse of 50 years on why the diagnosis sat unsynthesized in his records for 25 years, and what patient-held data really requires.