A lone baker cracking an egg at a small table under a spotlight on a dark, empty theater stage, symbolizing "judgment theater" and human oversight in healthcare AI adoption.

When the Machine Is Almost Always Right, Who Is Still Thinking?

Every few episodes we stop interviewing and start synthesizing. This is our seventh Reflections episode and our fifty-first overall, covering six conversations: Mika Newton on interoperability, Peter Embi on monitoring deployed AI, Vimla Patel on clinical reasoning, Renee Deehan on trustworthy consumer health AI, Christine Dymek on AI literacy, and Jeremy Harper on healthcare’s broken adoption curve.

We agreed at the top that we weren’t going to twist ourselves into knots manufacturing a single lesson. Three threads surfaced anyway.

Start with the good news, because there is some.

For years the roadmap for patient-mediated data access was written and didn’t work. The endpoints were coming, the regulations were coming, the patient would be able to pull their own record, and none of it functioned as advertised. Newton’s numbers say it functions now. Provider data exchange moved from roughly 30% in 2022 to 85 to 90% by 2025. What changed wasn’t a breakthrough in AI. Laws, policies and standards landed at the same moment the technology caught up, and the format that ended up carrying it was CDA. Everyone in the field would have told you FHIR. AI is still in the story, because writing transformation rules at that scale needed it. Leon’s reaction on the record: when something you’d quietly given up on starts working, that’s a real ground shift.

Human in the loop is doing more work than one phrase can contain.

Three of the six independently made it a design commitment. Deehan draws a hard line at taking the human out of a recommendation. Patel, whose research shows that an interface built for a novice actively degrades an expert, wants friction deliberately designed back in, a system that pauses and asks whether you really mean it. Harper wants a trained professional signing off on ambient notes, importing a practice healthcare already trusts from de-identification.

None of that touches the failure underneath. When a system is right 99% of the time, people accept the answer, and that’s automation bias rather than laziness. TSA screeners waved marginal bags through for years because a genuine threat almost never appears, which is precisely why training programs began planting one. The clinical version is the 80th note that looks like the 79 before it. A human in the loop doesn’t solve this. Amy Price, back on her own episode, gave us the sharper standard: a knowledgeable human who cares. We now need to add: a knowledgeable human who cares and who retains the capacity to make good decisions.

Nobody is watching after go-live, and the word for the fix is contested.

Embi coined algorithmovigilance and built VAMOS at Vanderbilt, an air traffic control tower for clinical AI that maps models to real outcomes and is being open-sourced. It’s one institution’s system, and healthcare still has no MedWatch for AI. Dymek’s answer runs upstream. If the regulatory floor is thin, your staff are the guardrail, which makes literacy the intervention. Leon pushed on the word, because metaphors import assumptions, and reading got schools and libraries as its institutional answer while AI literacy has nothing comparable. Peer reviewers on Steve’s own paper pushed from the other side, arguing that literacy has to be paired with competency. They’re right. You can be perfectly literate about a cardiac catheter and nowhere near competent to use one, and you don’t learn piano by being shown middle C.

Which leaves the closing worry. Friction is the honest fix and also the easiest thing to fake. Steve reached for the business-school story about the cake mix rebuilt so the baker had to crack a real egg. Historians have since dismantled that story, which is a useful reminder about plausible evidence in its own right, but the failure mode it names should worry us all. Add friction for the feeling of it and you get judgment theater, where people feel like they’re exercising judgment when all they’re doing is breaking eggs.

Listen to the full conversation.

You Might Also Enjoy

  • S1E47 with Vimla Patel — The cognitive science behind this block’s friction argument, including why an interface designed for novices drags expert performance down.
  • S1E49 with Christine Dymek — The full case for treating AI literacy as a policy problem rather than a training problem, and the open question about which institution answers it.
  • S1E50 with Jeremy Harper — Why healthcare skipped its early adopters, and why nobody can tell you how often an ambient note is wrong.