Jeff Chuang

When 25% of Diagnoses Are Wrong: AI Takes on Rare Pediatric Cancers

Pediatric sarcomas are rare enough that even expert pathologists might see a particular subtype only a handful of times in their careers. The result? About 25% of cases get inconsistent diagnoses—one pathologist says it’s one sarcoma type, another says it’s a different type requiring different treatment.

Jeff Chuang, Deputy Director of the JAX Cancer Center and Professor at Jackson Laboratory, saw this as a problem AI could help solve. His team built a classifier that achieves 96-98% accuracy in distinguishing between sarcoma subtypes.

The Data Problem

But here’s what makes this story interesting: the AI wasn’t the hard part.

“The hard part is getting all the different groups together to agree to share their data,” as our co-host Steve put it. Chuang’s team needed images from multiple institutions—COG, St. Jude, MGH, Yale—and that meant navigating institutional boundaries.

What actually made the collaboration work? In part, the nature of rare diseases. “It’s not like we are worrying about who’s going to corner the market on treating these kinds of patients,” Chuang explained. When patient populations are tiny, competition gives way to mutual interest.

But relationships mattered too. One researcher personally connected the institutions—”basically her friends,” Chuang acknowledged.

The IRB Surprise

One institution initially proposed shipping physical tissue slides. The folks at the institutional IRB raised eyebrows. Chuang’s response: why not just scan them? They already had a process for that.

Within months, the IRB challenge was solved. The entire dataset was assembled in four months—a timeline Chuang called “really not typical for IRB things.”

The lesson: sometimes proposing a simpler alternative unlocks doors that direct requests can’t.

Why 700 Images Were Enough

Counter to the “big data” assumption, Chuang’s team trained their classifier on only 700 images. How? Foundation models trained on 100K+ histology images had already learned to recognize general features. Those models washed out the batch effects—color variations, scanning differences—that would have torpedoed a traditional machine learning approach.

The final classification layer? Logistic regression. “Something that a high school student with access to any software library could run in 30 seconds.”

The sophistication lives in the foundation model. The clinical application layer can be deliberately simple.

What’s Next: 5,000 Colors

Chuang also gave us a glimpse of emerging technology: spatial transcriptomics. Traditional H&E staining shows tissue in “two colors”—essentially pink and purple. But new techniques can visualize gene expression spatially, showing 5,000 different markers at once.

Chuang compared it to Peter Jackson’s colorized WWI footage: “When you watch this, it’s like suddenly you were there.”

The Takeaway

For anyone building healthcare AI, Chuang’s experience reinforces a pattern we’ve seen repeatedly: start with the relationships and the clinical problem. The technology will follow. And when navigating institutional barriers, sometimes the breakthrough is simply suggesting the easier path.

Listen to the full conversation in Season 1, Episode 21 of Practical AI in Healthcare: