The Definitional Chaos Hiding Under Real-World Evidence
In 2008, about five companies sold access to real-world healthcare data. Today, that number exceeds one hundred—and the growth hasn’t made real-world evidence easier to produce. It’s made the definitional problem worse.
“Does the data mean what you think it means?” That’s how the FDA defines reliability in its 2021 guidance on real-world data. And according to Dr. Aaron Kamauu, CEO of Navidence, it’s the question most RWE studies fail to seriously address.
The One-Code Problem
Consider myocardial infarction. Include only acute MI codes in your query, and you might capture the patients you expect. But exclude the “old MI” code—a single diagnosis code—and you could lose up to 30% of your cohort.
“We’ve now partnered with a data source, they’ve run the queries, and they’ve actually shown that one code alone can have a large impact on the number of patients that they find,” Kamauu explained.
This isn’t a theoretical concern. It’s the kind of definitional variance that makes studies incomparable and regulatory submissions uncertain.
GLP-1 and the Eligibility Mess
The problem scales with complexity. Take GLP-1 eligibility for obesity treatment:
- The NHS guideline requires a higher BMI threshold than US prescribing guidelines
- Researchers have used 13 different conditions to define “cardiovascular comorbidity”
- Dyslipidemia alone can be defined by lab values, diagnosis codes, treatment with lipid-lowering medications, or some combination
“It may surprise some people that there was very little consistency in definitions,” Kamauu noted.
The Case for Proactive Definition Libraries
Kamauu’s company, Navidence, takes a different approach: pre-curating libraries of operational definitions by indication before studies begin.
“Rather than reacting to a research need, we pre-curate libraries of operational definitions, code lists, and medication lists—so we already have the seven different asthma exacerbation definitions ready to review.”
The value isn’t just efficiency. It’s scientific legitimacy through documented choice.
“Knowing there were seven other definitions you chose NOT to use helps justify your choice,” Kamauu explained. “It forces a decision-making process people often skip.”
When Silos Become Visible
At one pharma company, an oncology team needed to define a respiratory condition for a comorbidity analysis. They discovered—through a shared definition platform—that a respiratory team in medical affairs had already defined COPD. The two teams didn’t know each other.
“They were using different definitions and didn’t even know it. Because they were using our system, they were able to see, ‘our colleagues define COPD this way.’ They aligned.”
Same company. Different therapy areas. No coordination until a system made their choices visible.
Infrastructure, Not Glamour
This work isn’t glamorous AI. It’s the data stewardship backbone that makes everything else work—consistent with a theme we see on the podcast, that “mundane wins matter.”
AI is accelerating the process: definition library builds that once took one to two months now take weeks. But the core problem remains human: researchers who don’t know what definitional options exist, making choices they can’t defend.
The question isn’t whether your data is big enough. It’s whether it means what you think it means.
Listen to the full conversation with Aaron Kamauu on Practical AI in Healthcare.