Quick Answer
Real-world evidence uses AI to extract regulatory-grade clinical evidence from routine healthcare data, including EHR records, claims, and patient registries. The FDA accepted 86 RWE submissions in 2022, up from 34 in 2018. Key AI methods include natural language processing for data extraction, propensity score matching for causal inference, and synthetic control arms for rare diseases where placebo controls are unethical.
A Drug Approved Without a Traditional Trial
In September 2019, the FDA updated the label for Pfizer's Ibrance (palbociclib) to include male breast cancer, a population so small that enrolling a placebo-controlled randomised trial was considered practically and ethically impossible. The agency's decision rested instead on electronic health record data, insurance claims, and patient registry information, analysed using AI-assisted statistical methods. It was, regulators later noted, one of the clearest early demonstrations of what the agency now calls real-world evidence.
That decision has since become a reference point in a quiet transformation of how drugs, devices, and biologics earn and maintain their approvals in the United States and, increasingly, in Europe and Asia. The 21st Century Cures Act of 2016 specifically directed FDA to develop a framework for using real-world data and the evidence generated from it to support new indications, post-market safety commitments, and label expansions. A decade on, the consequences of that legislative mandate are materialising in ways that are reshaping both the pharmaceutical industry and the data infrastructure underpinning modern medicine.
Understanding RWE requires distinguishing it from a related but distinct concept. Real-world data (RWD) is the raw material: electronic health records, claims databases, patient registries, wearable outputs, and even social media. Real-world evidence is what emerges when those data are subjected to rigorous analytical methods, typically AI-assisted, and used to answer a defined clinical or regulatory question. The distinction matters because FDA does not accept data quality shortcuts, even when the underlying dataset is enormous. As the agency's own framework states, fit-for-purpose is the operative criterion, not merely large-scale. For deeper context on how AI systems handle the analytics layer, see our guide to AI healthcare analytics.
The FDA's Growing Acceptance: By the Numbers
The agency received 34 RWE submissions in fiscal year 2018, the year it published its foundational framework. By fiscal year 2022, that figure had risen to 86, according to FDA's Sentinel programme annual reports. The trajectory reflects both the maturation of the analytical methods required to process observational data and a growing confidence, within FDA's Office of Surveillance and Epidemiology and its Center for Drug Evaluation and Research, that well-designed RWE studies can produce reliable causal estimates under the right conditions.
The agency's Sentinel System is the institutional cornerstone of this shift. Built on claims and administrative data from more than 100 million patients across commercial insurers, Medicare, and Medicaid, Sentinel runs continuous queries looking for safety signals in approved products. When a new adverse event hypothesis emerges, epidemiologists can test it against the Sentinel database in weeks rather than the years a prospective trial would require. The rofecoxib episode, in which the COX-2 inhibitor Vioxx was withdrawn in 2004 after an estimated 28,000 excess cardiac events over five years on the market, as documented in David Graham and colleagues' landmark 2005 Lancet analysis, is now frequently cited as the cautionary case for what active RWE surveillance could have prevented.
The limitations of Sentinel are real and openly acknowledged by FDA. Claims data captures billing events, not clinical reality. A claim for a diagnosis does not confirm the diagnosis was accurate, that the drug was taken as prescribed, or that an outcome was correctly coded. Those data quality challenges are precisely where artificial intelligence has found its most commercially significant role in the RWE ecosystem.
The Platforms Turning Noise Into Evidence
In 2018, Roche acquired Flatiron Health for $1.9 billion, a price that reflected less the company's revenue than its data asset: structured clinical information on 2.4 million cancer patients drawn from 300 oncology practices across the United States. What makes Flatiron's dataset distinct from a simple EHR export is the layer of AI processing between the raw clinical notes and the structured data product. Oncology records are notoriously heterogeneous. A pathology report describing tumour response in a patient's third line of therapy is written in natural language, embedded in a PDF, and formatted differently at every institution. Flatiron's natural language processing pipeline extracts from those notes the specific clinical events that regulatory submissions require: line of therapy, response by RECIST criteria, date of progression, reason for discontinuation.
The accuracy of that extraction is not a solved problem. A 2021 study published in JAMA found that EHR-extracted laboratory values from five major US health systems had a 12 percent error rate when validated against chart abstraction as the gold standard. Flatiron and its competitors invest heavily in curation processes that combine NLP outputs with targeted human review, but sponsors submitting RWE to FDA are required to document both the error rate and the analytical impact of misclassification on their primary endpoint estimate.
TriNetX takes a different architectural approach, functioning as a federated network that allows sponsors to query de-identified patient populations across 150 participating health systems covering approximately 250 million patients without those health systems moving data off their own servers. This privacy-preserving model, which aligns with the federated learning principles described in more detail in our coverage of federated learning in healthcare, allows clinical trial feasibility analyses and retrospective cohort studies at a scale that would otherwise require complex data transfer agreements.
Aetion, a company spun out of Harvard Medical School with methodological roots in Miguel Hernan and James Robins's target trial emulation framework, has positioned itself explicitly in the regulatory submission market. Its Aetion Evidence Platform is built around the principle that every RWE analysis should be designed as if it were emulating a specific target trial, with pre-specified eligibility criteria, treatment strategies, follow-up periods, and outcome definitions. That discipline, Hernan and Robins argued in their foundational 2016 paper in the American Journal of Epidemiology, is what separates causal inference from descriptive epidemiology. IQVIA, the largest of the group, links claims and EHR data for roughly 400 million patients globally, offering regulatory, commercial, and clinical research services from a single integrated infrastructure.
The competitive dynamics among these platforms are shaping what kinds of AI methods become standard in regulatory submissions. Propensity score matching and inverse probability of treatment weighting have become near-universal tools for attempting to approximate the randomisation that observational data lacks. Both methods estimate the probability that a given patient received a particular treatment, then weight or match patients so that treated and untreated groups look similar on measured confounders. The critical qualifier is measured: unmeasured confounding, where sicker patients systematically receive different treatments for reasons not captured in the data, remains the central methodological vulnerability of all observational RWE. For a broader view of how patient-level data analytics is being applied at the population scale, see our reporting on AI in population health management.
Synthetic Control Arms: Evidence Without Placebo Patients
Among the most consequential applications of AI in RWE is the construction of synthetic control arms for rare disease drug development. The concept is straightforward: rather than randomising patients with a severe, often fatal condition to placebo, a sponsor constructs a comparator group from historical patients in natural history registries whose baseline characteristics are matched to the trial participants using AI-assisted propensity methods.
The FDA's 2020 approval of risdiplam (Evrysdi) for spinal muscular atrophy provides the clearest case study. SMA Type 1 causes severe muscle weakness and respiratory failure in infants, and the prospect of assigning any infant to placebo in a clinical trial was considered ethically untenable by the clinical community. Genentech and Roche submitted a single-arm trial for risdiplam in which the comparator was drawn from the SMArtCARE registry and the Pediatric Neuromuscular Clinical Research (PNCR) network. AI-assisted matching algorithms identified historical patients whose baseline motor function scores, disease duration, and support measures most closely resembled the trial population. FDA accepted the design, granting accelerated approval on the basis of a motor function endpoint and subsequent conversion to full approval as post-market evidence accumulated.
The approach is not without critics. Thomas Fleming, a biostatistician at the University of Washington and frequent FDA advisory committee member, has argued in published commentaries that historical controls are systematically biased by secular trends in supportive care. A child with SMA receiving optimal nutritional and respiratory support in 2019 may have substantially better outcomes than a historical control from 2010, regardless of the experimental drug's effect. The FDA's rare disease guidance acknowledges this limitation explicitly, requiring sponsors to conduct sensitivity analyses that stress-test the matching assumptions and to present any secular trend data available from the registry. The tension between the ethical imperative to minimise placebo exposure and the scientific imperative to minimise bias is precisely the kind of problem that does not resolve neatly.
The Causal Inference Debate Inside FDA Submissions
Behind the regulatory acceptance of RWE lies a methodological debate that has occupied academic epidemiologists and statisticians for decades and is now playing out in consequential regulatory decisions. Judea Pearl's do-calculus framework for causal inference, which formalises the conditions under which observational data can support causal conclusions using directed acyclic graphs, has become standard in academic RWE work. Traditional epidemiological methods, grounded in the potential outcomes framework of Donald Rubin, approach the same problem through a different mathematical lens but converge on similar practical requirements: transparent specification of assumptions, sensitivity analyses, and honest acknowledgment of what cannot be controlled.
The practical consequence for regulatory submissions is a set of analytical requirements that sponsors sometimes find commercially inconvenient. Pre-specification of the analysis plan before any data access is now FDA's firm expectation for most RWE submissions. Sponsors who query their dataset, find that their primary analysis did not produce the desired result, and then modify the analytical approach have introduced selection bias that FDA reviewers are trained to identify through submission metadata and audit trails. Aetion and several competitor platforms have built pre-specification workflow tools directly into their software, generating timestamped, cryptographically signed analysis protocols that can be submitted to FDA alongside the final results as provenance documentation.
The intersection of RWE analytical methods and the broader AI-driven analytics infrastructure in healthcare is explored in the QuanMed AI and Healthcare Guide, which covers how these tools fit into clinical decision-making more broadly. For the specific context of how RWE-derived insights translate into hospital operational decisions, the reporting on predictive analytics for hospital readmissions illustrates the downstream applications of the same observational data infrastructure.
Post-Market Surveillance: The Ongoing Experiment
The drug approval decision is not the end of FDA's interest in RWE. Post-market surveillance, the monitoring of approved drugs for unexpected adverse events at scale, is where RWE has the longest institutional history and arguably the most developed AI infrastructure. FDA's Sentinel System, operating since 2008 under authority granted by the FDA Amendments Act, continuously analyses claims and administrative data from more than 100 million covered lives. When a new safety hypothesis arises, Sentinel's distributed query network can run standardised analyses across participating data partners in days, with each partner running the code locally and returning only aggregate results, preserving patient privacy without sacrificing statistical power.
The Vioxx case has become the system's founding cautionary tale. Rofecoxib, approved in 1999, was withdrawn in September 2004 after the APPROVE trial demonstrated a doubling of cardiovascular risk. Graham and colleagues' 2005 Lancet analysis, drawing on Kaiser Permanente and Medicaid data, estimated that between 88,000 and 139,000 Americans had experienced excess cardiovascular events attributable to the drug over its five years on the market, with 28,000 of those cases likely fatal. A Sentinel-style active surveillance system operating at scale in 1999 would not have definitively prevented those outcomes, FDA officials have acknowledged, but it might have generated the hypothesis that led to earlier investigation. That acknowledgment is the political foundation on which Sentinel's budget has been consistently defended in congressional appropriations cycles.
The connection between RWE surveillance and the data provenance questions raised in discussions of blockchain for EHR data integrity is more than theoretical. Sentinel's distributed architecture relies on each data partner maintaining their own records with sufficient fidelity that queries produce consistent results across sites. Inconsistencies in how different insurers code the same clinical event, a problem that FDA's Common Data Model partially addresses by standardising terminology, remain a source of heterogeneity in multi-site analyses that AI methods can identify but cannot fully eliminate.
Wearables, Decentralised Trials, and the Next Wave
The definition of real-world data is expanding. FDA's 2023 guidance on digital health technologies in drug development formally acknowledged that patient-generated data from wearables, including continuous glucose monitors, implantable cardiac monitors, consumer-grade accelerometers, and digital spirometers, can be incorporated into RWE frameworks under the right validation conditions. The guidance requires sponsors to demonstrate that a digital measure captures a clinically meaningful outcome, that the device used to generate it has been validated against an appropriate gold standard, and that data collection is consistent enough across participants to support cross-sectional comparisons.
Decentralised clinical trials, which reached critical mass during the COVID-19 pandemic when in-person visits became impossible, represent a structural shift that is blurring the boundary between RCT and RWE. A trial that enrolls participants nationally, collects wearable data continuously in their homes, and relies on electronic patient-reported outcomes rather than clinic visits generates data with the prospective design of an RCT and the ecological validity of real-world observation. Pfizer's HERO-HF trial of ivabradine in heart failure, which used a decentralised design with remote monitoring, is one of the more cited early examples of this hybrid approach. The AI infrastructure required to manage, clean, and analyse continuous streams of wearable data at trial scale is substantially more complex than anything the clinical trial industry was operating five years ago, and several of the RWE platform companies are now actively expanding into this adjacent market.
The connection to broader drug discovery pipelines is direct. As described in our reporting on AI in drug discovery, the same patient-level data that generates RWE for regulatory purposes is increasingly being used earlier in the development process to identify biomarker-defined patient populations, predict which subgroups are most likely to respond, and prioritise compounds for advancement. The feedback loop between discovery, development, and post-market evidence is tightening, and AI is the connective tissue making that acceleration possible. The challenge for regulators, payers, and the patient advocacy community is ensuring that speed does not outpace the methodological rigour that makes the resulting evidence trustworthy.
Key Sources
- FDA. Framework for FDA's Real-World Evidence Program. FDA.gov, 2018. -- Legislative mandate under the 21st Century Cures Act; foundational regulatory framework for RWE acceptance.
- FDA. Submitting Documents Using Real-World Data and Real-World Evidence to FDA for Drugs and Biologics. Guidance for Industry. 2023. -- Operational guidance including submission count data; basis for the 86 submissions figure cited in text.
- Graham DJ, Campen D, Hui R, et al. Risk of acute myocardial infarction and sudden cardiac death in patients treated with cyclo-oxygenase 2 selective and non-selective non-steroidal anti-inflammatory drugs. Lancet. 2005;365(9458):475-481. -- Rofecoxib post-market RWE analysis estimating 28,000 excess cardiac events; foundational cautionary case for active surveillance.
- Corrigan-Curay J, Sacks L, Woodcock J. Real-world evidence and real-world data for evaluating drug safety and effectiveness. JAMA. 2018;320(9):867-868. -- FDA leadership overview of the regulatory and scientific standards for RWE; key reference for data quality requirements.
- Hernan MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. Am J Epidemiol. 2016;183(8):758-764. -- Target trial emulation framework; methodological foundation for the analytical discipline required in FDA RWE submissions.