3 of 1,357 Cleared AI Devices Have Public Outcome Trials. The Study Has No Non-AI Comparator.
A Review article in PLOS Digital Health, published 19 August 2026, which the authors describe as a regulatory evidence census and structured evidence-mapping analysis rather than a PRISMA-style meta-analysis. Abulibdeh and colleagues at Toronto, MIT, Johns Hopkins, Mbarara and Bergen took all 1,357 AI and machine-learning enabled devices cleared or approved by the FDA through 5 December 2025, using the FDA device database and the ACR Data Science Institute catalogue, then scraped ClinicalTrials.gov identifiers from 510(k) summary pages and linked those registrations to PubMed publications.
- The attrition is steep at every stage. Of 1,357 cleared devices, 34 (2.5%) were linked to registered prospective trials, 12 (0.9%) posted results, 12 (0.9%) reached peer-reviewed publication, and 3 (0.2%) evaluated patient-centred outcomes such as mortality, morbidity or readmission.
- The 34 trials that exist are small, domestic and industry-run. Roughly 73.5% enrolled fewer than 500 participants and a quarter fewer than 100. 68% ran only in the United States. 32 of 34 (94%) were industry-led. Nine reported any subgroup analysis, with race or ethnicity in 3 and language in none.
- Specialty concentration is extreme and inversely related to evidence. Radiology accounts for 78% of cleared devices (1,059 of 1,357) and has prospective trials for fewer than 1% of them. Cardiovascular and neurology sit at 9.5% and 9.7%. Anaesthesiology has 22 cleared devices and no registered prospective trials at all.
- Almost nothing tests treatment guidance. Among the 34 trials, 59% addressed diagnostic applications and 21% screening. Three studies, 9%, addressed therapeutic or treatment guidance.
The mechanism the authors identify is the predicate chain. A 510(k) submission demonstrates substantial equivalence to an existing device rather than independent clinical effectiveness, so a device cleared against a predicate that was itself never prospectively validated inherits that gap and passes it on. That is a structural property of the pathway rather than a failure of any individual clearance, and it explains why the shortfall concentrates in radiology, where the predicate population is largest and oldest. It also means the fix the authors propose, a staged framework requiring retrospective validation before clearance and prospective outcome trials at n of 500 and then 2,000, is a change to the pathway rather than to any company’s behaviour.
