Published Evidence

Healthcare AI Evidence Library

Published studies on healthcare AI, grouped by use case and sorted so trials and reviews come first. Every entry carries its journal, year and identifier so you can cite it directly.

You are in a meeting and someone asks whether there is any real evidence this works. Not a case study the vendor wrote. Actual published research.

This is where you get it, in a couple of minutes, with the journal and the identifier already attached so the citation survives being pasted into a board paper.

What is in this library

192 Studies listed
11 Trials and reviews
8 Subjects covered
6,428 Papers searched

Start With These

Trials, meta-analyses and systematic reviews carry more weight than a single site report, so they go first.

  1. Meta-analysis Intensive care medicine, 2020

    Machine learning for the prediction of sepsis: a systematic review and meta-analysis of diagnostic test accuracy.

    Filed under AI Sepsis and Deterioration Alerts ... cited 465 times

  2. Meta-analysis Gastrointestinal endoscopy, 2021

    Performance of artificial intelligence in colonoscopy for adenoma and polyp detection: a systematic review and meta-analysis.

    Filed under AI Polyp Detection in Colonoscopy ... cited 344 times

  3. Meta-analysis Endoscopy, 2021

    Artificial intelligence for polyp detection during colonoscopy: a systematic review and meta-analysis.

    Filed under AI Polyp Detection in Colonoscopy ... cited 166 times

  4. Meta-analysis Annals of internal medicine, 2023

    Real-Time Computer-Aided Detection of Colorectal Neoplasia During Colonoscopy : A Systematic Review and Meta-analysis.

    Filed under AI Polyp Detection in Colonoscopy ... cited 163 times

  5. Meta-analysis Computer methods and programs in biomedicine, 2019

    Prediction of sepsis patients using machine learning approach: A meta-analysis.

    Filed under AI Sepsis and Deterioration Alerts ... cited 144 times

  6. Systematic review Journal of neurointerventional surgery, 2020

    Artificial intelligence to diagnose ischemic stroke and identify large vessel occlusions: a systematic review.

    Filed under AI Stroke Triage ... cited 184 times

  7. Systematic review Frontiers in medicine, 2021

    Early Prediction of Sepsis in the ICU Using Machine Learning: A Systematic Review.

    Filed under AI Sepsis and Deterioration Alerts ... cited 103 times

  8. Systematic review NPJ digital medicine, 2025

    Systematic review of cost effectiveness and budget impact of artificial intelligence in healthcare.

    Filed under Cost and Return on Healthcare AI ... cited 70 times

  9. Systematic review World neurosurgery, 2022

    Artificial Intelligence for Large-Vessel Occlusion Stroke: A Systematic Review.

    Filed under AI Stroke Triage ... cited 55 times

  10. Systematic review Healthcare (Basel, Switzerland), 2025

    The Impact of AI Scribes on Streamlining Clinical Documentation: A Systematic Review.

    Filed under Ambient AI Documentation ... cited 51 times

Browse by Subject

Why a Citation Beats a Case Study

A vendor case study is written by the vendor. That does not make it false, and it does make it selective. Nobody publishes the site where it did not work.

Peer-reviewed work is not perfect either. Plenty of it is small, funded by the company, or run in a setting nothing like yours. But it has been through people whose job was to argue with it, and you can read the method instead of the highlight reel.

So the useful move in a procurement conversation is not to reject the case study. It is to put a real study next to it and ask why they differ.

Three Questions to Bring to Any Study Here

Reading one study properly beats skimming ten. These are the three things worth checking before a result changes your mind:

  1. Who was in it? A model that performed beautifully in one country's screening program, on one manufacturer's scanners, is telling you about that program. Look for the population and the equipment before you look at the accuracy figure.
  2. Compared with what? Better than nothing is easy. Better than your current process, staffed the way it is today, is the comparison that matters. A study that tested software against unassisted readers is answering a different question from one that tested it against double reading.
  3. What was actually measured? Detection rate, reading time, and patient outcome are three different claims, and a paper that improves the first says nothing automatic about the third.

How These Lists Are Built

Each subject is a saved search against Europe PMC, which indexes PubMed, MEDLINE, PMC and the main preprint servers. The search is pinned to the title and abstract fields rather than run loose across the full text, and the results are then ordered by citation weight.

That pinning is worth mentioning because the first version of this library did not have it. A loose search sorted by citations returned the most famous AI papers that merely mentioned the words somewhere, so a general image-segmentation paper turned up at the top of the colonoscopy list. It was correctly matched and completely useless. Tightening the fields fixed it.

Nothing on a study row is written here. Title, authors, journal, year, identifiers, citation count and the abstract all come straight from the source record, which is the only way you can trust a citation you did not look up yourself.

Study records come from Europe PMC, which indexes PubMed, MEDLINE, PMC and preprint servers. Titles, journals, years, identifiers, citation counts and abstracts are reproduced from the source record and are not rewritten here. A study being listed is not an endorsement of its conclusion, and citation count measures attention rather than quality. Read the paper before you cite it.