Short answer, before the detail: for most healthcare AI, nobody has done the sums. Across the whole published literature, formal economic evaluations of healthcare AI number in the hundreds while the research base runs to tens of thousands of papers. The gap is not evidence that AI does not pay. It is evidence that when a vendor says the return is proven, they are almost always pointing at a model somebody built, not a result somebody measured.
That matters because the modeled number and your number are different things. A published cost-effectiveness study is a careful argument about a specific health system, with assumptions its authors chose and wrote down. Borrowing its conclusion and dropping it into your business case throws away the part that made it credible.
Where the Economic Evidence Actually Sits
The use cases that have been costed properly are not the ones with the loudest marketing. They cluster around screening programs, where the math is tractable: a defined population, a repeatable test, a known cost per case found, and a health system that already tracks all three.
| Use case | Papers found | The lever you can move | Why the number is softer than it looks |
|---|---|---|---|
| Imaging triage and workflow | 442 | Time from scan to treatment, and transfers avoided | This count is inflated: imaging is a broad word and catches economic papers about scanners rather than software |
| Diabetic retinopathy screening | 65 | People screened who were not being screened before | Savings depend on your local specialist cost, which varies enormously |
| Colonoscopy polyp detection | 55 | Adenomas found per procedure, and the surveillance that follows | Extra small polyps add removal and follow-up cost immediately, benefit arrives years later |
| Mammography screening | 43 | Reader hours saved where double reading is the current standard | Most modeling assumes a European double-reading program, so it does not transfer to single-read settings |
| Revenue cycle and admin | 39 | Denial rate, days in accounts receivable, staff hours per claim | Almost all of this evidence is vendor-reported rather than peer-reviewed |
| Pathology | 37 | Turnaround time per case and avoided send-out testing | The scanning program usually costs more than the software and is often left out |
| Sepsis and deterioration alerts | 24 | Avoided intensive care admissions and length of stay | Attributing an avoided deterioration to the alert is genuinely hard, so estimates swing widely |
| Ambient documentation | 11 | After-hours minutes in the record, and clinician retention | The thinnest economic literature of any widely bought healthcare AI, despite the spend |
Read that table as a map of where somebody has bothered, not as a ranking of which tools pay best. A count is a screening figure: it tells you papers exist where the right words appear together. Some of those papers are rigorous evaluations and some mention cost in a single sentence. The refreshed count and the full study list live in the cost and return evidence library.
Four Numbers That Decide Every Healthcare AI Business Case
Most failed business cases fail in the same four places, and none of them are about the software. Get these right and the rest is arithmetic.
- What the current process costs today. Not the estimate, the measurement. If you cannot say what a radiologist minute, a denied claim, or an after-hours documentation hour costs you now, every figure downstream is decoration. This is the single most common gap and it is entirely within your control.
- The volume the lever applies to. A tool that saves ninety seconds per study is transformative at two hundred thousand studies a year and irrelevant at four thousand. Vendors quote the per-unit saving because it sounds the same at every size. Multiply it yourself.
- The downstream load it creates. Better detection means more follow-up. More alerts means more responses. More flagged patients means more clinic slots. This cost is real, it lands on a different budget from the software, and it is missing from almost every vendor model you will be shown.
- Who actually captures the saving. Time returned to a clinician is money only if a decision gets made about what fills it. If nothing is decided, the hour is absorbed and the return is zero, and that is a management failure rather than a technology one.
Ask for the Denominator
There is one question that does more work than any other in a vendor meeting. Ask what the number is divided by.
A thirty percent improvement in reading time is a different purchase depending on whether the base was the whole study list or the eleven complex cases they measured. A cost saving per patient screened means nothing until you know how many were screened and how many were eligible. A return within twelve months assumes a go-live date that assumes an integration timeline that assumes your team has capacity they probably do not have.
None of that is a trick question. Good vendors answer it immediately because they have the figures. The ones who move to a different slide have told you something useful too.
Start With Regulatory Status, Then Evidence, Then Money
The order matters. A product that is not cleared for the use you have in mind cannot have a return in that use, so check the FDA-cleared AI directory first and confirm the clearance covers what you were pitched. Then read what has actually been published about the use case in the evidence library. Only then build the number, using your own baseline rather than theirs.
Done in that order, a business case takes an afternoon and survives a finance review. Done in reverse, it takes a quarter and falls over on the first hard question.