The Evidence

The literature already told us.

Our thesis is the consistent finding of a decade of peer-reviewed work: clinical AI fails on generalization, deployment, attention, and adoption, not on benchmark accuracy. The literature on the clinical workday also names the workflows where that failure costs the most. A working reading room, organized the way we read it.

I
Generalization

Models don't travel.

Performance measured at home says little about performance anywhere else. External validation is where clinical AI meets reality, and where most of it stumbles.

  1. 01

    Yu, Mohajer & Eng · Radiology: Artificial Intelligence · 2022

    External validation of deep learning algorithms for radiologic diagnosis.

    A systematic review of radiology AI studies with external validation: in 81% of cases, performance dropped when the model left the institution that built it. Degradation isn't the exception; it's the base rate.

    81% degraded outside home institution
  2. 02

    Zech et al. · PLOS Medicine · 2018

    Variable generalization of a deep learning model for pneumonia detection in chest radiographs.

    The models didn't just learn pneumonia — they learned which hospital the X-ray came from, and leaned on it. What a model learns in one system may not be medicine; it may be the system itself.

  3. 03

    Wu et al. · Nature Medicine · 2021

    How medical AI devices are evaluated.

    An audit of FDA-cleared AI devices: most were evaluated at only a small number of sites, and prospective evaluation was rare. Regulatory clearance is not evidence of local fit. The burden of proof arrives with the deployment.

  4. 04

    Finlayson et al. · New England Journal of Medicine · 2021

    The clinician and dataset shift in artificial intelligence.

    Why models decay in place: populations drift, practice patterns change, coding conventions move. Maintenance is a clinical function, not an IT afterthought, and someone has to own it.

II
Deployment

The setting is part of the system.

The distance between a validated model and a working clinical tool is measured in workflows, governance, and trust. It's where the strongest models fail.

  1. 05

    Wong et al. · JAMA Internal Medicine · 2021

    External validation of a widely implemented proprietary sepsis prediction model.

    A sepsis model in use at hundreds of hospitals, evaluated independently: it missed two-thirds of the sepsis cases it was meant to catch, while 88% of its alerts were false alarms. Scale is not validation.

    88% false alarms in external validation
  2. 06

    Beede et al. · CHI Conference on Human Factors in Computing Systems · 2020

    A human-centered evaluation of a deep learning system for diabetic retinopathy screening.

    A model with excellent laboratory accuracy stumbled in real clinics — on lighting, connectivity, patient flow, and nurse workflows. The deployment environment is part of the system, whether you design for it or not.

  3. 07

    Sendak et al. · JMIR Medical Informatics · 2020

    Real-world integration of a sepsis deep learning technology into routine clinical care.

    A rare success story, and an honest accounting of what it took: years of socio-technical work on governance, workflow design, and clinician trust wrapped around the model. The model was the smallest part.

  4. 08

    Kelly et al. · BMC Medicine · 2019

    Key challenges for delivering clinical impact with artificial intelligence.

    The canonical map of the gap between retrospective accuracy and clinical impact: prospective evaluation, workflow integration, human factors, and monitoring. Most published models never cross it.

III
Attention

Attention is finite.

Every alert spends clinician attention, a budget that is already overdrawn. How, when, and to whom insight is delivered isn't packaging. It's the intervention.

  1. 09

    Nanji et al. · Journal of the American Medical Informatics Association · 2018

    Medication-related clinical decision support alert overrides in inpatients.

    Roughly three-quarters of medication alerts were overridden, and most overrides were judged appropriate. When the default clinician response to a system is "dismiss," the system is the problem.

    ~73% of alerts overridden
  2. 10

    Ancker et al. · BMC Medical Informatics and Decision Making · 2017

    Effects of workload and repeated alerts on alert fatigue in a clinical decision support system.

    Alert fatigue is dose-dependent: as volume and repetition rise, response falls, regardless of the alert's merit. Attention is a shared resource, and every tool that ignores that degrades all the others.

  3. 11

    Roshanov et al. · BMJ · 2013

    Features of effective computerised clinical decision support systems: meta-regression of 162 randomised trials.

    Across 162 trials, what separated systems that changed care from those that didn't was largely delivery: how advice was presented, routed, and integrated. The same insight, delivered differently, is a different intervention.

    162 randomised trials
IV
The Work

The work that eats the day.

Attention doesn't vanish; it gets spent. The literature on the clinical workday names the workflows that consume it: the inbox, the care gap, the event report, the note. These are the workflows we build in.

  1. 12

    Sinsky et al. · Annals of Internal Medicine · 2016

    Allocation of physician time in ambulatory practice: a time and motion study in 4 specialties.

    For every hour of direct face time with patients, physicians spent nearly two additional hours on the EHR and desk work. The workday's center of gravity isn't the exam room; it's the documentation and inbox work wrapped around it.

    ≈2 hours of desk work per hour of care
  2. 13

    Murphy et al. · JAMA Internal Medicine · 2016

    The burden of inbox notifications in commercial electronic health records.

    Primary care physicians received on the order of 77 inbox notifications a day, each one a small clinical decision delivered without ranking, routing, or context. The inbox is a clinical workflow that has never been designed as one.

    ~77 notifications a day
  3. 14

    McGlynn et al. · New England Journal of Medicine · 2003

    The quality of health care delivered to adults in the United States.

    The landmark accounting of the care gap: adults received barely half of recommended care. Two decades later, closing a gap is still chart-level manual work: finding it, confirming it, acting on it. That is why the gap persists.

    ≈55% of recommended care delivered
  4. 15

    Classen et al. · Health Affairs · 2011

    "Global trigger tool" shows that adverse events in hospitals may be ten times greater than previously measured.

    Systematic chart review found adverse events in roughly a third of admissions, around ten times more than voluntary reporting surfaced. Event review as practiced is a sampling of a sampling; most of the signal is never reviewed at all.

    ~10× more events than reporting catches
V
Adoption

Adoption is the outcome.

Use is accelerating faster than evaluation. The field's own institutions, from journals to regulators to global bodies, keep pointing at the same missing layer.

  1. 16

    Health Affairs · AHA national survey data · 2025

    Predictive AI use in US hospitals.

    65% of US hospitals report using predictive AI in the EHR, while local evaluation for accuracy and bias lags far behind adoption. Deployment has outrun validation at national scale.

    65% of hospitals use predictive AI
  2. 17

    van de Sande et al. · Intensive Care Medicine · 2021

    Moving from bytes to bedside: a systematic review of machine learning in the ICU.

    Of the flood of published ICU models, only a handful ever reached prospective clinical evaluation — and almost none demonstrated impact on care. The pipeline leaks between the paper and the bedside.

  3. 18

    U.S. Food & Drug Administration · 2021

    AI/ML-based Software as a Medical Device: action plan.

    The regulator's own convergence on our thesis: the open questions are real-world performance monitoring and transparency to users — how models behave after deployment, not just before it.

  4. 19

    World Health Organization · 2021

    Ethics and governance of artificial intelligence for health.

    Global guidance placing human oversight and contextual validation at the center of safe clinical AI: tools must be evaluated in the settings where they'll be used, by the people who'll use them.

A note on this page: annotations in bold are our editorial reading. The findings and figures belong to the cited authors, and citations are provided for locating the primary sources. This is a working list, not a complete bibliography; we keep it to work that shaped how we practice. Disagree with a reading, or think something's missing? We'd like to hear it: hello@pendentiv.ai.

This is the problem we organized around.

The literature names the gap and the workflows where it costs the most. The approach is how we close it, layer by layer, where care happens.

See the approach hello@pendentiv.ai