Paper library

Research

Every paper that appears here is approved by Adam personally. No auto-indexing. No content-farm preprints. Each entry is a deliberately-posted preprint with a documented protocol timestamp, shared for pre-submission peer commentary.

This page documents Adam's research in clinical AI. The methodology applied here generalizes to other quantitative domains, see the Methodology page for cross-domain framing, or the Services page for engagement options outside clinical AI.

Preprint sepsis falsification · 286,510 encounters · Protocol drafted 2026-01-30

A pre-registered falsification study of sepsis prediction benchmarks, and a published null result

Key finding, Across 286,510 ICU encounters in the MIMIC-IV and eICU-CRD public datasets (eICU-CRD covers ~209 US hospitals), reported benchmark accuracy depends on which sepsis definition generates the labels: performance holds under feature ablation on billing-code-derived labels but degrades on Sepsis-3 clinical criteria, and the gap replicates on a held-out cohort. The pooled clinical claim did not survive the pre-registered falsification protocol, and the null result was published rather than shelved. The implication is deliberately narrow: benchmark "accuracy" numbers mix administrative and biological signal, and evaluations should separate the two before clinical claims are made. This study does not claim any deployed tool is trained on billing codes.

[OSF support review pending, see verification mirror]

Read abstract
Background. Published sepsis AI models routinely report AUC > 0.90 on ICU benchmarks. We pre-registered a falsification protocol to test whether these benchmarks measure biological sepsis or administrative artifacts. Methods. 286,510 ICU encounters from MIMIC-IV and eICU-CRD. Labels were regenerated using three independent sepsis definitions (Sepsis-3 clinical criteria, biomarker thresholds, and billing-code derived). Model performance was compared across label regimes. Results. Models trained on billing-code labels retain AUC > 0.85 even when key biological features are removed, while models trained on Sepsis-3 clinical criteria degrade to AUC 0.68 under the same ablation. The gap replicates on eICU. Performance on the billing benchmark is largely explained by care-process features observed after sepsis onset. Conclusion. Benchmark performance reflects hospital coding practice as well as biology, and the pooled clinical claim did not survive falsification. Model evaluation should separate administrative from biological signal before clinical claims are made.
Preprint community-hospital workload · 136,864 encounters · Protocol drafted 2026-02-12

Community-hospital sepsis workload: a 136,864-encounter study of institution-type bias in ICU AI

Key finding, Sepsis AI trained on academic medical-center data underperforms by 11–19 AUC points when deployed to community hospitals, and the performance gap does not close with standard fine-tuning. Institution-type is a structural confound in most public sepsis datasets.

[OSF support review pending, see verification mirror]

Read abstract
Background. Public ICU datasets are dominated by academic medical centers, but the majority of US ICU admissions happen in community hospitals. Methods. 136,864 encounters from eICU-CRD stratified by hospital type. We trained sepsis models on AMC-only cohorts and evaluated on community cohorts, and vice versa. Results. Cross-institution-type transfer loses 11–19 AUC points. The gap persists after standard fine-tuning and is partially attributable to differences in nursing documentation cadence, lab-test ordering patterns, and antibiotic-administration timing, all of which are inputs to common sepsis models. Conclusion. Institution-type is a first-class covariate in ICU AI evaluation. FDA submissions relying on AMC-trained models should include community-hospital validation cohorts as a standard practice.
Preprint ICU mortality miscalibration · 201,905 encounters · Protocol drafted 2026-03-10

ICU mortality miscalibration: published estimates underestimate elderly risk by 66–168% across three conditions

Key finding, Across the MIMIC-IV (Beth Israel Deaconess) and eICU-CRD (208 US hospitals) public datasets, 201,905 total ICU encounters, published ICU mortality estimates for elderly patients are underestimates by 66.3% (AF), 131.7% (diabetes), and 168.4% (MI) simultaneously. The miscalibration replicates across two independent datasets and six condition × age strata.

[OSF support review pending, see verification mirror]

Read abstract
Background. Clinical risk stratification tools rely on published ICU mortality estimates that are often decades old, cohort-specific, or derived from meta-analyses with limited subgroup resolution. Methods. Pre-registered observational study, 201,905 ICU encounters from MIMIC-IV and eICU-CRD. Six primary strata: Diabetes, MI, AF, Seizure, PE × age > 70; COPD × age < 50. Observed mortality computed and compared to consensus-published estimates. Results. Five of six strata show divergences exceeding 60%, cross-validated on the independent dataset. The seizure × elderly stratum shows a +510% divergence, the largest literature gap in the series. COPD × young adults shows the inverse (−63.6%, suggesting published estimates are inflated by selection bias). Conclusion. Bedside risk stratification using literature-derived mortality estimates is systematically miscalibrated for elderly ICU patients. Clinical AI trained on or calibrated to these estimates inherits the same miscalibration.
Why only three papers? Because every paper on this page was either written by Adam or explicitly approved by him before publication. We do not aggregate the field's output. This is a library of what we stand behind. The larger pipeline of engine-generated findings, the ones not yet written up as peer-reviewed manuscripts, lives on the Findings dashboard.
Research partnerships

Academic and grant collaborations

We co-author, pre-register before analysis, publish our code and evidence hashes, and publish nulls. Independent verification of clinical AI is an active federal funding area, and we collaborate on grant-funded work with academic groups, health systems, and public-dataset maintainers. If you have a cohort, a benchmark, or a claim worth falsifying, talk to Adam.

Propose a collaboration →