← Research

MIT Critical Data · Ongoing collaboration

Reliability and Failure Analysis for Multimodal Clinical AI

Do model failures concentrate in the same patients or encounters across structured clinical data and chest radiographs? This project is building the linked cohort and reliability analyses needed to answer that question without treating model disagreement as evidence of clinical harm.

Clinical AICalibrationFailure analysisSubgroup auditing

Research question

Average performance can conceal repeated failure in a smaller group of patients.

Clinical AI is often evaluated one model and one task at a time. That design can miss a more consequential pattern: errors that repeatedly concentrate in the same patients, encounters, examinations, or care settings.

The project links reliability signals from emergency-department prediction models, chest-radiograph models, care phenotypes, and missingness patterns. The analysis distinguishes subject-level overlap from encounter-level and time-windowed overlap so that later inpatient information is not treated as though it were available at the intended prediction time.

Structured clinical models

Encounter-level predictions

Chest-radiograph models

Image-level disagreement

Patient and encounter reliability profile

Continuous scores, overlap, discordance, uncertainty

Care phenotypes

High-acuity clinical context

Missingness patterns

Documentation and measurement context

Cross-modal audit framework. Structured clinical predictions, radiograph-level disagreement, care patterns, and missingness are linked at the patient and encounter levels. Current work prioritizes reproducible linkage and sensitivity analysis before any shared reliability phenotype is defined.

Study design

Temporal alignment is part of the scientific question, not a preprocessing detail.

A patient may have several encounters and several radiographs, and clinical status changes over time. I am defining explicit linkage windows, the unit of analysis, and the information that would have been available at each prediction point.

These choices determine whether an apparent cross-modal pattern reflects persistent patient-level vulnerability, repeated measurements, a single episode of care, or information leakage.

PatientStable subject identifier
ED encounterPrediction time
RadiographOne or more image times
Admission windowCare context changes
Phenotype intervalTime-bounded interpretation

My contribution

Reliability checks belong in the primary analysis.

I am developing the linked cohort, documenting temporal assumptions, comparing continuous reliability scores, and testing whether conclusions persist across thresholds, task subsets, consensus rules, and aggregation choices. Calibration, bootstrap uncertainty, and subgroup auditing are specified as part of the analysis rather than added after a favorable result is selected.

CalibrationCompare reliability on a common scale

Structured clinical models

Confidence versus observed error

Chest radiographs

Task-level and image-level calibration

Cross-modal interpretation

Compare reliability on a common scale

Threshold sensitivityTest whether overlap depends on one cutoff

Structured clinical models

Risk-score operating points

Chest radiographs

Consensus quorum and task subset

Cross-modal interpretation

Test whether overlap depends on one cutoff

Linkage sensitivitySubject, encounter, and time-windowed overlap

Structured clinical models

Encounter definitions

Chest radiographs

Repeated-image aggregation

Cross-modal interpretation

Subject, encounter, and time-windowed overlap

UncertaintyStability of overlap and discordance

Structured clinical models

Bootstrap intervals

Chest radiographs

Score and ensemble variability

Cross-modal interpretation

Stability of overlap and discordance

SubgroupsDescriptive disparity without causal overreach

Structured clinical models

Patient and care strata

Chest radiographs

Exam and patient strata

Cross-modal interpretation

Descriptive disparity without causal overreach

Reliability checks across modalities. Primary conclusions should remain stable under reasonable changes to thresholds, task definitions, linkage windows, and aggregation rules.

Interpretation boundary

Model-label disagreement is an audit signal, not a patient outcome.

The immediate goal is to determine whether cross-modal vulnerability is stable, interpretable, and reproducible. Claims about harm, intervention, or deployment require a completed linked cohort, prespecified score definitions, a sensitivity plan, and downstream clinical outcomes that support those claims.

View MIT role