← Research

MIT Critical Data · Ongoing collaboration

Reliability and Failure Analysis for Multimodal Clinical AI

Can the same patients or encounters be difficult for AI systems across structured clinical data and chest radiographs? This project builds the linked cohort and reliability analyses needed to answer that question without turning model disagreement into an unsupported claim of clinical harm.

Clinical AICalibrationFailure analysisSubgroup auditing

Research question

Average performance can hide repeated failure in a smaller group of patients.

Clinical AI is usually evaluated one model and one task at a time. That approach can miss a more consequential pattern: errors that repeatedly concentrate in the same patients, encounters, examinations, or care settings.

This project links reliability signals from emergency-department prediction models, chest-radiograph models, care phenotypes, and missingness patterns. The analysis separates subject-level overlap from encounter-level and time-windowed overlap so that later inpatient information is not treated as if it were available at the intended prediction time.

Structured clinical modelsEncounter-level predictions
Chest-radiograph modelsImage-level disagreement
Patient and encounter reliability profileContinuous scores, overlap, discordance, uncertainty
Care phenotypesHigh-acuity clinical context
Missingness patternsDocumentation and measurement context
Cross-modal audit framework. Structured clinical predictions, radiograph-level disagreement, care patterns, and missingness are linked at the patient and encounter levels. Current work focuses on reproducible linkage and sensitivity analysis before defining a shared reliability phenotype.

Study design

Temporal alignment is part of the scientific question.

A patient may have several encounters and several radiographs, while care status changes over time. I am defining explicit linkage windows, the unit of analysis, and what information would have been available at each prediction point.

These decisions determine whether an apparent cross-modal pattern reflects patient vulnerability, repeated measurements, a care episode, or information leakage.

PatientStable subject identifier
ED encounterPrediction time
RadiographOne or more image times
Admission windowCare context changes
Phenotype intervalTime-bounded interpretation

My contribution

Reliability checks are part of the primary analysis.

I am developing the linked cohort, documenting temporal assumptions, comparing continuous reliability scores, and testing whether conclusions persist across thresholds, task subsets, consensus rules, and aggregation choices. Calibration, bootstrap uncertainty, and subgroup auditing are built into the analysis rather than added after a result is selected.

Reliability checkStructured clinical modelsChest radiographsCross-modal interpretation
CalibrationConfidence versus observed errorTask-level and image-level calibrationCompare reliability on a common scale
Threshold sensitivityRisk-score operating pointsConsensus quorum and task subsetTest whether overlap depends on one cutoff
Linkage sensitivityEncounter definitionsRepeated-image aggregationSubject, encounter, and time-windowed overlap
UncertaintyBootstrap intervalsScore and ensemble variabilityStability of overlap and discordance
SubgroupsPatient and care strataExam and patient strataDescriptive disparity without causal overreach

Reliability checks across modalities. Primary conclusions should remain stable under reasonable changes to thresholds, task definitions, linkage windows, and aggregation rules.

Interpretation boundary

Model-label disagreement is an audit signal, not a clinical outcome.

The immediate goal is to determine whether cross-modal vulnerability is stable, interpretable, and reproducible. Claims about harm, intervention, or deployment should wait until the linked cohort, score definitions, sensitivity plan, and downstream clinical outcomes support them.

View MIT role