MIT Critical Data · Ongoing collaboration
Reliability and Failure Analysis for Multimodal Clinical AI
Can the same patients or encounters be difficult for AI systems across structured clinical data and chest radiographs? This project builds the linked cohort and reliability analyses needed to answer that question without turning model disagreement into an unsupported claim of clinical harm.
Research question
Average performance can hide repeated failure in a smaller group of patients.
Clinical AI is usually evaluated one model and one task at a time. That approach can miss a more consequential pattern: errors that repeatedly concentrate in the same patients, encounters, examinations, or care settings.
This project links reliability signals from emergency-department prediction models, chest-radiograph models, care phenotypes, and missingness patterns. The analysis separates subject-level overlap from encounter-level and time-windowed overlap so that later inpatient information is not treated as if it were available at the intended prediction time.
Study design
Temporal alignment is part of the scientific question.
A patient may have several encounters and several radiographs, while care status changes over time. I am defining explicit linkage windows, the unit of analysis, and what information would have been available at each prediction point.
These decisions determine whether an apparent cross-modal pattern reflects patient vulnerability, repeated measurements, a care episode, or information leakage.
My contribution
Reliability checks are part of the primary analysis.
I am developing the linked cohort, documenting temporal assumptions, comparing continuous reliability scores, and testing whether conclusions persist across thresholds, task subsets, consensus rules, and aggregation choices. Calibration, bootstrap uncertainty, and subgroup auditing are built into the analysis rather than added after a result is selected.
| Reliability check | Structured clinical models | Chest radiographs | Cross-modal interpretation |
|---|---|---|---|
| Calibration | Confidence versus observed error | Task-level and image-level calibration | Compare reliability on a common scale |
| Threshold sensitivity | Risk-score operating points | Consensus quorum and task subset | Test whether overlap depends on one cutoff |
| Linkage sensitivity | Encounter definitions | Repeated-image aggregation | Subject, encounter, and time-windowed overlap |
| Uncertainty | Bootstrap intervals | Score and ensemble variability | Stability of overlap and discordance |
| Subgroups | Patient and care strata | Exam and patient strata | Descriptive disparity without causal overreach |
Reliability checks across modalities. Primary conclusions should remain stable under reasonable changes to thresholds, task definitions, linkage windows, and aggregation rules.
Interpretation boundary
Model-label disagreement is an audit signal, not a clinical outcome.
The immediate goal is to determine whether cross-modal vulnerability is stable, interpretable, and reproducible. Claims about harm, intervention, or deployment should wait until the linked cohort, score definitions, sensitivity plan, and downstream clinical outcomes support them.
View MIT role