Structured clinical models
Encounter-level predictions

MIT Critical Data · Ongoing collaboration
Do model failures concentrate in the same patients or encounters across structured clinical data and chest radiographs? This project is building the linked cohort and reliability analyses needed to answer that question without treating model disagreement as evidence of clinical harm.
Research question
Clinical AI is often evaluated one model and one task at a time. That design can miss a more consequential pattern: errors that repeatedly concentrate in the same patients, encounters, examinations, or care settings.
The project links reliability signals from emergency-department prediction models, chest-radiograph models, care phenotypes, and missingness patterns. The analysis distinguishes subject-level overlap from encounter-level and time-windowed overlap so that later inpatient information is not treated as though it were available at the intended prediction time.
Encounter-level predictions
Image-level disagreement
Continuous scores, overlap, discordance, uncertainty
High-acuity clinical context
Documentation and measurement context
Study design
A patient may have several encounters and several radiographs, and clinical status changes over time. I am defining explicit linkage windows, the unit of analysis, and the information that would have been available at each prediction point.
These choices determine whether an apparent cross-modal pattern reflects persistent patient-level vulnerability, repeated measurements, a single episode of care, or information leakage.
My contribution
I am developing the linked cohort, documenting temporal assumptions, comparing continuous reliability scores, and testing whether conclusions persist across thresholds, task subsets, consensus rules, and aggregation choices. Calibration, bootstrap uncertainty, and subgroup auditing are specified as part of the analysis rather than added after a favorable result is selected.
Confidence versus observed error
Task-level and image-level calibration
Compare reliability on a common scale
Risk-score operating points
Consensus quorum and task subset
Test whether overlap depends on one cutoff
Encounter definitions
Repeated-image aggregation
Subject, encounter, and time-windowed overlap
Bootstrap intervals
Score and ensemble variability
Stability of overlap and discordance
Patient and care strata
Exam and patient strata
Descriptive disparity without causal overreach
Reliability checks across modalities. Primary conclusions should remain stable under reasonable changes to thresholds, task definitions, linkage windows, and aggregation rules.
Interpretation boundary
The immediate goal is to determine whether cross-modal vulnerability is stable, interpretable, and reproducible. Claims about harm, intervention, or deployment require a completed linked cohort, prespecified score definitions, a sensitivity plan, and downstream clinical outcomes that support those claims.
View MIT role