Research program
Models should state not only what they predict, but what the evidence permits us to conclude.
My research asks how patient-level reliability, causal structure, and biological constraints should govern what a model is allowed to claim. A higher performance metric is not sufficient; the evidence must also be interpretable in relation to the scientific or clinical decision.
Overview
Three research areas, one evidentiary standard.
The work spans medical images, electronic health records, physiological measurements, and biomolecular data.
Across these settings, I use causal reasoning, calibrated uncertainty, mechanistic structure, and reproducible computation to determine whether a model supports the decision at hand and where its interpretation should stop.
Project portfolio
Selected projects in order of current research priority.
The project pages distinguish ongoing methodological development from completed analyses and published work.
Trustworthy multimodal clinical AI
Identify where errors concentrate, not only how often they occur.
At MIT Critical Data, I evaluate diagnostic-imaging models and link image-level errors to high-acuity clinical context. The unit of analysis is the examination, encounter, or patient, because cohort averages can conceal repeated failure in smaller groups.
- Calibration and uncertainty quantification
- Consensus-based harm scoring across model ensembles
- Subgroup, fairness, and instance-level error auditing
- Encounter-level linkage of MIMIC-CXR and MIMIC-IV
- Threshold and severity-sensitivity analysis
Causal inference & patient-specific decisions
Estimate structure, then quantify how stable the estimate is.
My causal-inference work combines ensemble Bayesian network learning, conditional-independence testing, MCMC exploration of graph structures, and bootstrap assessment of edge stability. BaMANI developed from my doctoral research on heterocellular networks in cancer.
The same evidentiary discipline guides my work on epilepsy-surgery outcomes. I interpret diagnostic performance alongside subgroup patterns, structural missingness, sensitivity analyses, and the clinical meaning of proposed thresholds.
- Bayesian and ensemble causal-network inference
- Bootstrap edge stability and multiple-testing control
- Penalized regression and prespecified sensitivity analysis
- Failure analysis for patient selection and treatment counseling
Mechanistic & multiscale scientific ML
Use biological structure to constrain prediction and simulation.
At Duke, I connect agent-based and differential-equation models with machine learning. Gradient-boosted trees, Gaussian-process surrogates, and physics-informed neural networks reduce the computational cost of exploring mechanistic systems while retaining explicit attention to uncertainty.
In the NIH-supported Duke and Weill Cornell congenital CMV project, these methods support maternal-fetal immune modeling, vaccine-efficacy analysis, and cross-species immune-cell alignment.
- Mechanistic simulators, ODEs, and agent-based models
- Gaussian-process and machine-learning surrogates
- Physics-informed neural networks
- Optimal transport with relaxed marginals
- Docker and SLURM/HPC reproducibility
Methodological backbone