Statistics Seminar · Fall 2026

Department of Mathematical Sciences
IU Indianapolis

Invited seminar talk

Uncertainty Quantification for Modern Machine Learning Predictions

Mengxin (Maxine) Yu, PhD

Assistant Professor of Statistics and Data Science · Washington University in St. Louis

Portrait of Maxine Yu
Date Tuesday, October 13, 2026
Time 12:15–1:15 PM Eastern Time
Online via Zoom ID 845 0989 4694 Password 113959 · Join the seminar

Abstract

Quantifying the uncertainty of black-box machine learning predictions is a core problem in modern statistics. Methods for predictive inference have been developed under a variety of assumptions, often—for instance, in standard conformal prediction—relying on the invariance of the distribution of the data under special groups of transformations such as permutation groups. Moreover, many existing methods for predictive inference aim to predict unobserved outcomes in sequences of feature-outcome observations. Meanwhile, there is interest in predictive inference under more general observation models (e.g., for partially observed features) and for data satisfying more general distributional symmetries beyond exchangeability (e.g., network, rotationally invariant, data with hierarchical structure). Here, we propose SymmPI, a unified methodology for predictive inference when data distributions have general group symmetries in arbitrary observation models. Our methods leverage the novel notion of distributional equivariant transformations, which process the data while preserving their distributional invariances. We show that SymmPI has valid coverage under distributional invariance and characterize its performance under distribution shift, recovering recent results as special cases. These methodologies are particularly relevant for cluster-randomized trials in clinical settings, where prediction reliability is essential. If time permits, I will also briefly present our work on evaluating uncertainty and confidence measures in large language models. We introduce a novel method called rank calibration, which enables the identification of reliable uncertainty measures across a range of tasks and LLM models.