On generative models for conformal prediction
Abstract
Details
- Title: Subtitle
- On generative models for conformal prediction
- Creators
- Zhenhan Fang
- Contributors
- Aixin Tan (Advisor)Jian Huang (Advisor)Kung-Sik Chan (Committee Member)Boxiang Wang (Committee Member)
- Resource Type
- Dissertation
- Degree Awarded
- Doctor of Philosophy (PhD), University of Iowa
- Degree in
- Statistics
- Date degree season
- Spring 2026
- DOI
- 10.25820/etd.008391
- Publisher
- University of Iowa
- Number of pages
- ix, 81 pages
- Copyright
- Copyright 2026 Zhenhan Fang
- Language
- English
- Date submitted
- 04/26/2026
- Description illustrations
- illustrations, tables, graphs
- Description bibliographic
- Includes bibliographical references (pages 74-77).
- Public Abstract (ETD)
When making predictions, it is important to communicate not just a single best guess but a range of plausible outcomes. For example, predicting where a taxi passenger will be dropped off involves uncertainty: the passenger could be heading to any of several destinations. A prediction region is a set of outcomes that is guaranteed to contain the true result with high probability, such as 90% of the time.
For problems with multiple output variables, such as predicting the location of a vehicle, forecasting temperature and humidity simultaneously, or predicting the opening and closing prices of a stock, constructing useful prediction regions is difficult. Simple shapes like rectangles or ellipses are easy to compute but often too large, covering many implausible locations. More flexible methods exist but produce irregular, fragmented regions that are hard to interpret.
This dissertation develops two new methods that use generative models, a class of machine learning models that learn to capture the likely outcomes, to build prediction regions that closely follow the shape of the data while maintaining statistical guarantees. The first method, CONTRA, uses a type of generative model called a normalizing flow to produce smooth, connected regions. The second method, TRACE, uses diffusion and flow matching models, which are more flexible and perform well even when the data has complex structure or many input variables. Both methods are validated on synthetic and real-world datasets, including taxi trip prediction and energy load forecasting, and consistently produce more compact and interpretable prediction regions than existing approaches.
These methods are broadly applicable to any setting that requires reliable multi-dimensional predictions. For example, in search and rescue, a tighter predicted area for a missing person’s location means a smaller region to search and a faster rescue; in autonomous driving, a more precise prediction of where a pedestrian might move helps the vehicle react safely without braking unnecessarily.
- Academic Unit
- Statistics and Actuarial Science
- Record Identifier
- 9985176974802771