Dissertation
Bayesian finite mixture regression models with cluster-specific variable selection
University of Iowa
Doctor of Philosophy (PhD), University of Iowa
Spring 2023
DOI: 10.25820/etd.007011
Abstract
It's common to encounter data from heterogeneous populations. In regression problems, subpopulations may exist and differ in terms of which sets of covariates (e.g. treatment) best predict the response, their effect sizes, as well as the distribution of these covariates. When the subpopulations were unidentified or potentially dependent on unobserved variables, clustering analysis is needed conjointly with the regression analysis, for which Bayesian mixture models provide a natural framework. So far, due to computing considerations, only certain infinite mixture models including the Dirichlet process mixture models have been used in this context, though they produce inconsistent estimation of the number of clusters.
Our work is a first in using Bayesian finite mixture models to analyze heterogeneous regression data. Specifically, we adopt a class of priors based on the Normalized Independent Finite Point Process (NIFPP), which includes the finite Dirichlet prior as a special case. In addition, by utilizing NIFPP priors with spike-and-slab base measures, we can construct mixture models that are capable of performing cluster-specific variable selection. We prove posterior consistency for the proposed models and develop and code new MCMC algorithms for them.
Various ways to do posterior inference are demonstrated, such as clustering, individual profiling, and prediction. Our empirical studies demonstrate that the proposed method outperforms existing Bayesian and non-Bayesian methods in terms of clustering, variable selection, and prediction. Finally, we apply our model to real-world datasets to showcase its practical applications.
Details
- Title: Subtitle
- Bayesian finite mixture regression models with cluster-specific variable selection
- Creators
- Zhen Wang
- Contributors
- Aixin Tan (Advisor)Patrick Breheny (Committee Member)Joyee Ghosh (Committee Member)Sanvesh Srivastava (Committee Member)Luke Tierney (Committee Member)
- Resource Type
- Dissertation
- Degree Awarded
- Doctor of Philosophy (PhD), University of Iowa
- Degree in
- Statistics
- Date degree season
- Spring 2023
- Publisher
- University of Iowa
- DOI
- 10.25820/etd.007011
- Number of pages
- xvi, 163 pages
- Copyright
- Copyright 2023 Zhen Wang
- Language
- English
- Date submitted
- 04/25/2023
- Date approved
- 05/01/2023
- Description illustrations
- illustrations, tables, graphs
- Description bibliographic
- Includes bibliographical references (pages 156-163).
- Public Abstract (ETD)
- In practical applications, it is typical to encounter data from heterogeneous population, meaning that the subjects come from different subpopulations with distinct characteristics. In regression problems, when both the response (e.g., clinical endpoints) and covariate variables (e.g., demographics) are present, subpopulations may differ in terms of which covariates can best predict the response, their effect sizes and the distributions of the covariates. Bayesian mixture models provide a natural framework to model heterogeneous data. While some infinite mixture models are popular due to computational advantages, they tend to overestimate the number of clusters. To address this issue, we adopt Bayesian finite mixture models to analyze heterogeneous regression data. Additionally, we explore novel priors on the mixture weights, some of which exhibit good properties both a priori and a posteriori. In regression problems, a critical aspect is selecting the most effective set of covariates for predicting the response. This becomes even more important in clustering analysis as removing redundant variables can often improve the accuracy of clustering. To tackle this issue, we propose using a spike-and-slab prior on the regression coefficients, which helps select the most useful covariates within each cluster. We show that our proposed models provide consistent estimates for the number of clusters and other parameters in the model. New Markov Chain Monte Carlo (MCMC) methods are developed for computing. We demonstrate the superior performance of the proposed models compared with existing ones through empirical studies, and show how they can be applied to real-world data for analysis.
- Academic Unit
- Statistics and Actuarial Science
- Record Identifier
- 9984437256902771
Metrics
30 Record Views