Logo image
Variable screening for high-dimensional data via Pearson's chi-square statistic
Dissertation   Open access

Variable screening for high-dimensional data via Pearson's chi-square statistic

Xingzhi Wang
University of Iowa
Doctor of Philosophy (PhD), University of Iowa
Autumn 2023
DOI: 10.25820/etd.006979
pdf
PhD (Statistics) thesis by Xingzhi Wang2.98 MBDownloadView
Open Access Free to read and download

Abstract

For solving the problem of continuous domain regression or discrete domain classification, variable/feature screening is an effective approach and to a certain extent even computationally indispensable when the number of predictor variables, p, is extremely large, as compared to n, the number of data cases. We are firstly concerned with the situations where both the predictor and the response variables are categorical variables and then extend the usage of the marginal utility statistic proposed beyond the categorical setting, which would further expand its applications in diverse fields including economics and medical studies, etc. For categorical feature screening, we adopt the framework introduced by Guo et al. (2023) to screening categorical variables via the Pearson's chi-square statistic. We derive a set of sufficient conditions for controlling the false discovery rate of the proposed method, in the ultrahigh dimensional setting. Our theoretical innovations are twofold: (i) we apply Bayesian model averaging to obtain the desired false discovery rate under the exchangeability assumption and (ii) do so by establishing the uniform consistency of the Pearson's chi-square statistics at a certain rate that allows for diverging number of categories and some cells with nearly zero cell probabilities. Furthermore, equipped with the idea of data binning, we widen the applicability of the Pearson's chi-square statistic beyond random variables with finite supports, through the creation of a novel population-level dependency measure which we call the Pearson's chi-square divergence, and whose potency may be carried towards multivariate variables or even more exotic random processes. Our main theoretical developments are focused upon: (i) elaborating the conditions that guarantee the uniform consistency of the Pearson's chi-square statistics under data-based binning to the Pearson's chi-square divergence and (ii) offering some closed-form formulae, in terms of the sample size, for the number of bins that could be empirically used for certain classes of marginal distributions for the covariates and the response. Finally, we report results from extensive numerical studies for contrasting the empirical performance of various marginal utilities for categorical/continuous type data and further illustrate them with real applications.

Details

Metrics

6 File views/ downloads
29 Record Views
Logo image