Dissertation
Theoretical understandings of robustness in machine learning
University of Iowa
Doctor of Philosophy (PhD), University of Iowa
Summer 2024
DOI: 10.25820/etd.007634
Abstract
In recent years, artificial intelligence (AI) and machine learning have played an increasingly important role in science and engineering fields, including in signal processing. Therefore, understanding the robustness of machine learning systems and designing robust machine learning systems are critical to signal processing and machine learning tasks such as those in self-driving cars, communications, and AI- aided medicine, where safety, reliability and security are of great concern.
To address this problem, in this thesis, we aim to contribute to theoretically understanding the robustness of machine learning models against compression of data, perturbations to data, outliers in data, and even adversarial attacks. Furthermore, we aim to apply our understandings to improve the robustness and performances of machine learning systems and signal processing systems employing data-driven models such as generative models.
For robustness of machine learning models against compression of data, we first formulate the problem of performing optimal data compression under the constraints that compressed data can be used for accurate classification in machine learning. We show that this translates to a problem of minimizing the mutual information between data and its compressed version under the constraint that error probability of classification or generalized cost is small when using the compressed data for machine learning. We then provide analytical and computational methods to characterize the optimal trade-off between data compression and classification error probability. To complete this, we provide an analytical characterization for the optimal compression v strategy for data with binary labels. Also, for data with multiple labels, we formulate a set of convex optimization problems, where the optimal trade-off between the classification error and compression efficiency can be obtained by numerically solving the formulated optimization problems. We further show the improvement of our formulations over the information-bottleneck methods in classification performance. We also propose two new loss functions that can be used in multi-class classification task, where there are asymmetric misclassification costs. The newly proposed loss functions work well for a general cost matrix. Furthermore, we propose a new classification algorithm, which decodes the data to the label corresponded to the smallest asymmetric cost.
For the robustness of machine learning models against outliers in data, we consider the problem of recovering signals modeled by generative models from linear measurements contaminated with sparse outliers. We propose an outlier detection approach for reconstructing the ground-truth signals modeled by generative models under sparse outliers. We establish theoretical recovery guarantees for reconstruction of signals using generative models in the presence of outliers, and give lower bounds on the number of correctable outliers. Our results are applicable to both linear generator neural networks and the nonlinear generator neural networks with an arbitrary number of layers. We propose an iterative alternating direction method of multipliers (ADMM) algorithm for solving the outlier detection problem via ℓ1 norm minimization, and a gradient descent algorithm for solving the outlier detection problem via squared ℓ1 norm minimization. We also conduct extensive experiments using variational auto-encoder and deep convolutional generative adversarial networks. The experimental results show that the signals can be successfully reconstructed under outliers using our approach and our approach outperforms the traditional Lasso and vi ℓ2 minimization approach.
For the robustness of machine learning models against adversarial attacks, we study the adversarial robustness of deep neural networks for classification tasks. We look at the smallest magnitude of possible additive perturbations that can change the output of a classification algorithm. We provide a matrix-theoretic explanation of the adversarial fragility of deep neural network for classification. In particular, our theoretical results show that neural networks’ adversarial robustness can degrade as the input dimension d increases. Analytically, we show that neural networks’ adversarial robustness can be only 1/√d of the best possible adversarial robustness. Our matrix-theoretic explanation is consistent with an earlier information-theoretic feature-compression-based explanation for the adversarial fragility of neural networks.
Details
- Title: Subtitle
- Theoretical understandings of robustness in machine learning
- Creators
- Jingchao Gao
- Contributors
- Weiyu Xu (Advisor)Palle Jorgensen (Committee Member)Qihang Lin (Committee Member)Raghu Mudumbai (Committee Member)
- Resource Type
- Dissertation
- Degree Awarded
- Doctor of Philosophy (PhD), University of Iowa
- Degree in
- Applied Mathematical and Computational Sciences
- Date degree season
- Summer 2024
- Publisher
- University of Iowa
- DOI
- 10.25820/etd.007634
- Number of pages
- xi, 132 pages
- Copyright
- Copyright 2024 Jingchao Gao
- Grant note
- I would like to thank Herbert-Hethcote Grant for providing travel support as well as book purchasing support. (iv)
- Language
- English
- Date submitted
- 06/23/2024
- Description illustrations
- illustrations, tables, graphs
- Description bibliographic
- Includes bibliographical references (pages 123-132).
- Public Abstract (ETD)
- In recent years, artificial intelligence (AI) and machine learning have played an increasingly important role in science and engineering fields, including in signal processing. Therefore, understanding the robustness of machine learning systems and designing robust machine learning systems are critical to signal processing and machine learning tasks such as those in self-driving cars, communications, and AI- aided medicine, where safety, reliability and security are of great concern. To address this problem, in this thesis, we aim to contribute to theoretically understanding the robustness of machine learning models against compression of data, perturbations to data, outliers in data, and even adversarial attacks. Furthermore, we aim to apply our understandings to improve the robustness and performances of machine learning systems and signal processing systems employing data-driven models such as generative models.
- Academic Unit
- Interdisciplinary Graduate Program in Applied Mathematical & Computational Sciences
- Record Identifier
- 9984698054302771
Metrics
2 File views/ downloads
5 Record Views