Empowering health data insights: tools for anomaly detection and multivariate clustering
Abstract
Details
- Title: Subtitle
- Empowering health data insights: tools for anomaly detection and multivariate clustering
- Creators
- Daren Kuwaye
- Contributors
- Hyunkeun Ryan Cho (Advisor)Emine Bayman (Committee Member)Joseph E. Cavanaugh (Committee Member) - University of Iowa, Statistics and Actuarial ScienceJacob Oleson (Committee Member)
- Resource Type
- Dissertation
- Degree Awarded
- Doctor of Philosophy (PhD), University of Iowa
- Degree in
- Biostatistics
- Date degree season
- Autumn 2023
- DOI
- 10.25820/etd.006978
- Publisher
- University of Iowa
- Number of pages
- xii, 102 pages
- Copyright
- Copyright 2023 Daren Kuwaye
- Comment
This thesis has been optimized for improved web viewing. If you require the original version, contact the University Archives at the University of Iowa: https://www.lib.uiowa.edu/sc/contact/.
- Language
- English
- Date submitted
- 12/04/2023
- Description illustrations
- Illustrations, tables, graphs, charts
- Description bibliographic
- Includes bibliographical references (pages 68-71).
- Public Abstract (ETD)
Medical records have come a long way from ancient stone tablets to today’s electronic medical records (EMRs) and wearable device data. These modern systems generate vast datasets with hidden insights and clustering techniques can be used to uncover them. However, these data have challenges: wearable data consist of multiple channels that are highly correlated, and EMRs have inconsistent measurement schedules. Furthermore, EMR data can contain anomalies from human mistakes. Therefore ensuring data quality is crucial due to the far-reaching implications of clustering in medical treatment, resource allocation, and policy.
This dissertation addresses these challenges with three tools. First, a method to identify and correct anomalies in EMR data is introduced. This method is designed to be conservative, so that it focus on anomalies arising from human mistakes. Extensive simulation studies are used to validate this approach and a real-world application shows the value in this tool to improve data quality.
Then, a technique for clustering EMRs with an inconsistent measurement schedule is presented. A statistical model is used to capture the individual trends and estimate covariance over time, and a proximity which uses this information to improve clustering accuracy is presented. The k-medoids algorithm, in combination with this proximity, precisely groups subjects with similar trends. Extensive simulation studies demonstrate its superiority to existing methods and we substantiate its practical application in two real-world case studies.
Finally, the challenge of high correlations in wearable device data is addressed by modeling the correlated channels simultaneously. This allows the correlation to be captured so that it can better inform calculating the proximity between subjects. This approach, combined with the k-medoids algorithm, improves the clustering of subjects with consistent trends across various channels. Extensive simulation studies showcase its superiority in handling strong correlations. We also provide a practical application demonstrated in an electroencephalography study.
- Academic Unit
- Biostatistics
- Record Identifier
- 9984546849502771