Logo image
Software for analyzing global family-based association studies: penalized linear mixed models for correlated genetic data with application to orofacial clefts
Dissertation   Open access

Software for analyzing global family-based association studies: penalized linear mixed models for correlated genetic data with application to orofacial clefts

Tabitha K. Peter
University of Iowa
Doctor of Philosophy (PhD), University of Iowa
Spring 2025
DOI: 10.25820/etd.007922
pdf
tabpeter_thesis_may056.78 MBDownloadView
Open Access Free to read and download

Abstract

Orofacial clefts are a major global public health problem, and improving health outcomes for infants and their families requires understanding the complex genetic mechanisms that contribute to the formation of such congenital anomalies. In turn, increasing our understanding of such genetic mechanisms requires that genetic data represent global populations. Improving human health requires having data that represents the full kaleidoscope of human diversity as well as family structure. The present work was motivated by a dataset at the intersection of orofacial clefts as a public health problem and the need for genetic data to represent families from global populations. The Pittsburgh Orofacial Cleft studies included an international, family-based, case-control genome-wide association study (GWAS) of orofacial cleft conditions designed to analyze the genetic factors associated with clefting in global populations. Cleft patients, most of whom were children, were recruited into the study along with their parents, siblings, and extended family members. Families were recruited from data collection sites that represented fourteen countries across five continents. These genetic data have a complex correlation structure, with family groups nested inside each data collection site. To analyze these data, we proposed a penalized linear mixed model. Chapter 2 identifies and address pitfalls that arise in this modeling process. We pay particular attention to developing an implementation of cross-validation that is appropriate for high-dimensional, correlated data setting. Cross validation is a leading method for model selection in penalized regression, making it an essential component of analysis. Chapter 3 describes an R package, plmmr, that implements penalized linear mixed models for high-dimensional, correlated data. This package is built to accommodate genome-scale data; users do not have to read their data into memory. The package is now available on CRAN, making it widely accessible. Combining the methodological and computational components, Chapter 4 analyzes GWAS data from the Pittsburgh Orofacial Cleft studies. Our analyses identifies new findings related to enrichment and protein-protein interaction, along with confirming the reproducibility of several past findings in the orofacial cleft literature. Finally, Chapter 5 integrates the perspectives and experiences of the orofacial cleft patient community into the work. In this qualitative analysis, we analyzed data from interviews with the parents/caregivers of children receiving cleft treatment at the Fundación Clínica Noel, the children's hospital that served as the leading recruitment center for the Pittsburgh GWAS study. This chapter theorizes an explanatory model of the etiology of cleft conditions in this population. Beyond the explanatory model, we make inferences about these families' perspectives towards genetic testing, creating a baseline of data to guide future return-of-results.
lasso regression Linear mixed models orofacial clefts Penalized regression

Details

Metrics

15 Record Views
Logo image