Journal article
A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings
PLoS computational biology, Vol.22(8), e1014625
08/2026
DOI: 10.1371/journal.pcbi.1014625
PMID: 42585225
Abstract
In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.
Details
- Title: Subtitle
- A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings
- Creators
- Erik D VonKaenel - Pacific Northwest National LaboratoryLisa M Bramer - Pacific Northwest National LaboratoryJavier E Flores - Pacific Northwest National LaboratoryThomas O Metz - Pacific Northwest National LaboratoryErnesto S Nakayasu - Pacific Northwest National LaboratoryBobbie-Jo M Webb-Robertson - Pacific Northwest National Laboratory
- Resource Type
- Journal article
- Publication Details
- PLoS computational biology, Vol.22(8), e1014625
- DOI
- 10.1371/journal.pcbi.1014625
- PMID
- 42585225
- NLM abbreviation
- PLoS Comput Biol
- ISSN
- 1553-7358
- eISSN
- 1553-7358
- Publisher
- PLoS
- Grant note
- UC4 DK063836 / NIDDK NIH HHS U01 DK063865 / NIDDK NIH HHS UL1 TR002535 / NCATS NIH HHS UC4 DK112243 / NIDDK NIH HHS U01 DK127786 / NIDDK NIH HHS HHSN267200700014C / NIDDK NIH HHS UC4 DK095300 / NIDDK NIH HHS UC4 DK063863 / NIDDK NIH HHS UC4 DK063829 / NIDDK NIH HHS U01 DK063861 / NIDDK NIH HHS
- Language
- English
- Date published
- 08/2026
- Academic Unit
- Biostatistics
- Record Identifier
- 9985218625102771
Metrics
2 Record Views