Logo image
A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings
Journal article   Open access   Peer reviewed

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings

Erik D VonKaenel, Lisa M Bramer, Javier E Flores, Thomas O Metz, Ernesto S Nakayasu and Bobbie-Jo M Webb-Robertson
PLoS computational biology, Vol.22(8), e1014625
08/2026
DOI: 10.1371/journal.pcbi.1014625
PMID: 42585225
url
https://doi.org/10.1371/journal.pcbi.1014625View
Published (Version of record) Open Access

Abstract

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.
Algorithms Machine Learning Software Benchmarking Classification Algorithms Computational Biology - methods Diabetes Mellitus, Type 1 - classification Diabetes Mellitus, Type 1 - genetics Diabetes Mellitus, Type 1 - metabolism Genomics - methods Humans Multiomics

Details

Metrics

2 Record Views
Logo image