Preprint
Improving Audio Event Recognition with Consistency Regularization
ArXiv.org
Cornell University
09/12/2025
DOI: 10.48550/arxiv.2509.10391
Abstract
Consistency regularization (CR), which enforces agreement between model predictions on augmented views, has found recent benefits in automatic speech recognition [1]. In this paper, we propose the use of consistency regularization for audio event recognition, and demonstrate its effectiveness on AudioSet. With extensive ablation studies for both small ( 20k) and large ( 1.8M) supervised training sets, we show that CR brings consistent improvement over supervised baselines which already heavily utilize data augmentation, and CR using stronger augmentation and multiple augmentations leads to additional gain for the small training set. Furthermore, we extend the use of CR into the semi-supervised setup with 20K labeled samples and 1.8M unlabeled samples, and obtain performance improvement over our best model trained on the small set.
Details
- Title: Subtitle
- Improving Audio Event Recognition with Consistency Regularization
- Creators
- Shanmuka SadhuWeiran Wang
- Resource Type
- Preprint
- Publication Details
- ArXiv.org
- DOI
- 10.48550/arxiv.2509.10391
- ISSN
- 2331-8422
- Publisher
- Cornell University; Ithaca, New York
- Language
- English
- Date posted
- 09/12/2025
- Academic Unit
- Computer Science
- Record Identifier
- 9984966540702771
Metrics
7 Record Views