Logo image
Adaptive Perturbation Selection for Contrastive Audio Decoding
Preprint   Open access

Adaptive Perturbation Selection for Contrastive Audio Decoding

Aaron Isidore Grace, Zhouyuan Huo and Weiran Wang
ArXiv.org
arXiv
06/30/2026
DOI: 10.48550/arxiv.2607.00247
url
https://doi.org/10.48550/arxiv.2607.00247View
Preprint (Author's original) This preprint has not been evaluated by subject experts through peer review. Preprints may undergo extensive changes and/or become peer-reviewed journal articles. Open Access

Abstract

Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding (CD) offers training-free mitigation, existing methods rely on blunt perturbations like masking or noise, leaving structured audio transformations unexplored. We explore this design space by evaluating a diverse library of targeted audio perturbations and adaptively selecting the optimal negative branch for each task and example. First, we improve upon earlier prompt engineering by showing that a simple binary yes/no constraint reduces the model's tendency to falsely confirm absent audio features. Second, evaluating our library across temporal, spectral, frequency, and amplitude domains reveals that optimal transformations are highly task-dependent; for instance, reversing the audio array disrupts temporal coherence, raising accuracy on the temporal order task from 74.7% to 81.4%. Finally, we trained a light-weight perturbation selector on model hidden states to dynamically route negative branches, yielding an additional +4.3% gain on the existence task.
Computer Science - Artificial Intelligence Computer Science - Sound

Details

Metrics

1 Record Views
Logo image