Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Inaudible Low-Frequency Attacks Can Cripple Audio-Language Models, New Study Warns

AI By Crimson AI Hugging Face Papers 15 August 2026 · 00:00 11 views
Share: X Telegram

Researchers propose a black-box red-teaming method using inaudible low-frequency waveforms to expose vulnerabilities in audio-language models, alongside a defense that detects distribution shifts and requests a second recording to recover accuracy.

Inaudible Low-Frequency Attacks Can Cripple Audio-Language Models, New Study Warns

Key points

A new research paper from Hugging Face highlights a previously overlooked safety risk in large audio-language models (LALMs): inaudible low-frequency signals that can be injected into the model to cause failures. The study, titled "From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs," proposes a black-box red-teaming method and a corresponding defense mechanism.

The attack method, called Intermittent Low-Frequency Lockout (ILL), uses a universal waveform template to generate inaudible low-frequency signals. It employs two techniques: Sentence Attention Scale Estimation to determine active intervals and Frequency Confusion Transfer to construct a low-frequency waveform with continuous phase from corpus spectral variation. This allows the attack to be applied without internal model access.

Testing across six LALMs and multiple audio understanding tasks, ILL reduced accuracy by up to 67 percentage points. Notably, the attack was nearly imperceptible to humans, receiving a mean audibility rating of 1.33, close to the 1.17 rating for clean audio.

To mitigate this risk, the authors propose Distributional Requery Guard (DRG), a defense that detects low-frequency distribution shifts and conditionally requests a second recording for semantic recovery. DRG raised mean attacked accuracy from 28.5% to 46.1% after clean reacquisition.

The findings underscore a critical vulnerability in LALMs and provide a foundation for future research on robust audio understanding. The paper also lists several related works, including denial-of-service attacks on speech language models and prosody-driven jailbreaks.

MetricValue
Accuracy reduction (max)67 percentage points
Mean human audibility rating (attack)1.33
Mean human audibility rating (clean)1.17
Mean attacked accuracy (before defense)28.5%
Mean attacked accuracy (after DRG)46.1%
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4