Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Massive Activations in Hybrid Linear Attention LLMs: New Study Reveals Pre-Attention Spikes and Inter-Spike Plateaus

AI By Crimson AI Hugging Face Papers 14 August 2026 · 00:00 15 views
Share: X Telegram

A new study from Hugging Face provides the first systematic analysis of massive activations in hybrid linear-attention LLMs, uncovering two distinct morphologies—pre-attention spikes and inter-spike plateaus—that are governed by cancellation timing and recover full-attention behavior at the limit.

Massive Activations in Hybrid Linear Attention LLMs: New Study Reveals Pre-Attention Spikes and Inter-Spike Plateaus

Key points

A new research paper from Hugging Face presents the first systematic study of massive activations (MAs) in layer-interleaved hybrid linear-attention (HLA) large language models. The authors identify two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming what they call pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP).

The study demonstrates that as full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology characteristic of full attention LLMs. This organization was found to recur across five linear attention architectures, six hybridization configurations, five data domains, and representative open-source hybrid models ranging from 1.2B to 397B parameters.

Controlled pretraining of GDN-based hybrids at scales up to 1.3B revealed that both morphologies emerge early in training and respond asymmetrically to output gating: full attention output gating strongly attenuates their absolute magnitudes without eliminating layerwise organization, whereas removing GDN gates yields comparatively modest amplification.

Mechanistically, the authors' systematic-outlier analysis supports a shared lifecycle account governed by the timing of MA cancellation. PAS follows a localized write-sink-cancel process, while the extended persistence of ISP is consistent with delayed cancellation. At the full attention limit, this account recovers the stable MA morphology seen in full attention LLMs. The code is available on GitHub.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4