Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Study: Optimal Data Repetition Scales Mildly with LLM Size

AI By Crimson AI Hugging Face Papers 17 August 2026 · 00:00 6 views
Share: X Telegram

A new paper from Hugging Face reveals that under proportional scaling of model size and training tokens, the optimal repetition of high-quality domain data increases mildly with scale and correlates with domain validation loss rather than unique data volume.

Hugging Face Study: Optimal Data Repetition Scales Mildly with LLM Size

Key points

As large language models (LLMs) scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (TPP). However, high-quality domain data is much harder to scale than general web data. As model size and the training-token budget increase, the fraction of high-quality data in the training mixture tends to decrease.

Repeating the available high-quality data provides an effective way to counteract this dilution, but excessive repetition may lead to overfitting. A new paper from Hugging Face studies this trade-off under practical LLM scaling, where the training-token budget grows proportionally with model size.

For a fixed domain, the researchers found that, surprisingly, at a fixed TPP, the optimal repetition count mildly increases with model size. Across different domains, the optimal repetition count is strongly negatively correlated with the final validation loss of a domain: domains with lower loss can generally benefit from more repetitions. In contrast, the amount of unique domain data is only weakly related to the optimal repetition count.

These findings suggest that repetition counts tuned on smaller proxy models with the same TPP can provide a practical estimate for larger models, offering a cost-effective strategy for optimizing pretraining data mixtures.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4