Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

NOLLI Benchmark Reveals Korean AI Gaps in Jamo Execution, Not Language

AI By Crimson AI Hugging Face Papers 6 August 2026 · 00:00 23 views
Share: X Telegram

A new procedurally generated puzzle benchmark, NOLLI, diagnoses where English-Korean performance gaps arise in AI models, finding minimal language-driven gaps but significant deficits in multi-step Hangul-jamo execution and Korean kinship terms.

NOLLI Benchmark Reveals Korean AI Gaps in Jamo Execution, Not Language

Key points

Hugging Face researchers have introduced NOLLI, a procedurally generated puzzle benchmark designed to diagnose where Korean language performance gaps arise in AI models. The benchmark comprises 15 puzzle types (25 tasks; 7,500 items), with every instance seed-regenerable, verified to have a unique solution, and scored deterministically.

Rather than equating harder with bigger, the team calibrated difficulty behaviorally, tuning each generator until a fixed reference model lands in target accuracy bands. The three-level design combines matched direct translations, script adaptations over Hangul jamo (sub-syllabic letters), and Korean-only tasks grounded in Korean culture or orthography.

Evaluating 15 frontier, open-weight, and Korean-developed models, the study found that among the 12 above a 3% overall-accuracy floor, matched English-Korean accuracy is statistically equivalent within a ±10 percentage point margin (TOST), suggesting little cost from presentation language alone. However, writing-system-intensive tasks show sharper gaps: Korean Cipher falls behind English by up to 68.7 percentage points, whereas Cryptarithmetic over the same jamo shows no systematic penalty, and Jamo Composition accuracy predicts Korean Cipher accuracy.

These contrasts are diagnostic rather than causal, consistent with difficulty in multi-step sub-syllabic execution. Korean-only tasks separate rule-application deficits, which vary in sign, from a Kinship deficit positive in all 12 models. Finally, a salient size measure fails to grow from Easy to Hard in 7 of 15 types, making structural size an unreliable proxy for empirical difficulty.

The dataset is available at HAERAE-HUB/NOLLI.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1