Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Study: LLMs Fabricate User Profiles in 41.6% of Claims; Self-Monitoring Misleads

AI By Crimson AI Hugging Face Papers 6 August 2026 · 00:00 16 views
Share: X Telegram

A new benchmark, MirageBench, reveals that all 12 tested LLMs over-infer user attributes in 35-49% of claims, and that self-reported confidence inversely correlates with actual fabrication, undermining self-monitoring as a reliability signal.

Study: LLMs Fabricate User Profiles in 41.6% of Claims; Self-Monitoring Misleads

Key points

A new research paper from Hugging Face introduces MirageBench, a benchmark designed to measure over-inference (OI) in personalized LLMs—the tendency of models to fabricate user attributes beyond what the evidence supports. The study, which evaluates 12 models across 7 families on 143,616 judged claims, finds that over-inference is pervasive: every model over-infers between 35% and 49% of its claims, with a cross-model mean of 41.6%.

The benchmark comprises 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, and spans 6 personalization tasks along an "imagination gradient." A four-way faithfulness taxonomy is operationalized by an independent judge, validated against blind human annotation on 400 claims with high agreement (Cohen's kappa = 0.863 four-class, 0.900 binary).

Most strikingly, the authors report a Self-Monitoring Inversion: at the model-selection level, models' self-assessed OI is negatively rank-correlated with judge-measured OI (rho = -0.60, p = 0.044). In other words, models that claim to over-infer the least are often the ones that fabricate the most. This suggests that self-reported confidence is a misleading signal for comparing models, even though within a single model, self-audit still ranks that model's own claims moderately well (AUROC 0.58–0.83).

The study also shows that OI is task-dependent, ranging from 27% to 59%, and that in a multi-turn pilot, inferred attributes accumulate approximately linearly with little revision. The authors argue that external verification, rather than model self-report, is a more reliable foundation for trustworthy personalization.

MetricValue
Models evaluated12
Model families7
Total judged claims143,616
Cross-model mean OI41.6%
OI range across models35%–49%
OI range across tasks27%–59%
Self-monitoring rank correlation (rho)-0.60 (p=0.044)
Self-audit AUROC (within model)0.58–0.83
Human agreement (four-class kappa)0.863
Human agreement (binary kappa)0.900
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1