Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Unveils DataSpace: A New Benchmark for Verifiable Data Agents in Complex Workspaces

AI By Crimson AI Hugging Face Papers 7 August 2026 · 00:00 25 views
Share: X Telegram

Hugging Face introduces DataSpace, a benchmark with 410 cross-language tasks and 7,439 artifacts (15.01 GB) across six formats, challenging data agents to produce verifiable tabular results from heterogeneous workspaces. Best accuracy reaches 66.34%, highlighting significant room for improvement.

Hugging Face Unveils DataSpace: A New Benchmark for Verifiable Data Agents in Complex Workspaces

Key points

Hugging Face has released DataSpace, a new benchmark designed to evaluate data agents on their ability to navigate heterogeneous workspaces and deliver verifiable tabular results. Unlike existing benchmarks that focus on isolated tasks like structured querying or retrieval, DataSpace requires agents to discover relevant evidence scattered across multiple formats and languages, integrate it, and produce complete tables.

The benchmark comprises 410 cross-language analytical tasks and 7,439 artifacts totaling 15.01 GB, spanning CSV, JSON, SQLite, Markdown, PDF, and video formats. Scenarios are drawn from financial, macroeconomic, and healthcare domains. DataSpace was also the official evaluation benchmark for the KDD Cup 2026 competition on Data Agents for Complex Data Analysis, which attracted 703 teams and 1,307 participants.

To ensure reliable evaluation, DataSpace uses an execution-grounded construction process with human review by 11 domain experts. A deterministic evaluator performs header-invariant column alignment, type- and precision-aware normalization, and order-aware row comparison, ensuring that results are judged consistently and without model bias.

Initial evaluations across six frontier multimodal models and five agent harnesses reveal that the best accuracy reaches only 66.34%, with 76 tasks missed by all tested models. Notably, changing the agent harness while keeping the backbone fixed creates a 15.36-point accuracy spread, and multimodal evidence integration and joins consistently reduce accuracy across all backbones. These findings indicate that DataSpace remains unsaturated and highlight key challenges for improving data-agent reliability.

The DataSpace benchmark, dataset, evaluator, baselines, and leaderboard are publicly available, inviting the community to test new models and agent designs. The paper is available on arXiv, with code and resources on GitHub and Hugging Face.

MetricValue
Tasks410
Artifacts7,439
Total size15.01 GB
FormatsCSV, JSON, SQLite, Markdown, PDF, video
Best accuracy66.34%
Accuracy spread (harness)15.36 points
Tasks missed by all models76
KDD Cup teams703
KDD Cup participants1,307
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1