Latest stories
Hugging Face Unveils FinanceComplexQA: A New Benchmark for Agentic Reasoning in Finance
FinanceComplexQA benchmarks agentic reasoning on industrial-grade financial documents, featuring 2,026 deep research tasks across...
OpenForgeRL: Open-Source Framework for Training Harness-Based AI Agents End-to-End
Hugging Face researchers introduce OpenForgeRL, an open-source framework that enables end-to-end reinforcement learning for AI age...
WorldWeaver: Streaming Multi-Agent Diffusion Model with Cross-Agent State Registers
Hugging Face researchers introduce WorldWeaver, a streaming multi-agent video diffusion model that uses learnable world state regi...
Hugging Face Researchers Propose Color Pass-Through: End-to-End Camera-Display Calibration
A new end-to-end learned framework treats camera and display as a coupled system, achieving over 2x improvement in color reproduct...
Hugging Face Researchers Propose Structured Dynamics Model to Disentangle Camera and Object Motion from Videos
A new self-supervised method, the Structured Dynamics Model (SDM), explicitly separates camera motion from object motion in videos...
LLMs Struggle to Track Evolving User Intent, New Study Finds
A new framework reveals that LLMs suffer significant performance drops when user intent evolves across conversation turns, exposin...
Tencent WorkBuddy Bench: A Contamination-Resistant Multi-Domain Coding-Agent Benchmark
Tencent introduces WorkBuddy Bench, a multi-domain benchmark for coding agents with a contamination-resistant task construction me...
NVIDIA and Hugging Face Introduce NOOA: Object-Oriented Agents in Native Python
NVIDIA and Hugging Face present NOOA, a Python framework that treats AI agents as native Python objects, merging prompts, tools, a...
SANA-Video 2.0: Hybrid Linear Attention Cuts Video Generation Cost, Matches Softmax Quality
NVIDIA and Hugging Face introduce SANA-Video 2.0, a hybrid video diffusion transformer that combines linear and softmax attention...
ProVisE: New Benchmark Lets Image-Generation Models Show Spatial Intelligence in Pixels
Researchers introduce ProVisE, a framework that evaluates image-generation models on spatial tasks by letting them answer directly...
VCSD: Visual Contrastive Self-Distillation Boosts VLMs Without External Teachers
Researchers propose Visual Contrastive Self-Distillation (VCSD), a novel on-policy self-distillation method that improves vision-l...
ReferTrack: A New Paradigm for Embodied Visual Tracking Achieves SOTA on EVT-Bench
Hugging Face researchers introduce ReferTrack, a referring-then-tracking paradigm that grounds embodied visual tracking using a si...