Research Papers
Tencent WorkBuddy Bench: A Contamination-Resistant Multi-Domain Coding-Agent Benchmark
Tencent introduces WorkBuddy Bench, a multi-domain benchmark for coding agents with a contamination-resistant task construction me...
NVIDIA and Hugging Face Introduce NOOA: Object-Oriented Agents in Native Python
NVIDIA and Hugging Face present NOOA, a Python framework that treats AI agents as native Python objects, merging prompts, tools, a...
SANA-Video 2.0: Hybrid Linear Attention Cuts Video Generation Cost, Matches Softmax Quality
NVIDIA and Hugging Face introduce SANA-Video 2.0, a hybrid video diffusion transformer that combines linear and softmax attention...
ProVisE: New Benchmark Lets Image-Generation Models Show Spatial Intelligence in Pixels
Researchers introduce ProVisE, a framework that evaluates image-generation models on spatial tasks by letting them answer directly...
VCSD: Visual Contrastive Self-Distillation Boosts VLMs Without External Teachers
Researchers propose Visual Contrastive Self-Distillation (VCSD), a novel on-policy self-distillation method that improves vision-l...
ReferTrack: A New Paradigm for Embodied Visual Tracking Achieves SOTA on EVT-Bench
Hugging Face researchers introduce ReferTrack, a referring-then-tracking paradigm that grounds embodied visual tracking using a si...
K12-KGraph: Curriculum-Aligned Knowledge Graph Boosts Educational LLMs
Researchers introduce K12-KGraph, a curriculum-aligned knowledge graph from Chinese textbooks, along with a 23,640-question benchm...
Hugging Face Unveils AREX: A Recursively Self-Improving AI Agent for Deep Research
Hugging Face introduces AREX, a family of recursively self-improving agents that alternate between evidence gathering and constrai...
Anchor-Align: New Method Boosts VLA Robot Generalization by Preserving Pretrained Representations
Researchers propose Anchor-Align, a finetuning method for vision-language-action (VLA) policies that prevents representation drift...
New Scaling Laws Reveal Hypernetworks as a Scalable Alternative for Knowledge Injection in LLMs
A new study from Hugging Face introduces scaling laws for hypernetwork-based knowledge injection in LLMs, demonstrating that hyper...
ActiveVision Benchmark Reveals MLLMs Collapse on Active Observation Tasks
A new benchmark, ActiveVision, tests whether multimodal large language models (MLLMs) can perform active visual observation—repeat...
Hugging Face Paper Proposes Rubric-Oriented Document Set Selection Beyond Relevance
A new framework, SetwiseEvalKit, evaluates document sets as a whole—measuring coverage, conflict, and complementarity—rather than...