Latest stories
OpenAI Disrupts Cambodia-Based Scam Operation Using ChatGPT
OpenAI has disrupted a criminal scam operation based in Cambodia that leveraged ChatGPT to facilitate investment, romance, gamblin...
Manus Integrates ElevenLabs to Add Voice to AI Tasks
Manus, now part of Meta, announces a new integration with ElevenLabs, enabling users to add realistic voice and sound to their AI-...
Qwen 3.8-Max: Alibaba's 2.4T-Parameter Model Sets New Bar in Autonomous Coding and Real-World Work
Alibaba's Qwen 3.8-Max, a 2.4T-parameter MoE model, demonstrates unprecedented autonomy in coding and work tasks, from building a...
OpenAI Unveils GPT-Live: Real-Time Voice AI with Turnless Speech Model
OpenAI introduces GPT-Live, a realtime voice interaction system built on a turnless speech model and low-latency architecture, ena...
Hugging Face Research: CSCR Reallocates Token Credit to Improve Long-CoT Reasoning
A new paper proposes Counterfactual Sensitivity Credit Reallocation (CSCR), a simple extension of GRPO that reduces credit for hig...
EMBL AI Librarian: A Natural-Language Knowledge Layer for Life-Science Agents
Hugging Face researchers unveil EMBL AI Librarian, a knowledge layer that lets life-science AI agents query Europe PMC in natural...
RL^2-VLA: Adaptive Steering Boosts VLA Robots in Out-of-Domain Tasks
Researchers introduce RL^2, an adaptive inference-time steering framework that applies reinforcement learning on VLA latents only...
ODEWorld: Continuous-Time World Modeling via Physical-Time Flow
Researchers introduce ODEWorld, a continuous-time latent world model that learns an ODE-based velocity field in physical time, ena...
Hugging Face Researchers Introduce EVR: A New Reward Model for Consistent Multi-Reference Image Editing
A new paper from Hugging Face proposes the Multi-dimensional Evaluation-Verification Reward (EVR) to improve multi-reference image...
ExtractBench: New Benchmark Measures Accuracy, Cost, and Grounding in Enterprise Document Extraction
Hugging Face researchers introduce ExtractBench, the first benchmark to jointly score value accuracy, record completeness, groundi...
CriPO: Self-Distillation Fixes Two Hidden Failure Modes in Rubric-Based RL
Hugging Face researchers introduce Criterion-Distilled Policy Optimization (CriPO), an on-policy framework that tackles both Unexp...
Hugging Face Researchers Introduce CAPA: A Benchmark for Personalized Ambiguity Adaptation in Coding Assistants
A new benchmark, CAPA, evaluates how well coding assistants adapt to recurring, user-specific ambiguities across sessions, aiming...