Latest stories
TimeLens2: A Generalist Video Temporal Grounding MLLM That Pinpoints When Events Occur
TimeLens2 treats temporal evidence as an interval set, using a new dataset and a temporal Wasserstein reward to achieve state-of-t...
Hugging Face Unveils RynnBrain 1.1: A Family of Embodied Foundation Models Up to 122B Parameters
RynnBrain 1.1, a new family of embodied foundation models from Hugging Face, scales from 2B to 122B parameters and introduces cont...
Hugging Face Unveils Cura 1T: A Specialized LLM for Agentic Healthcare
Cura 1T, a healthcare-specialized LLM trained via a human-gated self-evolution loop, achieves top scores on 5 of 6 hardest healthc...
xHC: Expanded Hyper-Connections Enable Residual Stream Scaling Beyond N=4
Hugging Face researchers propose xHC, a method that expands the residual stream of Transformers to N=16 parallel streams, overcomi...
Xiaomi-Robotics-1: A VLA Foundation Model Trained on Over 100K Hours of Real-World Data
Xiaomi introduces Xiaomi-Robotics-1, a vision-language-action foundation model pre-trained on over 100,000 hours of real-world man...
Hugging Face Unveils Loopie: A Breakthrough in Looped Transformers with MoE
Hugging Face introduces Loopie, a series of looped Mixture-of-Experts Transformer models that outperform vanilla baselines under t...
RESOURCE2SKILL: Turning Tutorial Videos and Multimodal Resources into Executable Agent Skills
Microsoft and Hugging Face introduce RESOURCE2SKILL, a framework that distills multimodal resources like tutorial videos, code rep...
RAGU: Open-Source GraphRAG Engine Uses Compact 7B Model to Outperform Larger LLMs
Researchers introduce RAGU, a modular GraphRAG engine that separates extraction from consolidation, and Meno-Lite-0.1, a 7B model...
Qwen3.5-Max-Preview Debuts on Arena with Strong Preliminary Results
Alibaba's Qwen team has released the preview of Qwen3.5-Max on the Arena platform, showcasing impressive performance in early eval...
Qwen-Image-Edit-2511: Enhanced Consistency and LoRA Integration
Alibaba's Qwen team releases an improved image editing model with better character consistency, multi-person group photo fusion, b...
Alibaba’s Qwen3-TTS Introduces Voice Cloning and Voice Design Models
Qwen3-TTS family expands with two new models: Qwen3-TTS-VD-Flash for voice design via natural language instructions and Qwen3-TTS-...
Qwen-Image-Layered: Decomposing Images into Editable RGBA Layers
Alibaba's Qwen introduces a model that decomposes images into multiple RGBA layers, enabling independent manipulation of each laye...