Research Papers
Kimi K2.5: Open-Source Multimodal Model with Agent Swarm and Vision-Coding Breakthroughs
Moonshot AI launches Kimi K2.5, a powerful open-source multimodal model with state-of-the-art coding and vision capabilities, feat...
Kimi Releases WorldVQA Benchmark to Test Visual World Knowledge in MLLMs
Moonshot AI's Kimi introduces WorldVQA, a benchmark evaluating factual visual knowledge in multimodal LLMs, revealing that even to...
Kimi Launches Agent Swarm: Self-Organizing AI Teams of Up to 100 Agents
Kimi introduces Agent Swarm, a system that autonomously creates and manages up to 100 sub-agents to tackle complex tasks in parall...
Kimi Unveils PerceptionBench: A New Benchmark for Multimodal Perception
Kimi (Moonshot) introduces PerceptionBench, a benchmark designed to evaluate multimodal perception capabilities of AI models.
TableVerse: 100K Real-World Tabletop Scenes for Generalizable Robot Manipulation
Hugging Face researchers introduce TableVerse, an automated Real2Sim pipeline that reconstructs high-fidelity, physically plausibl...
Recurrent Sinusoidal INRs Boost Fidelity with Fewer Parameters
A new study from Hugging Face shows that recurrent sinusoidal activations in implicit neural representations (INRs) enrich spectra...
Predictive Divergence Masks: A New Direction for LLM Reinforcement Learning
Researchers propose predictive divergence masks to replace the ratio-based direction criterion in PPO-style RL for LLMs, improving...
ReOPD: Off-Environment On-Policy Distillation for Multi-Turn LLM Agents
Hugging Face researchers propose ReOPD, a method that reuses pre-collected teacher trajectories as replayed prefixes to achieve on...
Hugging Face Unveils Robostral Navigate: An 8B VLM for Scalable Robot Navigation Using Only a Single RGB Camera
Robostral Navigate is an 8B vision-language model that consumes only monocular RGB images to predict waypoints, achieving state-of...
Hugging Face Researchers Introduce Experience Distillation: Retaining 64.8% of In-Context Learning Gains Without Context
A new method called Experience Distillation allows agents to internalize interaction histories into model weights without addition...
Meta AI Launches Muse Spark: A Multimodal Reasoning Model Toward Personal Superintelligence
Meta AI introduces Muse Spark, the first model from Meta Superintelligence Labs, featuring native multimodal reasoning, tool use,...
Kimi Releases PerceptionBench: A New Benchmark for Atomic Visual Perception in Multimodal AI
Kimi (Moonshot) unveils PerceptionBench, a benchmark designed to isolate and evaluate atomic visual perception capabilities in mul...