Latest stories
Moonshot AI Unveils WorldVQA Benchmark to Test Visual World Knowledge in Multimodal LLMs
Moonshot AI releases WorldVQA, a benchmark with 3,500 image-question pairs designed to measure factual visual knowledge in multimo...
Kimi Launches Agent Swarm: 100 AI Agents Self-Organize to Tackle Complex Tasks
Kimi (Moonshot) unveils Agent Swarm, a research preview that lets K2.5 deploy up to 100 parallel sub-agents that self-organize int...
Moonshot AI Open-Sources Kimi K2.6 with Advanced Coding and Agent Swarm Capabilities
Moonshot AI has released Kimi K2.6, an open-source model featuring state-of-the-art coding, long-horizon execution, and agent swar...
Moonshot AI Launches PerceptionBench to Isolate and Measure Atomic Visual Perception in MLLMs
Moonshot AI releases PerceptionBench, a new benchmark that evaluates multimodal large language models on ten atomic visual percept...
Moonshot AI Unveils Kimi K3: World's First Open 3T-Class Model with 2.8T Parameters
Moonshot AI releases Kimi K3, a 2.8T-parameter open model with native vision and a 1M-token context, claiming frontier-level perfo...
OpenAI Unveils 'The Defender's Window': A New Framework for AI-Driven Cybersecurity
OpenAI highlights the dual-edged nature of AI in cybersecurity, urging defenders to adopt proactive measures and leverage AI to st...
OpenAI Joins PORTS-Pike Project to Boost Southern Ohio Jobs
OpenAI has joined the PORTS-Pike project, expanding its community investment and supporting thousands of jobs in Southern Ohio.
OpenAI Backs 14 Independent AI Policy Projects for the Intelligence Age
OpenAI is funding 14 independent projects to explore new AI policy ideas aimed at expanding economic opportunity and strengthening...
MMDiff: New Framework Lets Researchers Isolate and Control Features in Multimodal AI Models
Researchers introduce MMDiff, a framework using multimodal sparse autoencoders to identify and control specific features in multim...
Hugging Face Study: Optimal Data Repetition Scales Mildly with LLM Size
A new paper from Hugging Face reveals that under proportional scaling of model size and training tokens, the optimal repetition of...
Hugging Face Paper: Claim-Level Verification Boosts Reasoning Efficiency
A new training-free method, Claim-Level Reliability Assessment (CLR), improves LLM reasoning accuracy by verifying critical claims...
Second Thought: Parallel Reasoning During Agent Idle Time Cuts Sequential Decoding by Up to 43%
A new training-free framework from Hugging Face, Second Thought, runs auxiliary reasoning branches in parallel while LLM agents wa...