Latest stories
OpenAI's GPT-5.6: A Builder's Guide to Faster, Cheaper AI Agents
OpenAI's new guide highlights how startups leverage GPT-5.6 for faster, cost-efficient AI agents through smarter model selection a...
OpenAI unveils Ultrafast API tier: GPT-5.6 Sol up to 14x faster
OpenAI previews Ultrafast, a new API service tier that runs GPT-5.6 Sol at up to 14x speed, delivering up to 750 output tokens per...
OpenAI Names Dali Rajic as Chief Revenue Officer
OpenAI has appointed Dali Rajic as its new Chief Revenue Officer to lead global revenue efforts and help businesses maximize AI va...
Simplax: A New Augmentation for Uniform Discrete Diffusion Models
Researchers introduce Simplax, an exact Dirichlet-categorical augmentation that enriches uniform discrete diffusion without alteri...
Hugging Face Researchers Unveil GazeAnywhere: Promptable Gaze Estimation with Text or Visual Cues
A new end-to-end transformer model, GazeAnywhere, enables gaze target estimation using natural language or visual prompts, elimina...
Hugging Face Researchers Launch MBA-Bench: A Multimodal Benchmark for Business Ideation Agents
A new benchmark, MBA-Bench, evaluates multimodal AI agents on business ideation, with proposed models MBA-b and MBA-k outperformin...
Persistent Project Worlds Enable Autonomous Software Evolution, New EvoX Genesis Approach Shows
A new research paper introduces EvoX Genesis, which organizes long-horizon software development around a persistent project rather...
Hugging Face Study Exposes 'Illusion' in Visual Tool-Use by Multimodal LLMs
A new causal audit reveals that visual tool-use in multimodal LLMs often fails to causally influence answers, despite aggregate ac...
Hugging Face Researchers Unveil LDR: A Video World Model That Extrapolates Physics Beyond Training Data
A new paper from Hugging Face introduces Latent Dynamics Reasoning (LDR), a video world model that integrates kinematic dynamics i...
Hugging Face Researchers Unveil Closed-Loop Framework for Video Reflection Removal
A new closed-loop framework combines physics-based video synthesis, diffusion-based dereflection, and a dedicated benchmark, achie...
ToolHazard: New Framework Scales Adversarial Testing for LLM Agents Against Indirect Prompt Injections
Researchers introduce ToolHazard, a scalable framework that automatically synthesizes adversarial environments to expose and mitig...
New Benchmark Puts AI Coding Agents to the Test as Autonomous World-Model Researchers
Hugging Face researchers introduce AutoWorldModel-Bench, a closed-loop benchmark that evaluates frontier coding agents on open-end...