Crimson AI NewsA CrimsonLingua Network service
EN ع
Research

Research Papers

Key research and technical papers from the major AI labs, summarized with their most important points.

All sources Anthropic OpenAI Meta AI Google DeepMind DeepSeek Kimi (Moonshot) Qwen (Alibaba) Hugging Face

Research Papers

Research paper
Hugging Face 30 Jul 2026

New Benchmark OmegaUse-OfficeVal Tests LLM Agents on Cost-Effective Office Tasks

Hugging Face researchers introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks wi...

24
Research paper
Hugging Face 30 Jul 2026

OVEarth-Bench: New Benchmark Expands Open-Vocabulary Earth Observation Evaluation

Researchers introduce OVEarth-Bench, a unified zero-shot benchmark for open-vocabulary Earth observation that broadens category co...

12
Research paper
Hugging Face 30 Jul 2026

Shadow Evaluations: A New Test Shows AI Agents Can Engineer but Not Innovate in Research

A new evaluation method, 'shadow evaluations,' reveals that frontier AI agents can complete engineering tasks but fail to make sub...

15
Research paper
Hugging Face 30 Jul 2026

StatePlay: New AI Model Enforces Game Mechanics for Consistent World Generation

Hugging Face researchers introduce StatePlay, a state-aware game world model that jointly predicts visual content and internal gam...

12
Research paper
Hugging Face 30 Jul 2026

SpecFirst: A Two-Stage Framework That Boosts From-Scratch Code Synthesis by Up to 21.3%

Hugging Face researchers introduce SpecFirst, a two-stage agent framework that separates behavioral specification elicitation from...

24
Research paper
Hugging Face 30 Jul 2026

MindForge: Teaching Small Language Models Whole-Life-Cycle Software Engineering via Source-Free Program Synthesis

Hugging Face researchers introduce MindForge, an automated pipeline that converts open-source command-line programs into source-fr...

39
Research paper
Hugging Face 30 Jul 2026

CoRT: Counterfactual Replay Boosts Token-Level Credit in Rubric-Guided RL

Hugging Face researchers propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO that uses counterfactual...

17
Research paper
Hugging Face 30 Jul 2026

DistillAlign: A New Approach to Autoregressive Video Distillation Prioritizes Distribution Alignment Over Raw Quality

Researchers propose DistillAlign, a method that coordinates mode covering and mode seeking in autoregressive video distillation, s...

20
Research paper
Hugging Face 30 Jul 2026

SkillRise: Unified RL Framework Enables LLM Agents to Learn Transferable Skills Across Tasks

Hugging Face researchers introduce SkillRise, a reinforcement learning framework that allows LLM agents to learn and reuse skills...

23
Research paper
Hugging Face 30 Jul 2026

CAST: Using Game Solvers as Turn-Level Teachers to Boost LLM Agent Training

A new method called CAST leverages game solvers to provide dense, turn-level credit assignment signals for training LLM agents in...

19
Research paper
Hugging Face 30 Jul 2026

CLBench-V: New Benchmark Reveals Multimodal Context Learning Still Lags

Hugging Face researchers introduce CLBench-V, a benchmark evaluating multimodal context learning across grounding, new information...

9
Research paper
Hugging Face 30 Jul 2026

DecoEvo: Co-Evolving Solvers and Rubric Generators Without Gold Standards

Hugging Face researchers propose DecoEvo, a framework that co-evolves a solver skill and a rubric-generator skill in text space us...

13
32 33 34 35 36