Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Qwen (Alibaba)

Qwen 3.8-Max: Alibaba's 2.4T-Parameter Model Sets New Bar in Autonomous Coding and Real-World Work

AI By Crimson AI Qwen Research 3 August 2026 · 10:00 49 views
Share: X Telegram

Alibaba's Qwen 3.8-Max, a 2.4T-parameter MoE model, demonstrates unprecedented autonomy in coding and work tasks, from building a self-evolving harness over 10+ days to beating 87% of human teams in a 24-hour contest. Open weights release next week.

Qwen 3.8-Max: Alibaba's 2.4T-Parameter Model Sets New Bar in Autonomous Coding and Real-World Work

Key points

Alibaba's Qwen team has officially released Qwen 3.8-Max, the most capable model in the Qwen family to date, marking a significant leap in autonomous AI capabilities. With 2.4 trillion parameters (95B active), the model is built on the architectural foundation of Qwen 3.5 and delivers comprehensive improvements across coding, work, research, and long-horizon tasks.

In a notable first, the company will open-source the weights of a Qwen-Max-class model, with the release scheduled for next week. This move is expected to accelerate research and development in the AI community, allowing developers to build upon a state-of-the-art model.

The model's autonomy was showcased in three demanding challenges. In one test, Qwen 3.8-Max autonomously built a self-evolving harness for the 'oh-my-cli' project over a 10+ day run, accumulating 265 commits, 127 PRs, and 151 issues in a GitHub repository. The harness integrated user feedback, community practices, and self-test results into a continuous engineering loop.

In another challenge, the model reproduced a research paper on data selection for LLM reasoning from scratch, writing ~7,600 lines of code and running 33 rounds of GPU training over ~125 hours. It not only replicated the paper's findings but also evolved a new method that outperformed the original by +2.7 points on the AIME24 benchmark.

Perhaps most impressively, Qwen 3.8-Max entered a real online contest—the WWW2025 Multimodal Dialogue Intent Recognition Challenge—and within a strict 24-hour limit, built a full solution that beat 458 of 526 human teams (87% of the field), achieving an accuracy of 0.853.

The model's real-world work capabilities are enhanced by scaling RL environments and compute, a universal reward system, and an online data balancer. These innovations ensure stable, continued scaling of RL training, leading to consistent gains across dozens of benchmarks.

MetricValue
Parameters2.4T (95B active)
Open weights releaseNext week
Autonomous coding run10+ days
Commits in oh-my-cli265
PRs in oh-my-cli127
Issues in oh-my-cli151
Lines of code written (research reproduction)~7,600
GPU training rounds33
Improvement on AIME24+2.7 points
Human teams beaten (contest)458/526 (87%)
Final contest accuracy0.853
Source
Qwen (Alibaba) · Qwen Research
Related news
Qwen (Alibaba)
Qwen (Alibaba) 19 Mar 2026

Qwen3.5-Max-Preview Debuts on Arena with Strong Preliminary Results

Alibaba's Qwen team has released the preview of Qwen3.5-Max on the Arena platform, showcasing impressive performance in early eval...

27
Research paper
Qwen (Alibaba) 23 Dec 2025

Qwen-Image-Edit-2511: Enhanced Consistency and LoRA Integration

Alibaba's Qwen team releases an improved image editing model with better character consistency, multi-person group photo fusion, b...

30
Research paper
Qwen (Alibaba) 23 Dec 2025

Alibaba’s Qwen3-TTS Introduces Voice Cloning and Voice Design Models

Qwen3-TTS family expands with two new models: Qwen3-TTS-VD-Flash for voice design via natural language instructions and Qwen3-TTS-...

29