Benchmark score rise is verified by Artificial Analysis, but internal recursive self-improvement mechanisms remain vendor-reported without independent audit
Alibaba claims AI upgraded itself 33 times
An independent test confirmed a score rise to 45, but the automated learning process behind it remains unverified.
In a nutshell
Alibaba claimed a breakthrough in self-improving artificial intelligence after its Qwen3.8-Max model lifted its independent benchmark score from 40 to 45, but while the higher score is publicly confirmed, outside analysts note the 33 unassisted training cycles behind it remain an unverified corporate claim.
Highlights
- Alibaba said Qwen3.8-Max completed 33 automated self-improvement cycles within one month.
- Independent benchmarking service Artificial Analysis confirmed the model's score rose from 40 to 45.
- The updated score returned Qwen3.8-Max to top ranking among domestic Chinese models.
- The software completed 60 hours of unassisted optimization to shrink a chip module by 42%.
- No third-party observer or technical whitepaper has confirmed the automated training mechanism.
Qwen3.8-Max Performance and Engineering Disclosures
| Metric | Value |
|---|---|
| Artificial Analysis score before test | 40 |
| Artificial Analysis score after test | 45 |
| Reported autonomous self-improvement cycles | 33 |
| Autonomous chip optimization duration | 60 hours |
| Autonomous tool calls in chip design | 10000 |
| Chip bus module area reduction | 42% |
- Confirmed benchmark gains alongside vendor-reported figures from Alibaba's software runs.
- Artificial Analysis — An independent platform that evaluates and rates artificial intelligence models.
- Shows the gap between publicly audited performance metrics and unverified internal engineering claims.
From the Editor’s Diary
A confirmed benchmark score proves software has improved, but without open test data, claims that the machine taught itself remain unproven corporate marketing.
Who's involved
Alibaba Group
Chinese multinational technology company developing the Qwen family of artificial intelligence models
goal → Aiming to lead global computing by proving its software can reason and teach itself
Eddie Wu
Chief Executive Officer of Alibaba Group
goal → Directing company strategy across computer chips, cloud networks, and advanced intelligence models
Artificial Analysis
Independent evaluation service that tests and scores artificial intelligence systems
goal → Publishing standard performance ratings across public and commercial software models
Kyle Chan
Technology researcher studying Chinese artificial intelligence developments
goal → Reviewing technical claims and engineering disclosures from Chinese computing laboratories
In short
Alibaba Group says its flagship artificial intelligence software upgraded its own code 33 times in one month without human direction, claiming a leap toward computers that teach themselves.
The claim makes an industry-wide push toward fully autonomous software development the most likely next step in the global technology race.
Whether that breakthrough actually happened is too early to tell: an independent rating confirmed the model gained points, but only Alibaba's own word supports the claim that the machine did the training on its own.
Previously in this story
Alibaba falls behind rivals in benchmark test
7 September 2026An independent test places the group's flagship model behind Chinese rivals Zhipu AI and Moonshot AI on automated tasks.
How it unfolded
Alibaba Unveils Automated AI Training Strategy
At the company's Apsara Conference in Hangzhou, chief executive Eddie Wu and cloud executives laid out their long-term computing plans, pointing to software loops where Qwen3.8-Max completed 33 iterative improvement cycles in one month.
Independent Evaluators Confirm Score but Question Method
Outside researchers verified that the updated model snapshot earned a score of 45 on the independent Artificial Analysis ranking, returning it to first place among Chinese domestic systems, but technical observers warned that the 33-cycle autonomous mechanism rests entirely on Alibaba's corporate word without public logs or outside replication.
Where things stand
The story sits between a verified test result and an unproven training claim. Artificial Analysis lists the September snapshot of Qwen3.8-Max at 45 on its version 4.3.2 test index, confirming the software performs better than previous editions.
What remains unanswered is whether the system genuinely improved its own core design or merely ran automated fine-tuning guided by human-written rules. Researchers are waiting for a technical paper from Alibaba explaining the compute budget, safety boundaries, and step-by-step test data.