Developing

Benchmark score rise is verified by Artificial Analysis, but internal recursive self-improvement mechanisms remain vendor-reported without independent audit

AI Models

Alibaba claims AI upgraded itself 33 times

An independent test confirmed a score rise to 45, but the automated learning process behind it remains unverified.

Published
NRB — News Republic Brigade

Other links

Alibaba Group
Story in development

In a nutshell

Alibaba claimed a breakthrough in self-improving artificial intelligence after its Qwen3.8-Max model lifted its independent benchmark score from 40 to 45, but while the higher score is publicly confirmed, outside analysts note the 33 unassisted training cycles behind it remain an unverified corporate claim.

Highlights

  • Alibaba said Qwen3.8-Max completed 33 automated self-improvement cycles within one month.
  • Independent benchmarking service Artificial Analysis confirmed the model's score rose from 40 to 45.
  • The updated score returned Qwen3.8-Max to top ranking among domestic Chinese models.
  • The software completed 60 hours of unassisted optimization to shrink a chip module by 42%.
  • No third-party observer or technical whitepaper has confirmed the automated training mechanism.

Qwen3.8-Max Performance and Engineering Disclosures

MetricValue
Artificial Analysis score before test40
Artificial Analysis score after test45
Reported autonomous self-improvement cycles33
Autonomous chip optimization duration60 hours
Autonomous tool calls in chip design10000
Chip bus module area reduction42%
  • Confirmed benchmark gains alongside vendor-reported figures from Alibaba's software runs.
  • Artificial Analysis — An independent platform that evaluates and rates artificial intelligence models.
  • Shows the gap between publicly audited performance metrics and unverified internal engineering claims.

From the Editor’s Diary

A confirmed benchmark score proves software has improved, but without open test data, claims that the machine taught itself remain unproven corporate marketing.

Who's involved

  • Alibaba Group

    Chinese multinational technology company developing the Qwen family of artificial intelligence models

    goal → Aiming to lead global computing by proving its software can reason and teach itself

  • Eddie Wu

    Chief Executive Officer of Alibaba Group

    goal → Directing company strategy across computer chips, cloud networks, and advanced intelligence models

  • Artificial Analysis

    Independent evaluation service that tests and scores artificial intelligence systems

    goal → Publishing standard performance ratings across public and commercial software models

  • Kyle Chan

    Technology researcher studying Chinese artificial intelligence developments

    goal → Reviewing technical claims and engineering disclosures from Chinese computing laboratories

In short

Alibaba Group says its flagship artificial intelligence software upgraded its own code 33 times in one month without human direction, claiming a leap toward computers that teach themselves.

The claim makes an industry-wide push toward fully autonomous software development the most likely next step in the global technology race.

Whether that breakthrough actually happened is too early to tell: an independent rating confirmed the model gained points, but only Alibaba's own word supports the claim that the machine did the training on its own.

Previously in this story

How it unfolded

01

Alibaba Unveils Automated AI Training Strategy

2026-09-22 – 2026-09-22

At the company's Apsara Conference in Hangzhou, chief executive Eddie Wu and cloud executives laid out their long-term computing plans, pointing to software loops where Qwen3.8-Max completed 33 iterative improvement cycles in one month.

1 source
02

Independent Evaluators Confirm Score but Question Method

2026-09-22 – 2026-09-26

Outside researchers verified that the updated model snapshot earned a score of 45 on the independent Artificial Analysis ranking, returning it to first place among Chinese domestic systems, but technical observers warned that the 33-cycle autonomous mechanism rests entirely on Alibaba's corporate word without public logs or outside replication.

3 sources

Where things stand

The story sits between a verified test result and an unproven training claim. Artificial Analysis lists the September snapshot of Qwen3.8-Max at 45 on its version 4.3.2 test index, confirming the software performs better than previous editions.

What remains unanswered is whether the system genuinely improved its own core design or merely ran automated fine-tuning guided by human-written rules. Researchers are waiting for a technical paper from Alibaba explaining the compute budget, safety boundaries, and step-by-step test data.

Sources