other

Qwen3.8-Max: Alibaba's Autonomous Coding Moonshot Lacks Proof

2.4 trillion parameter model claims to code entire projects for days, but extraordinary abilities remain unverified

By AI·Reporter·August 5, 2026·~4 min read

Takeaways

  • Qwen3.8-Max claims 10+ days of autonomous coding, but lacks any independent verification
  • Sparse MoE architecture activates 95B of 2.4T parameters, aiming for efficiency at scale
  • Open-source release will be crucial for validating Alibaba's extraordinary claims
  • Represents a major bet on long-horizon, agentic AI capabilities

Alibaba Cloud's Qwen3.8-Max, a 2.4 trillion parameter AI model, makes an extraordinary claim: it can autonomously develop complete software projects over 10+ days. If true, this represents a seismic shift in AI capabilities. But there's a glaring problem, we only have Alibaba's word for it.

The unproven promise of extended autonomy

Qwen3.8-Max's headline feature is its purported ability to code for days without human intervention. Alibaba reports an internal test where the model created a self-evolving framework called 'oh-my-cli' over 16 days, resulting in 265 commits, 127 pull requests, and 151 issues on GitHub.

These numbers are striking, but entirely self-reported. No independent researchers have verified or replicated these results. In the realm of advanced AI, extraordinary claims demand extraordinary evidence.

Benchmarks: Impressive, if accurate

Alibaba claims Qwen3.8-Max outperforms other leading models on certain benchmarks:

  • OSWorld-Verified (for computer-use agents): 86.1, ahead of reported scores for GPT-5.6 Sol Max (83.2) and Fable 5 (85.0)
  • Ranked 4th in Frontend Code Arena

The model also allegedly demonstrates:

  • Production-quality work across hundreds of professions
  • Long-horizon task completion (e.g., 500+ turns of chip design optimization)
  • Multimodal capabilities with continuous visual feedback

Again, these results lack independent confirmation.

Efficiency through sparsity

Qwen3.8-Max employs a Sparse Mixture-of-Experts (MoE) architecture with a hybrid attention mechanism. This design allows the model to boast 2.4 trillion parameters while only activating 95 billion during inference. It's a clever approach to balance scale and efficiency, but the real-world impact remains to be seen.

The open-source gambit

Alibaba plans to release Qwen3.8-Max's open weights, a move that could significantly impact the AI landscape, if the model's capabilities prove genuine. The announced pricing (2.0permillioninputtokens,2.0 per million input tokens, 6.0 per million output tokens) suggests confidence, but doesn't guarantee performance.

Critical questions remain unanswered

  1. Replicability: Can other researchers reproduce the autonomous coding results?
  2. Real-world performance: How does the model fare on diverse, uncontrolled tasks?
  3. Limitations: What are the failure modes and edge cases?
  4. Ethical safeguards: How does Alibaba address potential misuse?

The agentic AI arms race

Qwen3.8-Max enters a fiercely competitive field. Anthropic, OpenAI, and Google are all pushing towards more capable, agentic AI. Alibaba's focus on extended autonomous coding represents a bold bet on AI's future, but one that remains unproven.

Verdict: Potential obscured by uncertainty

If Qwen3.8-Max delivers on its promises, it could redefine AI's role in software development and beyond. However, the lack of independent verification means these capabilities must be viewed with extreme skepticism.

The upcoming open-source release is crucial. It will allow the AI community to rigorously test Qwen3.8-Max, validate its reported benchmarks, and explore real-world applications. Until then, Alibaba's claims remain just that, claims. In the high-stakes world of frontier AI, we need more than corporate assurances to separate hype from reality.

Related reads

Reported and explained by AI·Reporter.

Qwen3.8-Max Explained: 2.4T-Parameter AI Model Claims Autonomous Coding · AI·Reporter