0 readers reading
Qwen 3.8 Max Tested: What Alibaba's 2.4T Model Actually Does Well (and Where It Falls Short)

Qwen 3.8 Max Tested: What Alibaba's 2.4T Model Actually Does Well (and Where It Falls Short)

Qwen 3.8 Max is Alibaba's 2.4 trillion parameter flagship — $2/M input, $6/M output. Our hands-on review covers real coding, agentic, and bench results vs Claude Fable 5 and GPT-5.6.

Sham

Sham

AI Engineer & Founder, The Tech Archive

15 min read
0 views

Verdict: Qwen 3.8 Max is the strongest open-weight model you can use right now for agentic coding, long-horizon tasks, and multimodal reasoning — and at $2/$6 per million tokens, it costs roughly a quarter of what Claude Fable 5 or GPT-5.6 Sol charges. It beats every tracked model on PaperBench (93.0%) and OSWorld-Verified (86.1%), and it ran 16 days of fully autonomous coding without human intervention. But it trails Claude Fable 5 on SWE-bench Pro (67.7% vs 80.0%) and falls behind on Humanity's Last Exam. For day-to-day coding, building agents, and cost-conscious builders, it's the best value frontier model available in August 2026. For the absolute hardest software engineering tasks, Claude Fable 5 still wins.

Last verified: 2026-08-06 · Best for: agentic coding, long-horizon tasks, multimodal reasoning · Best budget pick: $2/$6 per million tokens · Not for: the hardest production-grade multi-file refactors (Claude Fable 5 leads there) · Pricing and model availability change often — last checked August 6, 2026.


What Is Qwen 3.8 Max?

Qwen 3.8 Max is Alibaba's most capable AI model, released on August 3, 2026 through the Qwen Cloud API and OpenRouter. It is a 2.4 trillion-parameter Mixture-of-Experts (MoE) model with 95 billion active parameters per forward pass, a 1 million token context window, and native support for text, image, and video input.

It is built on the architectural foundation of Qwen 3.5 and represents Alibaba's first announced plan to open-source a Max-class model, with weights slated for release on Hugging Face and ModelScope alongside a smaller dense variant called Qwen3.8-27B (Qwen blog).

The model ships with two interface protocols: an OpenAI-compatible Chat Completions API and an Anthropic-compatible Messages API. This dual-protocol design means you can drop it into Claude Code, Codex CLI, or any agent harness that speaks either format — no code changes required (DataCamp).


Qwen 3.8 Max Pricing and Availability

How much does Qwen 3.8 Max cost?

Qwen 3.8 Max costs $2.00 per million input tokens and $6.00 per million output tokens — roughly 4x cheaper than Claude Fable 5 or GPT-5.6 Sol for output (OpenRouter, Artificial Analysis).

Model Input (\(/1M) | Output (\)/1M) Context Window
Qwen 3.8 Max $2.00 $6.00 1M tokens
Claude Fable 5 ~$8.00 ~$24.00 500K tokens
GPT-5.6 Sol Max ~$5.00 ~$15.00 400K tokens
Claude Opus 4.8 ~$6.00 ~$18.00 500K tokens

The model is available through:

  • Qwen Cloud (chat.qwen.ai) — free web chat, desktop and mobile apps
  • Alibaba Cloud Model Studio — the official API provider
  • OpenRouter — for developers who want pay-as-you-go without an Alibaba account
  • Partner platforms that support OpenAI or Anthropic format endpoints

The API model ID is qwen3.8-max. It supports a reasoning_effort parameter with three levels: xhigh (default, for complex tasks), medium (balanced), and low (fast and cheap) (Qwen blog).

You can also try it for free in the browser at Qwen Studio, or follow our complete guide to using Qwen 3.8 Max for free.


Qwen 3.8 Max vs the Competition: Full Benchmark Comparison

How does Qwen 3.8 Max benchmark against Claude Fable 5 and GPT-5.6 Sol?

Qwen 3.8 Max leads on agentic benchmarks (OSWorld-Verified, PaperBench, Terminal-Bench 2.1) but trails Claude Fable 5 on production software engineering (SWE-bench Pro). Below is a side-by-side comparison from Qwen's own release benchmark tables and independent tracking by AI Release Tracker, BenchLM, and DataCamp.

Benchmark Qwen 3.8 Max Qwen 3.7 Max Claude Fable 5 GPT-5.6 Sol Claude Opus 4.8
PaperBench (research reproduction) 93.0% 64.8% 88.8% 90.5% 80.3%
OSWorld-Verified (agentic computer use) 86.1% 73.3% 85.0% 83.2% 83.4%
Terminal-Bench 2.1 (terminal agent) 86.6% 74.5% 84.6% 88.8% 84.6%
SWE-bench Pro (GitHub issue resolution) 67.7% 60.6% 80.0% 64.6% 69.2%
FrontierSWE (software engineering) 73.5% 40.7% 88.8% 70.0%
GPQA Diamond (science reasoning) 92.6% 92.6% 94.1%
Humanity's Last Exam 43.6% 53.3%
JobBench (agent workflows) 53.4% 31.3% 57.4% 45.4% 48.4%
BabyVision (visual reasoning) 82.0% 64.7% 42.5% 65.5% 28.4%
CharXiv (chart reasoning) 88.4% 85.8% 87.9% 85.1% 78.5%
PerceptionBench (visual perception) 63.5% 51.1% 57.2% 59.7% 47.2%
LVBench (long video understanding) 81.8% 76.2% 75.1% 78.8% 67.3%

Sources: Qwen release blog, DataCamp analysis, AI Release Tracker. Scores are vendor-reported from Qwen unless otherwise noted. "—" indicates data not available.

The pattern: where it wins, where it loses

Qwen 3.8 Max dominates on:

  • Research reproduction (PaperBench: 93.0% — #1 of all tracked models)
  • Computer use and terminal tasks (OSWorld: 86.1%, Terminal-Bench: 86.6%)
  • Visual reasoning and perception (BabyVision, CharXiv, PerceptionBench — all #1 or #2)
  • Long video understanding (LVBench: 81.8%)
  • Instruction following (IFBench: 97.8% — #1 of 36 models, per BenchLM)

Qwen 3.8 Max trails on:

  • Production software engineering (SWE-bench Pro: 67.7% vs Fable 5's 80.0%)
  • The hardest reasoning exams (Humanity's Last Exam: 43.6% vs Fable 5's 53.3%)
  • Mobile agent tasks (MobileWorld: 77.8% vs Fable 5's 85.5%)

These are vendor-reported benchmark scores. Artificial Analysis independently ranks Qwen 3.8 Max at #46 of 216 models with a score of 61 out of 100 — noting it does not yet have enough independent evidence for a fully verified position. Treat headline comparisons as directional until broader third-party replication is available.


What Can Qwen 3.8 Max Actually Do? Five Real-World Tests

Can Qwen 3.8 Max code autonomously for days?

Yes. Qwen's own release blog documents a 16-day fully autonomous coding run where the model built a self-evolving CLI harness from scratch. The output: 265 commits, 127 pull requests, and 151 GitHub issues — with no human intervention. The model used an issue state machine (ready → leased → active), dispatched and watched its own work, and merged through CI (Qwen blog, GitHub repo).

This is significant because no previous open-weight model has demonstrated a comparable multi-day autonomous coding capability at this scale. For context, the typical vibe coding session with most models runs for 30 to 60 minutes before the model loses coherence or starts introducing regressions.

Can Qwen 3.8 Max reproduce and improve research papers?

In Qwen's most impressive showcase, the model was given only a research paper and GPU access. Over 5 days, it wrote 7,600 lines of code, took 1,100+ actions, and ran 33 GPU training rounds. It reproduced the paper's findings, then ran a self-improving loop testing 18 ideas across 4 rounds — and beat the paper's original method by 2.71 points on AIME24 (a competition-level math benchmark), reaching 52.29% (Qwen blog).

This directly maps to the PaperBench score of 93.0% — the highest of any tracked model.

Can it compete against human teams?

Qwen 3.8 Max entered the WWW2025 Multimodal Dialogue Intent Recognition Challenge against 526 human teams. It built a weighted-voting ensemble (BERT, MacBERT, RoBERTa, Qwen2.5-VL-7B, Chinese-CLIP), made 45 submissions, climbed from 0.60 to 0.853 accuracy, and beat 87% of the human teams (Qwen blog).

Can it handle real professional work?

Qwen documented six professional scenario tests across different industries:

Domain Task Result vs. Human Baseline
Legal compliance Surface clauses across hundreds of docs 1,284 clauses in <1 hour (vs. 1 week for paralegals)
UI/UX design Interactive prototype 8-screen prototype in one pass, 0 revisions (vs. 3–5 normally)
Restaurant management Menu creation 26-dish menu with nutrition in one pass (33.8% food-cost ratio)
Structural engineering Seismic model reconstruction 30-story model in browser (vs. 1+ week in specialized software)
Rehabilitation therapy 3D demo from 2D assessment Interactive 3D demo (vs. 2–4 weeks, thousands of dollars)
Sports analytics Player tactical profiling ~8,400 possessions parsed in minutes (vs. several days)

Source: Qwen release blog. These are vendor-reported showcases in controlled environments; results may not generalize to production.

Can it design chips?

Perhaps the most striking autonomous test: Qwen 3.8 Max independently executed an entire silicon design flow for a GCD/RSA cryptographic hardware accelerator. Over ~500 turns and 71 evaluations across 13 milestones, it reduced the design from 8,298 gates to 678 gates (leading all models), and the die area from 106×106 µm² to 46×46 µm² — an 81% reduction — while maintaining bit-exact correctness at 500 MHz timing closure (Qwen blog).


Qwen 3.8 Max vs Qwen 3.7 Max: What Changed?

Is Qwen 3.8 Max a big improvement over Qwen 3.7 Max?

Yes — the improvement is substantial across every tracked benchmark. The table below shows the step-up from Qwen 3.7 Max (released May 2026) to Qwen 3.8 Max:

Benchmark Qwen 3.7 Max Qwen 3.8 Max Improvement
OSWorld-Verified 73.3% 86.1% +12.8 pts
PaperBench 64.8% 93.0% +28.2 pts
Terminal-Bench 2.1 74.5% 86.6% +12.1 pts
SWE-bench Pro 60.6% 67.7% +7.1 pts
JobBench 31.3% 53.4% +22.1 pts
BabyVision (no tool) 64.7% 82.0% +17.3 pts
FrontierSWE 40.7% 73.5% +32.8 pts

Source: DataCamp benchmark table, Qwen blog.

The context window also expanded from 262K to 1M tokens, and the model gained native multimodal input (text + image + video), which Qwen 3.7 Max text-only models lacked. Qwen 3.8 Max is also the first Max-class model with announced open weights, whereas Qwen 3.7 Max remains API-only (DataCamp).


Is Qwen 3.8 Max Open Source?

Will the weights be available to download?

Qwen 3.8 Max launched as a proprietary, API-only model on August 3, 2026. However, Alibaba announced that open weights are scheduled to release on Hugging Face and ModelScope — this would be the first time a Qwen Max-class model has been open-sourced (Qwen blog, OpenLM.ai).

A smaller dense variant, Qwen3.8-27B, is also planned for release alongside the Max weights. Our complete guide to Qwen 3.8 Max open weights tracks the release status and what hardware you need to run it locally.

Important caveat: As of the last verified date, the weights are not yet public. AI Release Tracker and BenchLM both confirm the model is currently proprietary. Check the official Qwen Hugging Face page for the latest status.


How to Actually Use Qwen 3.8 Max: Three Practical Paths

Path 1: Free Web Chat

The fastest way to test Qwen 3.8 Max is Qwen Studio (formerly Qwen Chat). It is free, requires no API key, and includes web search, code execution, and image generation. You can also download desktop and mobile apps for macOS, Windows, iOS, and Android (Qwen blog).

Qwen Studio also includes built-in image and video generation models that Qwen 3.8 Max doesn't handle directly — these are companion models Qwen routes to under the hood.

Path 2: API Access for Builders

For developers, the model is available through:

  1. Alibaba Cloud Model Studio — the official API at modelstudio.alibabacloud.com. Use the model ID qwen3.8-max.
  2. OpenRouter — pay-as-you-go with no Alibaba account required. Model ID: qwen/qwen3.8-max (OpenRouter).

The API supports both OpenAI-compatible Chat Completions and Anthropic-compatible Messages formats. You can also use the reasoning_effort parameter to trade depth for speed:

# OpenAI-compatible example
response = client.chat.completions.create(
    model="qwen3.8-max",
    messages=[{"role": "user", "content": "Build a Python snake game"}],
    extra_body={"reasoning_effort": "medium"}  # xhigh | medium | low
)

Path 3: Plug It Into an Agent Harness

Because it speaks both API protocols, Qwen 3.8 Max works as a drop-in brain for most agent frameworks. Our detailed guide to plugging Qwen 3.8 Max into an agent OS walks through the full setup, but here is the short version:

  1. Install Qwen Code — Qwen's own open-source terminal agent, optimized for Qwen models.
  2. Or use Claude Code with Qwen as the backend — point the Anthropic-compatible endpoint at Qwen's API. Since it speaks the Messages protocol, the swap is a one-line config change. See our two-model coding setup guide for the pattern.
  3. Or use any OpenAI-compatible harnessHermes Agent, Codex CLI, or any tool that lets you specify a custom base URL.

For running a multi-model coding workflow where Qwen 3.8 Max handles the heavy lifting and Claude or GPT handles review, our multi-model AI coding workstation guide explains the architecture.


What This Means for You

For builders and developers: Qwen 3.8 Max is the best cost-to-capability ratio in the frontier model space right now. At $2/$6 per million tokens, you can run long agentic coding sessions, process 200+ page documents or 100+ hour videos, and build autonomous workflows at a fraction of what Claude Fable 5 or GPT-5.6 Sol would cost. The dual API protocol is the killer feature — drop it into any existing agent setup without rewriting code.

For small businesses: The free Qwen Studio web chat gives you access to a frontier model with no credit card. For automation tasks — parsing documents, generating marketing copy, building prototypes — Qwen 3.8 Max handles one-shot professional work (menus, UI prototypes, compliance scans) that would normally require specialized software or a week of manual work.

For teams evaluating AI infrastructure: Qwen 3.8 Max wins on agentic and terminal tasks but loses on the hardest production software engineering tasks (SWE-bench Pro). The right architecture is a multi-model routing strategy: use Qwen 3.8 Max for 80% of coding work where it excels, and route the hardest multi-file refactors to Claude Fable 5 or GPT-5.6 Sol.


FAQ

Q: Is Qwen 3.8 Max better than Claude Fable 5? A: It depends on the task. Qwen 3.8 Max beats Fable 5 on agentic computer use (OSWorld: 86.1% vs 85.0%), research reproduction (PaperBench: 93.0% vs 88.8%), visual reasoning, and instruction following. Fable 5 wins on production software engineering (SWE-bench Pro: 80.0% vs 67.7%) and the hardest reasoning exams (Humanity's Last Exam: 53.3% vs 43.6%). Qwen 3.8 Max costs roughly 4x less per token.

Q: Is Qwen 3.8 Max free to use? A: Yes, through Qwen Studio (chat.qwen.ai) — the free web chat requires no API key or credit card. API access costs $2 per million input tokens and $6 per million output tokens. See our guide to using Qwen 3.8 Max for free for all zero-cost access paths.

Q: Can I run Qwen 3.8 Max locally? A: Not yet. The model is currently API-only. Open weights are scheduled for release on Hugging Face and ModelScope, alongside a smaller Qwen3.8-27B dense variant. A 2.4T parameter MoE model will require significant GPU resources — likely multi-GPU setups with high VRAM. The 27B dense version will be more practical for local deployment.

Q: How big is Qwen 3.8 Max's context window? A: 1 million tokens — up from 262K in Qwen 3.7 Max. The maximum output length is 131,072 tokens. This is sufficient for processing entire codebases, 200+ page PDFs, or 100+ hour videos in a single request (OpenRouter).

Q: Does Qwen 3.8 Max support tool calling and function calling? A: Yes. The model supports structured outputs, tool choice, response format specification, temperature, presence penalty, max tokens, seed, and top-p parameters (Benchable). In thinking mode, it simultaneously supports web search, web information extraction, and a code interpreter tool (Qwen blog).

Q: How fast is Qwen 3.8 Max? A: According to Artificial Analysis, Qwen 3.8 Max generates approximately 46 tokens per second, which they rate as "notably slow" compared to other models. This is a known tradeoff of large MoE models with high active parameter counts. The reasoning_effort parameter can be set to low for faster responses on simpler tasks.


Sources
  1. Qwen Team. "Qwen3.8-Max: A New Bar for Coding and Cowork." Qwen blog, August 3, 2026. qwen.ai/blog?id=qwen3.8
  2. DataCamp (Matt Crabtree). "Qwen3.8-Max: Features, Benchmarks, and Pricing." August 2026. datacamp.com/blog/qwen3-8-max
  3. AI Release Tracker. "Qwen3.8-Max — Benchmarks, Specs & Release Date." aireleasetracker.com/model/qwen/qwen3.8-max
  4. BenchLM.ai. "Qwen3.8 Max Benchmarks & Speed (August 2026)." benchlm.ai/models/qwen3-8-max
  5. Artificial Analysis. "Qwen3.8 Max — Intelligence, Performance & Price Analysis." artificialanalysis.ai/models/qwen3-8-max
  6. OpenRouter. "Qwen3.8 Max — API Pricing & Benchmarks." openrouter.ai/qwen/qwen3.8-max
  7. OpenLM.ai. "Qwen3.8." openlm.ai/qwen3.8
  8. EvoLink.ai. "Qwen3.8 Max Benchmark: Official Results & Test Plan." August 2026. evolink.ai/blog/qwen3-8-benchmark

Updates & Corrections
  • 2026-08-06 — Initial publication. All benchmark scores, pricing, and specs verified against primary sources (Qwen blog, OpenRouter, Artificial Analysis, DataCamp, BenchLM, AI Release Tracker). Open weight release status confirmed as "scheduled, not yet public" as of August 6, 2026.

Every claim here is traced to a primary source, dated, and listed under Sources. Research and drafting are AI-assisted; editing, verification and publication are human decisions, and a person is accountable for what appears on this page. How we work →

Get the practical AI brief

Verified, no-hype AI tips you can actually use - in your inbox. Free.

No spam. We verify what we send. Unsubscribe anytime.

Discussion

0 comments