2026-08-03Model Release

Qwen3.8-Max Now Available on QwenCloud for Long-Horizon Coding, Professional Work, and Multimodal Tasks

With 2.4 trillion total parameters, 95 billion active parameters, and a 1-million-token context window, the MoE model is built for long-horizon coding, professional work, multimodal understanding, and reliable end-to-end task delivery.

1. Overview

Qwen3.8-Max is now available on QwenCloud. Built on the Qwen3.5 architecture, the model uses a Mixture-of-Experts design with 2.4 trillion total parameters and 95 billion active parameters. It accepts text, image, and video input and supports a 1-million-token context window. Across coding, professional work, research, and long-horizon tasks, Qwen3.8-Max is designed to plan multi-step work, use tools, respond to execution feedback, verify results, and deliver code, reports, analyses, and interactive applications.

Qwen3.8-Max will also be the first Qwen Max-class model with open weights. The weights are planned for release next week.

Read more: Qwen3.8 full release article

image.png

Figure 1. Qwen3.8-Max performance overview.

2. Model Details

Qwen3.8-Max: A Flagship Model for End-to-End Delivery

Model name / version: Qwen3.8-Max (qwen3.8-max)

Positioning: A flagship model for long-horizon coding, professional work, research, and multimodal tasks, covering planning, tool use, execution, and verification across long contexts.

Specifications

ItemQwenCloud
Input modalitiesText, image, video
Output modalityText
Context window1M tokens
Maximum input991K tokens
Maximum output131K tokens
Maximum input (thinking)983K tokens
Maximum output (thinking)131K tokens
Maximum reasoning262K tokens
TPM2M tokens per minute
RPM15K requests per minute

For current pricing, rate limits, and API capabilities, see Model details and the API Reference.

Qwen3.8-Max is now available for online testing and API access.

3. Model Capabilities

Coding and Research

Multi-Day Autonomous Coding

Qwen3.8-Max can start from an empty folder and connect requirement intake, issue dispatch, code generation, testing, previews, and recovery into one execution process. In one autonomous coding run, the model operated for approximately 16 days and completed 265 commits, 127 pull requests, and 151 issues. Each update was followed by builds, unit tests, end-to-end tests, and desktop lifecycle checks, with failures routed back for correction.

Research Reproduction and Iterative Improvement

In a research reproduction task lasting approximately 125 hours, Qwen3.8-Max started with a paper and GPU access, then built the data processing, training, and evaluation workflow. It wrote roughly 7,600 lines of code, took more than 1,100 actions, and completed 33 GPU training runs. The model used approximately 37 hours to reproduce six main findings, then spent another 88 hours testing 18 improvement ideas. The resulting method improved the AIME24 score by 2.7 points over the paper's approach.

A 24-Hour Multimodal Competition

In the WWW 2025 Multimodal Dialogue Intent Recognition Challenge, Qwen3.8-Max completed rule interpretation, data preparation, model training, fusion, and submission within 24 hours. Across 45 submissions, accuracy rose from 0.60 to 0.853. The final result ranked ahead of 458 of the 526 teams, placing it in the top 13%.

Read more: Full coding and research case studies

Professional Work

Qwen3.8-Max is designed for multi-step professional workflows in legal services, finance, design, and engineering analysis. Training environments scale along three dimensions—tasks, workspaces, and agent harnesses—to cover multi-file, multi-step, and multi-day work. The model can operate across QwenWork, Claude Code, Codex, OpenClaw, and Hermes.

image.png

Figure 2. Aggregate performance across more than 10 work-oriented benchmarks as the number of real RL environments increases.

image.png

Figure 3. Qwen3.8-Max performance across different agent harnesses.

Representative Professional Workflows

  • Corporate compliance review: Read hundreds of documents and identified 1,284 relevant clauses in under one hour.

  • UI/UX design: Generated an eight-page, high-fidelity interactive prototype with a consistent design system.

  • Structural engineering: Reconstructed a 30-story office building's seismic model in a browser from a single drawing, with interactive access to periods, base shear, and inter-story drift ratios.

  • Sports analytics: Converted approximately 8,400 offensive and defensive possessions per player into tactical profiles and coaching reports.

  • Quantitative research: Expanded six short factor descriptions into 50 research directions per category, dispatched approximately 330 sub-agents, and completed around 6,000 backtests. Selected factors achieved excess Sharpe ratios from 0.64 to 1.48.

Read more: Full professional-work and quantitative-research case studies

Long-Horizon Tasks

Chip Design Across Hundreds of Feedback Cycles

Qwen3.8-Max can work through RTL implementation, simulation, synthesis, and physical design using Iverilog, Yosys, and OpenROAD. For a GCD/RSA cryptographic accelerator, the model completed approximately 500 interactions, 71 evaluations, and 13 major stages, reducing its first working design from 8,298 gates to 678. Physical area fell from 106 × 106 μm² to 46 × 46 μm², an 81% reduction. Total wire length decreased from 33,369 μm to 4,187 μm, with timing closure achieved at 500 MHz.

A 365-Day Business Simulation

E-Commerce Bench simulates a full year of online retail across 12 store types, 60 product categories, nearly 600 suppliers, and 7,000 products. Starting with ¥100,000, the model handled product selection, supplier negotiation, inventory, dynamic pricing, returns, seasonal demand, supply-chain disruptions, and 152 fraudulent merchants. After more than 2,000 interactions, Qwen3.8-Max finished with ¥416,252 in cash—a 4.16× return and 38% more than the second-place result.

image.png

Figure 4. Year-end balance trajectory and final ranking in the 365-day e-commerce simulation.

Read more: Full chip-design and business-simulation case studies

Multimodal Agents

Qwen3.8-Max accepts text, images, and video and can use visual information during planning, execution, and verification:

  • Long-document analysis: Understand text, charts, and layouts across financial reports and complex PDFs exceeding 200 pages, then produce structured reports or web deliverables.

  • Long-video understanding: Process videos exceeding 100 hours and organize people, events, timestamps, and scenes into a searchable relationship structure.

  • Visual production: Turn personal footage into a vlog, convert a problem into an instructional animation, reconstruct a frontend project from a screenshot, or build a Blender 3D scene from a floor plan.

  • Visual verification: Read screenshots and rendered output during generation, then revise interfaces, 3D models, or interactive applications based on visual feedback.

image.png

Figure 5. Generating a 3D scene from a floor plan and matching furniture through visual feedback.

Read more: Full multimodal-agent case studies and videos

User Feedback

Qwen3.8-Max has been evaluated in coding and design tools, long-horizon agent products, scientific and engineering research, and document-intensive professional workflows. Feedback has focused on execution speed, dynamic planning, stable tool use, multimodal understanding, and end-to-end task completion. The cards below reproduce a selection of feedback based on hands-on testing and production use.

image.png

Figure 6. Selected user feedback across coding, scientific research, long-horizon agents, and document intelligence.

Read more: Additional user feedback

Complete Evaluation Results

Qwen3.8-Max is evaluated across coding agents, general agents, professional capabilities, multimodal reasoning, visual agents, document and office intelligence, real-world understanding, and video intelligence. Figure 1 provides a summary view; the complete results are listed below. For the broader evaluation context and original presentation, see the Qwen3.8 full release article.

Text, Coding, and Agents

BenchmarkOpus4.8Fable5GPT5.6 Sol (max)Qwen3.7-MaxQwen3.8-Max
Coding Agent
Terminal Bench 2.184.684.688.874.586.6
SWE-bench Pro69.280.064.660.667.7
DeepSWE 1.159.070.073.021.656.6
NL2Repo-Bench69.4----47.255.9
FrontierSWE70.088.8--40.773.5
MLS-Bench-Lite42.849.946.231.741.0
PaperBench80.388.890.564.893.0
AndroidBench69.884.574.056.575.1
QwenSWEBench84.086.373.563.480.7
QwenQoderBench62.763.153.836.858.4
QwenReactBench16941770156415381724
QwenSVGBench16481690175814991713
General Agent
CoWorkBench72.375.971.564.674.8
WorkSpaceBench66.868.765.661.467.7
JobBench48.457.445.431.353.4
SkillsBench65.170.973.561.270.2
Agents' Last Exam (Pass / Score)27.0 / 45.1-- / --30.6 / 53.611.8 / 31.127.0 / 52.4
Automation-Bench (Pass@1)27.229.129.714.227.3
Toolathlon Verified (Pass@1)76.277.974.949.772.5
WideSearch72.981.2--75.281.9
HLE w/ tools57.964.558.053.556.2
General Capabilities
GPQA Diamond92.092.694.192.492.6
HLE45.753.347.241.443.6
IFBench62.263.572.779.182.8
$OneMillion-Bench (expert score)41.855.953.844.452.5
HealthBench52.4--55.354.560.2
PLawBench69.670.272.358.973.2
PRBench-Legal52.757.657.648.557.6
PRBench-Finance51.955.855.546.858.3
MRCR v2 256K (8-needle)83.2--93.886.792.9
LongBench v269.1--67.165.366.3

Evaluation Notes

  1. Fable5 results may involve fallbacks.

  2. Terminal Bench 2.1: Evaluated with Claude Code (avg@10), using a 5-hour timeout and max_tokens=131,072. For all other models, the table reports the best published score across harnesses: Claude Opus 4.8 and Claude Fable 5 with Terminus 2 from Artificial Analysis, and GPT-5.6 Sol with Codex.

  3. SWE-bench Pro: Evaluated with the Claude Code harness, temp=1.0, top_p=0.95, and a 256K context window. Problematic tasks were corrected and all baselines were evaluated on the refined benchmark.

  4. DeepSWE 1.1: Evaluated with the Claude Code and mini-SWE-agent harnesses, temp=1.0, top_p=0.95, and a 256K context window. The table reports the higher score across the two harnesses; Qwen3.8-Max performs best on Claude Code.

  5. NL2Repo-Bench: Evaluated with the Claude Code harness. To prevent reward hacking, Bash commands that attempt to access the specific repository—such as pip download, pip install, and git clone—were disabled.

  6. FrontierSWE: Evaluated with the Claude Code harness. All other available MEAN@5 results are taken from the official FrontierSWE leaderboard as of August 3, 2026. Dominance scores were recomputed from the raw scores using the official evaluation script. “--” indicates that no official MEAN@5 result was available as of that date.

  7. MLS-Bench-Lite: Evaluated with Claude Code using a 5-hour timeout and max_tokens=131,072. All other model scores are taken from the official leaderboard.

  8. PaperBench: Evaluated in the BasicAgent setting under Code-Dev mode, judged by Claude Opus 4.6, and averaged over 3 runs, with a maximum of 12 hours per run.

  9. AndroidBench: Evaluated on the public 95-task subset, reporting avg@3 scores.

  10. QwenSWEBench: An in-house coding benchmark for software engineering capabilities. Evaluated with the Claude Code harness, reporting avg@3 with an 8-hour timeout, max_tokens=32,768, temperature=1.0, and a 256K-token context window.

  11. QwenQoderBench: An in-house coding benchmark for user experience on Qoder. Evaluated with the Claude Code harness, reporting avg@5 with a 6-hour timeout, max_tokens=32,768, temperature=1.0, and a 256K-token context window.

  12. QwenReactBench: An in-house React project-building benchmark using Claude Code as the harness, bilingual in English and Chinese, spanning 7 categories, with auto-rendering, a multimodal judge, and BT/Elo rating.

  13. QwenSVGBench: An in-house SVG code-generation benchmark, bilingual in English and Chinese, with auto-rendering, a multimodal judge, and BT/Elo rating.

  14. CoWorkBench: An in-house cowork benchmark for long-horizon tasks across computer science, finance, law, medicine, and other productivity domains.

  15. SkillsBench: Evaluated on the public SkillsBench v1.1 benchmark across 87 tasks, reporting the average score over 3 runs per task. Opus 4.8 and Fable 5 were evaluated on Claude Code, GPT-5.6 Sol on Codex, and the Qwen series on OpenCode. All results are from this evaluation.

  16. Automation-Bench: Evaluated on the public 600-task subset.

  17. WideSearch: External models were evaluated with the Claude Code harness and Qwen models with the Qwen-Agent harness, reporting average item-F1 over 4 runs.

  18. $OneMillion-Bench: Evaluated using gemini-3.1-pro-preview.

  19. PLawBench: Evaluated using gemini-3.1-pro-preview.

  20. Empty cells (--): Scores are not yet available or are not applicable.

Multimodal, Visual, and Document Intelligence

01| Benchmark | Opus4.8 | Fable5 | Gemini3.1-Pro | GPT5.6-Sol | Qwen3.7-Plus | **Qwen3.8-Max** |
02| --- | --- | --- | --- | --- | --- | --- |
03| **Multimodal Reasoning** | | | | | | |
04| MMMU-Pro | 75.6 | 81.2 | 80.5 | 83.0 | 79.0 | 82.3 |
05| MathVision | 87.1 / 97.1 | 92.7 / 98.6 | 87.4 / 95.7 | 90.8 / 97.8 | 90.3 / -- | 95.2 / 97.7 |
06| BabyVision | 28.4 / 81.2 | 42.5 / 90.5 | 55.9 / 68.3 | 65.5 / 88.9 | 64.7 / 70.4 | 82.0 / 91.3 |
07| HLE-VL (w/ Tools) | \-- | \-- | 43.9 | 51.2 | 25.6 | 52.2 |
08| ZeroBench (Pass@5) | 17.0 / 34.0 | 20.0 / 46.0 | 17.0 / 23.0 | 22.0 / 35.0 | 19.0 / 19.0 | 24.0 / 49.0 |
09| ZeroBench-Sub | 31.1 | 37.1 | 36.5 | 46.7 | 41.0 | 48.5 |
10| LogicVista | 76.7 | 85.7 | 82.6 | 89.7 | 84.3 | 91.9 |
11| HiPhO | 69.3 | 78.6 | 85.4 | 86.8 | 84.1 | 90.0 |
12| PhyX | 54.2 | 71.7 | 79.4 | 79.1 | 80.0 | 83.5 |
13| SLAKE | 75.9 | 86.6 | 82.9 | 85.1 | 83.2 | 90.8 |
14| MedXpertQA-MM | 71.7 | 80.0 | 80.7 | 81.5 | 71.0 | 80.4 |
15| PMC-VQA | 59.2 | 63.2 | 62.5 | 62.3 | 63.4 | 66.2 |
16| **Visual Agent & Coding** | | | | | | |
17| OSWorld-Verified | 83.4 | 85.0 | 76.2 | 83.2 | 73.3 | 86.1 |
18| OSWorld 2.0 | 20.6 / 54.8 | \-- / 66.1 | 7.8 / 30.6 | \-- / 62.6 | 2.8 / 21.5 | 19.4 / 46.7 |
19| ScreenSpot Pro | 82.3 | 87.3 | 68.1 | 81.3 | 79.0 | 84.5 |
20| WebArena-Verified | 67.9 | 71.3 | 64.3 | 69.7 | 55.3 | 66.8 |
21| AndroidWorld | 75.0 | 88.8 | 70.7 | 77.6 | 81.0 | 85.3 |
22| MobileWorld | 67.5 | 85.5 | 58.1 | 76.9 | 51.2 | 77.8 |
23| ClawEval-MM | 73.3 / 73.8 | 81.2 / 77.5 | 50.5 / 55.2 | 81.2 / 78.9 | 57.4 / 60.1 | 77.2 / 74.8 |
24| Vision2Web | 62.4 | 70.5 | \-- | 62.1 | 42.1 | 69.0 |
25| QwenBlenderBench | 62.4 | 69.5 | 23.0 | 68.6 | 41.5 | 69.9 |
26| Parametric CAD Bench | 85.1 | 87.5 | 73.5 | 86.2 | 73.8 | 91.5 |
27| RecreationBench | 48.0 | 56.1 | 16.2 | 47.6 | 30.2 | 51.7 |
28| PresentBench | 80.9 | 79.8 | 55.4 | 82.9 | 65.7 | 79.6 |
29| **Document & Office Intelligence** | | | | | | |
30| CharXiv (RQ) | 78.5 / 89.9 | 87.9 / 93.5 | 84.4 / 89.9 | 85.1 / 89.1 | 85.8 / 85.9 | 88.4 / 93.5 |
31| OmniDocBench 1.5 | 86.5 | 89.5 | 90.0 | 86.7 | 91.4 | 92.1 |
32| OCR-Bench-V2 (EN/ZH) | 53.9 / 55.3 | 65.3 / 58.1 | 64.6 / 58.2 | 69.0 / 57.3 | 70.7 / 67.1 | 74.2 / 68.3 |
33| CC-OCR-Bench-V2 | 60.3 | 72.4 | 68.9 | 68.0 | 72.7 | 79.6 |
34| MTVQA-Test | 48.1 | 41.6 | 54.3 | 52.7 | 51.2 | 56.6 |
35| MADQA | 86.8 | 86.0 | 81.1 | 87.8 | 87.1 | 91.8 |
36| QwenVisualOffice | 34.5 | 32.4 | 39.6 | 29.5 | 32.4 | 44.6 |
37| **Real-World & Spatial Understanding** | | | | | | |
38| RealWorldQA | 76.6 | 85.9 | 83.5 | 83.7 | 86.9 | 88.0 |
39| ERQA | 57.2 | 70.0 | 68.0 | 70.0 | 69.8 | 77.8 |
40| LingoQA | 73.8 | 77.4 | 66.8 | 72.6 | 83.4 | 84.8 |
41| SURDS | 62.2 | 79.4 | 64.0 | 63.0 | 77.2 | 77.8 |
42| **Visual Perception & Grounding** | | | | | | |
43| SimpleVQA | 67.3 | 73.4 | 73.1 | 66.6 | 70.3 | 75.0 |
44| WorldVQA | 33.9 | 53.5 | 54.0 | 45.1 | 43.9 | 53.2 |
45| MMStar | 76.7 | 80.5 | 84.0 | 82.5 | 83.2 | 85.9 |
46| PerceptionBench | 47.2 | 57.2 | 56.2 | 59.7 | 51.1 | 63.5 |
47| CountQA | 41.3 | 63.1 | 72.8 | 68.6 | 77.0 | 82.4 |
48| RefAdv-S | 61.7 | 68.6 | 71.9 | 69.2 | 73.0 | 80.2 |
49| Dense200 | 20.8 | 31.1 | 69.7 | 55.3 | 60.7 | 87.0 |
50| COCO | 50.7 | 56.4 | 72.4 | 61.2 | 74.2 | 78.7 |
51| VisFactor | 30.1 | 54.5 | 39.8 | 62.8 | 42.8 | 60.8 |
52| VLMsAreBiased | 43.8 | 61.2 | 74.1 | 59.8 | 36.6 | 88.3 |
53| **Video Intelligence & Agents** | | | | | | |
54| VideoMME (w/ Sub.) | 85.4 | \-- | 86.7 | 89.5 | 88.0 | 90.4 |
55| VideoMME v2 (w/ Sub.) | 49.0 | 52.2 | 66.9 | 71.1 | 59.7 | 68.3 |
56| VideoMMMU | 75.3 | 81.2 | 85.3 | 85.0 | 85.4 | 88.7 |
57| MMVU | 67.4 | 72.0 | 77.9 | 81.2 | 76.6 | 82.4 |
58| MLVU (M-Avg) | 53.4 | \-- | 84.7 | 87.6 | 87.4 | 90.8 |
59| TVBench | 61.5 | \-- | 73.0 | 83.2 | 78.2 | 81.9 |
60| LVBench | 67.3 | \-- | 75.1 | 78.8 | 76.2 | 81.8 |
61| LVBench (w/ Mem.) | 84.3 | 90.1 | \-- | 84.2 | 74.5 | 85.6 |
62| EgoLife (w/ Mem.) | 78.3 | 82.3 | \-- | 70.8 | 68.8 | 80.3 |
63| VideoDR (w/ Search) | 65.6 | 77.1 | \-- | 71.3 | 41.0 | 73.2 |
64
65Evaluation Notes
  1. MathVision, BabyVision, CharXiv (RQ), and ZeroBench: Scores are reported as “without CI / with CI.” A small number of incorrect ground-truth annotations in MathVision and CharXiv (RQ) were corrected following manual verification.

  2. MathVision: Qwen models were evaluated using a fixed prompt, such as “Please reason step by step, and put your final answer within \boxed{}.” For other models, the table reports the higher score from runs with and without the \boxed{} formatting requirement.

  3. MMMU-Pro: Results for Gemini3.1-Pro and GPT5.6-Sol are taken from model reports or system cards. All other models were evaluated in-house.

  4. ClawEval-MM: Scores are reported as “Pass@3 / average score.” Pass@3 measures the percentage of tasks passed in at least 1 of 3 trials, while average score is the mean across all 3 trials.

  5. Vision2Web: Scores are averaged across the frontend, webpage, and website categories, using the Claude Code harness and gpt-5.4-2026-03-05 as the judge.

  6. HLE-VL (w/ Tools): Evaluated with Code Interpreter (CI) and Search. Scores for the tool-enabled versions of Gemini3.1-Pro and GPT5.6-Sol were measured end-to-end through their native tool-calling APIs.

  7. OSWorld 2.0: Scores are reported as “binary / partial.” The binary score is the percentage of tasks receiving the full task reward, while the partial score aggregates partial rewards across all tasks.

  8. ScreenSpot Pro: Scores for Opus4.8 and Fable5 are taken from system cards. The Fable5 result refers to the corresponding Mythos Preview score. All other models were evaluated in-house.

  9. WebArena-Verified: Scores are reported using the WebArena grader within the OSWorld scaffold.

  10. RecreationBench: An internal long-horizon application-recreation benchmark for hybrid-agent capabilities across Ubuntu, macOS, Windows, Android, and the web.

  11. PerceptionBench: Scores for comparison models are taken from the benchmark's release report, while Qwen3.8-Max was evaluated in-house.

  12. VideoMME (w/ Sub.) and VideoMME v2 (w/ Sub.): Evaluated with subtitles enabled.

  13. QwenBlenderBench and QwenVisualOffice: Both are internal benchmarks.

  14. LVBench and EgoLife (w/ Mem.): Evaluated using a memory system built with Qwen-MM-Plugins, enabling fine-grained, long-horizon video memory.

  15. VideoDR (w/ Search): Evaluated with access to a search tool.

  16. Empty cells (--): Scores are not yet available or are not applicable.

The benchmarks use different tasks, tool environments, settings, and scoring methods, so scores should not be compared directly across rows.

4. Developer Access

Qwen3.8-Max can be called directly through the API or connected to agents, coding tools, IDEs, desktop clients, and workflow platforms.

Call the OpenAI-Compatible API

A first request takes three steps: create a QwenCloud API key, store it in the DASHSCOPE_API_KEY environment variable, and call the OpenAI-compatible endpoint. Do not hard-code or commit the API key in source code.

01import os
02from openai import OpenAI
03
04client = OpenAI(
05 api_key=os.getenv("DASHSCOPE_API_KEY"),
06 base_url="https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
07)
08
09response = client.chat.completions.create(
10 model="qwen3.8-max",
11 messages=[
12 {"role": "user", "content": "Plan and implement a multi-step coding task."}
13 ],
14 extra_body={"enable_thinking": False},
15)
16
17print(response.choices[0].message.content)

This example disables thinking for a basic request. For more complex work, enable thinking and adjust reasoning_effort as needed. See Build with QwenCloud, First API call, and the Qwen3.8-Max API Reference for complete instructions and additional languages.

Connect Agents and Developer Tools

QwenCloud provides setup guides for common agent frameworks, coding assistants, IDEs, desktop clients, and workflow tools:

For example, the OpenClaw integration guide covers installation, credentials, Base URL selection, and the qwen3.8-max model definition. OpenClaw reads its configuration from ~/.openclaw/openclaw.json. After making a partial update that preserves existing settings, restart the gateway and open either the browser console or terminal interface:

01openclaw gateway restart
02openclaw dashboard
03# Or use the terminal interface:
04openclaw tui

Inside the terminal interface, use /model to inspect or switch models and /think <level> to set the reasoning depth. Examples that use auth.mode: none are intended only for single-machine local use; shared or remote deployments should enable token authentication with openclaw doctor --fix.

Reasoning and Built-in Tools

  • Reasoning control: Qwen3.8-Max supports hybrid thinking. Set enable_thinking per request, and use reasoning_effort with xhigh, medium, or low; the default is xhigh. For multi-turn tasks that need earlier reasoning to be included in the next input, set preserve_thinking to true.

  • API compatibility: QwenCloud provides OpenAI-compatible Chat Completions and Responses APIs, with Python, Node.js, cURL, and additional language options documented in the developer guides.

  • Built-in tools: Available tools include Code Interpreter, web search, web content extraction, text-to-image search, and image-to-image search. Availability varies by API; see the API Reference for details.

5. Try Qwen3.8-Max

Qwen3.8-Max: Model details | Try AI | API Reference