Qwen3-Max

Alibaba · Qwen3

Alibaba's older 262K-context Qwen API tier, scheduled to retire in favor of Qwen3.7 Max.

Type
language
Context
262K tokens
Max Output
66K tokens
Status
current
Input
$1.2/1M tok
Output
$6/1M tok
API Access
Yes
License
proprietary
reasoning multilingual tool-use enterprise coding agentic
Released January 2026 · Updated July 10, 2026

Overview

Freshness note: Model capabilities, limits, and pricing can change quickly. This profile is a point-in-time snapshot last verified on July 10, 2026.

Qwen3-Max is an older Alibaba Cloud API model in the Qwen3 family. It was the high-capability production tier for reasoning, coding, and multilingual assistant work before the Qwen3.6 and Qwen3.7 waves.

Qwen3.7 Max is now the hosted flagship and official replacement. Alibaba has scheduled qwen3-max and its January 23 snapshot for retirement on October 10, 2026. The route remains active today, so this entry stays current until that retirement takes effect.

Capabilities

Qwen3-Max is strongest in:

  • Complex reasoning and long-form analytical tasks.
  • Strong Chinese-English multilingual performance.
  • Coding and technical support workflows.
  • Enterprise assistants that need tool integration and structured outputs.
  • Domain-specific adaptation through prompt and retrieval design.

It is commonly used as a premium route, with smaller Qwen variants handling bulk traffic.

Technical Details

Alibaba Cloud model docs currently list for Qwen3-Max:

  • Up to 262,144 token context.
  • Up to 65,536 output tokens on the stable non-thinking route, with preview thinking routes using different limits and pricing.
  • Distinct pricing and token limits by mode and deployment route.
  • API-first availability through DashScope/Alibaba Cloud model services.

Because mode settings can affect both output limits and pricing, production routing should make model mode explicit rather than implicit. This model is also commonly evaluated in bilingual benchmark suites because mode and prompt strategy can affect Chinese and English quality differently.

Pricing & Access

Representative international pricing (per 1M tokens, <=32K input tier):

  • Input: $1.20
  • Output: $6.00

Higher input-length tiers and mode variations can change effective cost.

Access paths:

  • Alibaba Cloud DashScope / Model Studio APIs
  • Related managed deployment surfaces in Alibaba Cloud ecosystems

For predictable cost, teams should enforce context caps and separate lightweight routes from premium routes.

Best Use Cases

Keep Qwen3-Max when you need:

  • High-quality Chinese-English reasoning in one model.
  • Production-grade technical support assistants.
  • Complex agent steps that need stronger planner behavior.
  • Better multilingual output quality than lower-tier options.

For new integrations, use Qwen3.7 Max as the quality tier or a current Plus/Flash model for lower-cost traffic. Existing deployments should validate the replacement before the October 10 retirement rather than changing aliases without regression tests.

Comparisons

  • Qwen3.7 Max (Alibaba): Current 1M-context successor with stronger agent positioning and the official migration path.
  • Qwen 3.6 Max Preview (Alibaba): Newer than Qwen3-Max but also scheduled for retirement on October 10, 2026.
  • GLM-5.2 (Z.ai): MIT-licensed 1M-context alternative for teams that need public weights rather than an Alibaba-hosted route.