Qwen3.7 Max

Alibaba · Qwen3

Alibaba's proprietary 1M-context Qwen flagship for reasoning, coding, productivity, and long-horizon agents.

Type
language
Context
1M tokens
Max Output
66K tokens
Status
current
Input
$2.5/1M tok
Output
$7.5/1M tok
API Access
Yes
License
proprietary
chinese multilingual reasoning coding agentic tool-use long-context prompt-caching hosted
Released May 2026 · Updated July 10, 2026

Overview

Freshness note: Model capabilities, regional availability, and pricing can change quickly. This profile is a point-in-time snapshot last verified on July 10, 2026.

Qwen3.7 Max is Alibaba’s current highest-capability hosted Qwen model. The stable qwen3.7-max route launched globally on May 21, 2026 and targets difficult reasoning, coding, office productivity, tool use, and long-horizon autonomous execution.

This is a proprietary Model Studio model rather than an open-weight Qwen checkpoint. Its practical role is the escalation tier above Qwen3.7 Plus and Qwen3.6 Flash when quality matters more than price or latency.

Capabilities

Qwen3.7 Max supports hybrid thinking, with reasoning enabled by default but available in both thinking and non-thinking modes on current stable routes. Alibaba documents function calling, prompt caching, built-in tools, and OpenAI-compatible Responses and Chat Completions APIs.

The 1M-token context is useful for large repositories, long document sets, and persistent agent sessions. Through the Responses API, supported built-in tools include web search, web scraping, and code interpreter. Tool and structured-output support can differ by endpoint and dated snapshot, so production integrations should pin and test the exact route they use.

Technical Details

Official anchors at this snapshot:

  • Stable model ID: qwen3.7-max, currently mapped by Alibaba’s pricing page to qwen3.7-max-2026-05-20.
  • 1,000,000-token context window and 65,536-token maximum output.
  • Up to 256K reasoning budget in current Qwen Cloud documentation.
  • Text input and text output on the stable alias.
  • Thinking and non-thinking modes, function calling, context caching, and built-in tools.

Alibaba also exposes qwen3.7-max-2026-06-08, a separately pinned snapshot with text, image, and video input. The stable alias remains a language-model route because it still points to the text-only May snapshot at this check. Alibaba does not publish weights or full architecture details for the Max tier.

Pricing & Access

Standard international Model Studio pricing is 2.50per1Minputtokensand2.50 per 1M input tokens and 7.50 per 1M output tokens. Regional rates, cache discounts, batch discounts, free quotas, and temporary Token Plan promotions can differ, so procurement estimates should use the deployment region’s live pricing page.

Access options include Alibaba Cloud Model Studio, OpenAI-compatible and Anthropic-compatible APIs, the Responses API, Token Plan Team Edition, and supported third-party gateways. Exact model and snapshot availability varies by region.

Best Use Cases

Choose Qwen3.7 Max for difficult bilingual reasoning, architecture and code review, repository-scale analysis, office automation, and agent workflows that need a large evidence window. It is also a strong escalation route in Qwen-native stacks where Plus or Flash handles routine traffic.

It is less suitable when you need public weights, predictable low-cost bulk inference, or multimodal behavior through the stable alias. Use the June 8 snapshot only after validating its distinct API and modality behavior.

Comparisons

  • Qwen 3.6 Max Preview: Older 262K-context preview scheduled for retirement; Qwen3.7 Max is the current migration target.
  • Kimi K2.7 Code: Open-weight multimodal coding specialist at a lower API price; Qwen3.7 Max is the broader proprietary reasoning and productivity tier.
  • GLM-5.2: MIT-licensed 1M-context open-weight alternative; Qwen3.7 Max offers Alibaba’s managed tool and Model Studio ecosystem.