Home/AI News/
Releases
RELEASE
Breaking2 hours ago ยท 4 min read

Alibaba ships Qwen3.7 Flash with 1M context and video input

Alibaba's new vision model reads video across a million token window at $0.03 per million input tokens, though the headline rate only holds under 32K.

AM
Amit Malvi
News Desk, HiggsChat
๐• Share
cover image ยท 1600ร—686
MODEL
Qwen3.7 Flash
CONTEXT
1,000,000 tokens
STATUS
Live on Alibaba API

Alibaba published Qwen3.7 Flash on 27 July 2026, a vision language model that reads text, image and video across a 1,000,000 token context window. Input starts at $0.03 per million tokens and output at $0.13. Alibaba serves it directly, and no other million token model on OpenRouter that accepts video costs less.

ModelContextInput per MOutput per MAccepts
Qwen3.7 Flash1,000,000$0.03$0.13Text, image, video
Qwen3.7 Plus1,000,000$0.32$1.28Text, image
Qwen3.7 Max1,000,000$1.48$4.42Text

What changed

Flash is now the multimodal tier of the Qwen 3.7 line rather than the cut down one. It takes video input, which Qwen3.7 Plus does not, and it does so at roughly a tenth of Plus pricing. Max remains text only and costs about fifty times more per million input tokens.

Against its own predecessor the drop is steeper. Qwen3.6 Flash lists at $0.19 per million input tokens on the same model listing, so the new version cuts entry input pricing by about 84 percent while keeping the million token window and the video input.

Output ceiling is 65,536 tokens per request. Alibaba describes the model as suited to "multimodal agents, visual coding, search, and computer interaction". Object recognition and spatial understanding are the two capabilities the model page puts forward.

What it costs

The $0.03 headline is a short prompt rate, and that detail is missing from most of the listings carrying the number.

Pricing runs on three tiers keyed to how much you actually send. Under 32,000 prompt tokens you pay $0.03 per million in and $0.13 out. Past 32,000 it moves to $0.10 and $0.40. Past 256,000 it moves again to $0.20 and $0.80.

Prompt sizeInput per MOutput per M
Under 32K tokens$0.03$0.13
32K to 256K tokens$0.10$0.40
Above 256K tokens$0.20$0.80

So a request that actually uses the million token window costs $0.20 per million input tokens, which is 6.7 times the advertised rate. That is still low. It is not the number on the label.

Cached input reads are $0.006 per million and cache writes $0.038 per million at the base tier, which matters more than usual on a model whose selling point is long context.

Who this matters for

Anyone processing video or long screen recordings at volume gets the clearest benefit. Of the 368 models listed on OpenRouter, 107 carry a context window of at least a million tokens, and only 35 of those accept video input at all. Qwen3.7 Flash is the cheapest of those 35 on input price, ahead of the nearest batch tier alternative at $0.05 and the nearest standard tier alternative at $0.10.

Teams building visual agents that click through interfaces are the second group, since object recognition and spatial understanding are what the model is tuned for.

Anyone writing plain text should ignore it. A vision model priced for video is not a better writer than a text model, and Qwen3.7 Max exists for that job.

Availability

The model runs on Alibaba Cloud Model Studio and through OpenRouter, which routes to Alibaba as the sole provider. Alibaba's own Chinese documentation lists it in the cn-beijing and ap-southeast-1 regions, reachable through OpenAI compatible, DashScope and Anthropic compatible endpoints.

It is not in HiggsChat today. HiggsChat, an all-in-one AI chatbot that puts every leading text, image and video model under one subscription, currently carries 14 chat models including GPT, Claude, Gemini, Grok, DeepSeek and Sonar, with no Qwen tier among them. For video work the six video models on the platform generate footage rather than read it, which is the opposite job to the one Qwen3.7 Flash is built for.

If your work is reading video at scale, go to Alibaba or OpenRouter directly today. If it is producing video and text on one balance, HiggsChat plans start at $15 a month.

What is still unclear

Alibaba has published no benchmark scores alongside the listing. There is no reported result on any public evaluation, so the claimed strengths in object recognition and spatial understanding are unverified by anyone outside the company.

The English documentation has not caught up. A search of Alibaba's English pricing page returned no occurrences of the model ID, while the Chinese equivalent carries 64. Pricing outside China is therefore visible through resellers rather than through Alibaba's own English pages.

Rate limits are unpublished, so whether the full million token window is reachable at standard account tiers is not known.

There are no open weights. No official repository exists on Hugging Face, which matches the proprietary release pattern of Qwen3.7 Plus.

Sources

  • OpenRouter, Qwen3.7 Flash model page and models API, accessed 29 July 2026
  • Alibaba Cloud Model Studio documentation, Chinese, accessed 29 July 2026
  • Alibaba Cloud Model Studio pricing documentation, English, accessed 29 July 2026
  • HiggsChat model lineup and pricing, accessed 29 July 2026
DR

Amit Malvi

Covers model launches and the AI industry for the HiggsChat News Desk. Cuts through the hype to what actually ships.

More from the wire

qwen3-7-flash-1m-context-video
July 29, 2026
4
Qwen3.7 Flash ships with 1M context and video
Alibaba released Qwen3.7 Flash with a 1,000,000 token context window, video input and $0.03 per million input tokens at the entry tier.
https://higgschat.com/chat-models
https://www.linkedin.com/in/amit-malvi-am7/
https://x.com/i_am_amitmalvi
This is some text inside of a div block.
Dithered abstract signal pattern in the HiggsChat violet to teal gradient, with the words Qwen3.7 Flash centredAmit Malvi