Qwen3.8-Omni-Flash adds million-token context and lower API costs

Qwen3.8-Omni-Flash adds million-token context and lower API costs

Alibaba has released Qwen3.8-Omni-Flash, a native omnimodal model for text, images, audio, and video with a context window of up to 1 million tokens. The model is designed for large documents, extended conversations, and long-form audiovisual content, with improvements to evidence retrieval and processing efficiency.

For audiovisual tasks, Qwen3.8-Omni-Flash can generate controllable descriptions, gather evidence from long recordings, reason across video sections, analyze meetings, assist with follow-up tasks, and produce video-focused research reports. On OmniVideoBench, its score increased from 63.4 to 67.8 compared with its predecessor, while token consumption fell 45.7%, from 145,736 to 79,117 tokens.

Alibaba says audio input costs have been reduced by 98%, with audio and visual input costs down by more than 93%. The company also reports performance close to Gemini 3.8 Flash across several long-context workloads, although results vary by inference mode. Qwen3.8-Omni-Flash is available through Qwen Chat on web and mobile, while developers can access it through DashScope using OpenAI-compatible client libraries.

by Mauricio B. Holguin

Add as a preferred source on Google
No comments so far, maybe you want to be first?
Gu