Alibaba Releases Qwen3.8-Omni-Flash Omni-Modal Model: 26% Average Gain Across 30 Benchmarks, Audio-Video Input Costs Plummet 98%

9.18 On September 18, Alibaba's Qwen officially launched its next-generation native omni-modal model Qwen3.8-Omni-Flash, now available on the Qwen AI platform. The model processes text, images, audio, and video simultaneously, supports a 1 million-token context window, and delivers significant omni-modal gains over its predecessor while maintaining comparable text capabilities.

On September 18, Alibaba's Qwen officially launched its next-generation native omni-modal model Qwen3.8-Omni-Flash, now available on the Qwen AI platform. The model processes text, images, audio, and video simultaneously, supports a 1 million-token context window, and delivers significant omni-modal gains over its predecessor while maintaining comparable text capabilities.

2026-09-19_110826_071

Across-the-Board Gains: 26% Average Improvement Across 30 Benchmarks

Across 30 benchmarks, Qwen3.8-Omni-Flash improved over the previous Qwen3.5-Omni-Plus by an average of over 26%. Performance was particularly strong in audio-video agents, coding, and long-horizon tasks:

  • WildClawBench-MM: +36.5 points

  • AgenticVBench: +22.3 points

  • UniClawBench: 69.6 points

On foundational capabilities, LongAudioSpan improved by 8.3 points, OmniVideoBench by 9.6 points, and OmniCap-IF's CSR/ISR by 8.5/14.1 points. In meeting scenarios, AliMeeting's DER/cpWER dropped from 88.11/89.61 to 3.35/17.18, a qualitative leap in speaker diarization and transcription accuracy.

The company says its audio-video capabilities approach Gemini 3.8 Flash, with overall audio capabilities surpassing it.

2026-09-19_110926_976

Pricing "Cut in Half, Then Halved Again": Audio Input Costs Down 98%

Pricing is highly competitive: text, image, audio, and video input costs just ¥0.8 per million tokens. API pricing for hourly audio input dropped by over 98%, and hourly audio-video input by over 93%, dramatically lowering the barrier for long audio-video processing.

Long Audio-Video Understanding: Token Consumption Down 45.7%

To support long audio-video processing, the company expanded Qwen-MM-Plugins and open-sourced Qwen-Live Harness—the former for on-demand perception, tool invocation, and execution in long workflows; the latter for real-time, continuous omni-modal interaction.

In Agentic long audio-video understanding, OmniVideoBench accuracy rose from 63.4 to 67.8, while token consumption dropped from 145,736 to 79,117—a reduction of approximately 45.7%. The model supports on-demand description, agentic proactive evidence gathering, meeting task advancement, and in-depth video research reports.

In meeting scenarios, the model supports up to one hour of audio-video input, performing speaker diarization, transcription, and identity matching, generating meeting minutes and action items, and analyzing project risks. With tool invocation, it can also send emails, organize tasks, or write code.

Full-Chain Content Production: From Music2MV to Full-Length Film Commentary

On the audio-video production side, Music2MV understands song structure, rhythm, and emotion, outputting sentence-level lyrics with timestamps. Short drama translation handles character recognition, colloquial translation, voice cloning, audio remixing, and final quality checks. For full-length film commentary, users provide a film and a one-sentence request; the agent handles omni-modal understanding, key plot extraction, commentary planning, voiceover, music, editing, and final quality control.

In a self-optimization experiment, Qwen3.8-Omni-Flash attempted to improve the Sichuan dialect recognition of Qwen2.5-Omni-3B within 12 hours. It autonomously selected evaluation sets, ran four rounds of experiments building 3,413 training samples, and reduced character error rate from 25.79% to 15.30%—a relative drop of about 40.7%.

Real-Time Version Debuts "Sound Localization"

The accompanying Qwen3.8-Omni-Flash-Realtime perceives and responds simultaneously to audio-video streams, invoking tools with real-time context. The company calls it the first omni-modal model supporting "sound localization" —fusing spatial audio and visual information to determine sound source direction and distance. It also supports real-time spoken language practice, jointly modeling pronunciation and semantics, and can inject identity, style, business knowledge, and interaction rules via Skills.

Qwen3.8-Omni-Flash marks the transition of omni-modal models from "processing multiple modalities simultaneously" to "coordinating multiple modalities to complete tasks." Two details matter more than the 26% average gain: first, token consumption dropped 45.7%, showing the model learned to "look selectively"; second, the real-time version's "sound localization" shows AI beginning to use spatial auditory information like humans. As omni-modal interaction costs drop low enough and capabilities become natural enough, the boundary between AI and the physical world will open further.