On June 2, Alibaba's Tongyi Qianwen team officially released the Qwen3.7-Plus multimodal interactive hybrid agent model. This new model, positioned as an "intelligent base that unifies vision and language," deeply integrates visual understanding based on Qwen3.7's powerful text and Agent capabilities, realizing end-to-end closed loop of GUI operations, CLI calls, code generation and self-verification. In the Vision Arena, the authoritative visual model list, Qwen3.7-Plus has helped Ali rank among the top five in the world and first in China.

Core positioning: able to see, think, and do
The core feature of Qwen3.7-Plus is that it can "see, think, and do"-it can understand the graphical interface, operate applications, generate code and deliver results. It retains and enhances Qwen3.7's complete agent capabilities in text, coding, tool use and productivity workflow, while fully upgrades vision-language fusion capabilities, supports image, video, screen, web and text input, and can be completed in GUI (Graphical User Interface), CLI (Command Line Interface) and tool environments.
In the text test,Qwen3.7-Plus is close to Max-level model performance and remains highly competitive in coding agents, general purpose agents, reasoning, instruction compliance and multilingual tasks. Multimodal tests show that the model has significant improvements in visual reasoning, tool calls and task execution links, and has made significant progress in evaluations such as BabyVision, MathVision, ScreenSpot Pro, OSWorld-Verified, and AndroidWorld.

Disruptive measurement: 11 hours of independent closed-loop development of real apps
In order to verify the actual implementation capabilities, the Ali team asked Qwen3.7-Plus has independently completed the full-link development of an English word learning app. The Hybrid-Agent agent system built based on this model runs continuously and stably for more than 11 hours without human intervention throughout the process. It has generated more than 10,000 lines of code and triggered more than 1,000 Agent calls, completely covering the core aspects of the entire software development life cycle:
- Automatic generation of requirements documents
- Automatic code writing
- Automated installation and deployment
- Test case creation
- GUI automated testing
- Multi-scenario parallel testing
- Automatic update of product descriptions
- Automatic version iteration evolution
This measured result marks AI agents are a key leap from "assisted programming" to "full-stack independent development."
Multimodal programming: From screenshots to runnable code
Qwen3.7-Plus shows strong potential in the field of visual programming. In the desktop application scenario, this model can independently interact with macOS native "Stocks" applications, understand the UI layout and functional details, automatically generate SwiftUI source code, connect to the LongBridge Real Market API to obtain real-time data, automatically compile, build, and launch a replica application. Subsequently, 10 functional verification tests were independently performed and all passed, finally fully reproducing the dark theme, column layout and real-time market interactive experience of the native application.
Furthermore,Qwen3.7-Plus also supports:
Complex image analysis: understanding spatial information such as subway line maps
Search-enhanced visual question and answer: Multimodal reasoning combined with online search
Image/Video to SVG Vector Code: Vision-Driven Web Design
Browser Agent scenario: Automatically complete Alibaba Cloud ECS Cloud Virtual Machine procurement and closed loop of operation and maintenance links
Evaluation performance: Vision Arena is among the top five in the world, and domestic programming capabilities reach the top
On the global authoritative visual model list In Vision Arena, Alibaba successfully ranked among the top five in the world and first in China with the strong performance of Qwen3.7-Plus.

recalling At the release pace of the Qwen 3.7 series, Ali showed an amazing iteration speed. On May 18, Arena AI officially announced the results of Qwen3.7-Max-Preview and Qwen3.7-Plus-Preview: the former ranked 13th in the world overall in the text field, 7th in the mathematics track, and 10th in the programming track, both of which are among the top ten in the world; The latter won 16th place in the visual field, pushing the Ali Laboratory Visual Track overall ranking into the top five.
Earlier Qwen3.6-Plus has caused a sensation in the developer community-on April 4, OpenRouter data showed that its daily calls exceeded 1.4 trillion Tokens, ranking first on the new list, with a growth rate of 711%. In the Arena programming sub-list, Ali has become the second-largest AI organization in the world with Qwen3.6-Plus, second only to OpenAI.
Cross-framework generalization and open access
Qwen3.7-Plus has officially provided services to the outside world through the Alibaba Cloud Refining platform and can also be experienced in Qwen Studio. The model supports OpenAI-compatible API and Anthropic protocol calls, and can maintain stable cross-framework generalization performance whether deployed through frameworks such as Claude Code, OpenClaw, or Qwen Code.
This open strategy continues the Ali Thousand Questions seriesThe "open source + commercial" two-wheel drive route. Previously, Qwen3.6-Plus has supported millions of Token native contexts, deep adaptation of native Agent frameworks (OpenClaw, Claude Code, OpenCode, etc.), and multi-language adaptation in 119 languages.
With With the official launch of Qwen3.7-Plus, the Alibaba Thousand Questions series is shifting from a "model ability competition" to a new stage of "implementing agent productivity." Behind Vision Arena's top five rankings in the world is the landmark moment when domestic multi-modal large models compete on the same stage and occupy a place on the visual agent track for the first time.