Ali released Qwen3.7-Plus: Independent development of 11-hour Wan-line code APP, visual intelligence ranks among the top five in the world

6.2 It retains and enhances Qwen 3.7's complete agent capabilities in the use of text coding tools and productivity workflows. At the same time, it fully upgrades vision-language fusion capabilities, supports image, video, screen, web pages and text input, and can be completed in the GUI graphical user interface CLI command line interface and tool environment.

On June 2, Alibaba's Tongyi Qianwen team officially released the Qwen3.7-Plus multimodal interactive hybrid agent model. This new model, positioned as an "intelligent base that unifies vision and language," deeply integrates visual understanding based on Qwen3.7's powerful text and Agent capabilities, realizing end-to-end closed loop of GUI operations, CLI calls, code generation and self-verification. In the Vision Arena, the authoritative visual model list, Qwen3.7-Plus has helped Ali rank among the top five in the world and first in China.


2026-06-02_104847_763

 

Core positioning: able to see, think, and do

 

The core feature of Qwen3.7-Plus is that it can "see, think, and do"-it can understand the graphical interface, operate applications, generate code and deliver results. It retains and enhances Qwen3.7's complete agent capabilities in text, coding, tool use and productivity workflow, while fully upgrades vision-language fusion capabilities, supports image, video, screen, web and text input, and can be completed in GUI (Graphical User Interface), CLI (Command Line Interface) and tool environments.

 

In the text test,Qwen3.7-Plus is close to Max-level model performance and remains highly competitive in coding agents, general purpose agents, reasoning, instruction compliance and multilingual tasks. Multimodal tests show that the model has significant improvements in visual reasoning, tool calls and task execution links, and has made significant progress in evaluations such as BabyVision, MathVision, ScreenSpot Pro, OSWorld-Verified, and AndroidWorld.


2026-06-02_105555_701

 


Disruptive measurement: 11 hours of independent closed-loop development of real apps

 

In order to verify the actual implementation capabilities, the Ali team asked Qwen3.7-Plus has independently completed the full-link development of an English word learning app. The Hybrid-Agent agent system built based on this model runs continuously and stably for more than 11 hours without human intervention throughout the process. It has generated more than 10,000 lines of code and triggered more than 1,000 Agent calls, completely covering the core aspects of the entire software development life cycle:

 

- Automatic generation of requirements documents

- Automatic code writing

- Automated installation and deployment

- Test case creation

- GUI automated testing

- Multi-scenario parallel testing

- Automatic update of product descriptions

- Automatic version iteration evolution

 

This measured result marks AI agents are a key leap from "assisted programming" to "full-stack independent development."

 


Multimodal programming: From screenshots to runnable code

 

Qwen3.7-Plus shows strong potential in the field of visual programming. In the desktop application scenario, this model can independently interact with macOS native "Stocks" applications, understand the UI layout and functional details, automatically generate SwiftUI source code, connect to the LongBridge Real Market API to obtain real-time data, automatically compile, build, and launch a replica application. Subsequently, 10 functional verification tests were independently performed and all passed, finally fully reproducing the dark theme, column layout and real-time market interactive experience of the native application.

 

Furthermore,Qwen3.7-Plus also supports:

  • Complex image analysis: understanding spatial information such as subway line maps

  • Search-enhanced visual question and answer: Multimodal reasoning combined with online search

  • Image/Video to SVG Vector Code: Vision-Driven Web Design

  • Browser Agent scenario: Automatically complete Alibaba Cloud ECS Cloud Virtual Machine procurement and closed loop of operation and maintenance links

 


Evaluation performance: Vision Arena is among the top five in the world, and domestic programming capabilities reach the top

 

On the global authoritative visual model list In Vision Arena, Alibaba successfully ranked among the top five in the world and first in China with the strong performance of Qwen3.7-Plus.


2026-06-02_105118_861

 

recalling At the release pace of the Qwen 3.7 series, Ali showed an amazing iteration speed. On May 18, Arena AI officially announced the results of Qwen3.7-Max-Preview and Qwen3.7-Plus-Preview: the former ranked 13th in the world overall in the text field, 7th in the mathematics track, and 10th in the programming track, both of which are among the top ten in the world; The latter won 16th place in the visual field, pushing the Ali Laboratory Visual Track overall ranking into the top five.

 

Earlier Qwen3.6-Plus has caused a sensation in the developer community-on April 4, OpenRouter data showed that its daily calls exceeded 1.4 trillion Tokens, ranking first on the new list, with a growth rate of 711%. In the Arena programming sub-list, Ali has become the second-largest AI organization in the world with Qwen3.6-Plus, second only to OpenAI. 


Cross-framework generalization and open access

 

Qwen3.7-Plus has officially provided services to the outside world through the Alibaba Cloud Refining platform and can also be experienced in Qwen Studio. The model supports OpenAI-compatible API and Anthropic protocol calls, and can maintain stable cross-framework generalization performance whether deployed through frameworks such as Claude Code, OpenClaw, or Qwen Code.

 

This open strategy continues the Ali Thousand Questions seriesThe "open source + commercial" two-wheel drive route. Previously, Qwen3.6-Plus has supported millions of Token native contexts, deep adaptation of native Agent frameworks (OpenClaw, Claude Code, OpenCode, etc.), and multi-language adaptation in 119 languages.

 

With With the official launch of Qwen3.7-Plus, the Alibaba Thousand Questions series is shifting from a "model ability competition" to a new stage of "implementing agent productivity." Behind Vision Arena's top five rankings in the world is the landmark moment when domestic multi-modal large models compete on the same stage and occupy a place on the visual agent track for the first time.