On August 20, Alibaba officially released Qwen-UI-Agent, a real-world-centric foundation GUI agent model covering mobile devices, desktops, web browsers, and DeepSearch environments. The model matches or surpasses flagship models like GPT-5.6 Sol and Claude Opus 4.8 across multiple core benchmarks, transforming AI into a "digital executor" capable of understanding screens and operating software.
Real-Device Training: Bridging Simulation and Reality
Unlike most GUI agents trained in simulated environments, Qwen-UI-Agent was built on a real-device training infrastructure covering over 100 real smartphones and 150+ applications for task construction, trajectory collection, and model training and evaluation.
The team established the MobileWorld-Real benchmark (400+ tasks across 100+ apps), fully bridging the gap from simulation to real-world deployment. MobileWorld, a mobile agent benchmark jointly released by Alibaba and multiple institutions at ACL 2026, features tasks averaging 27.8 steps with 62.2% cross-app tasks, significantly harder than previous AndroidWorld.
Performance Leadership: Comprehensive SOTA Achievements
Qwen-UI-Agent achieved new SOTA records across multiple benchmarks:
| Benchmark | Qwen-UI-Agent | Comparative Performance |
|---|---|---|
| MobileWorld | 82.1% | Outperforms GPT-5.6 Sol by 12.0%, Claude Opus 4.8 by 14.6%, Seed 2.1 Pro by 8.9% |
| MobileWorld-Real (Real Device) | 92.2% | Exceeds Gemini 3.1 Pro, Claude Opus 4.8, GPT-5.6 Sol, Seed 2.1 Pro |
| AndroidDaily | 97.5% | Near-perfect score |
| OSWorld-Verified | 79.5% | Surpasses GPT-5.5, Gemini 3.1 Pro, Seed 2.1 Pro |
| OSWorld-v2 Partial | 40.0% | Saves 58% execution steps vs baseline |
| WebArena | 73.6% | Ranked #1 among all compared models |
| ScreenSpot-Pro | 81.5% | Sets new SOTA on 4 additional grounding benchmarks |

Hybrid Actions and Batch Execution: Efficient "Work" Like a Human
Qwen-UI-Agent supports standard GUI clicks, command-line (CLI) operations, and batch action output in a single decision step.
In desktop tasks, CLI actions account for nearly half of all operations, with approximately 40% output in batches, significantly reducing execution trajectories. For example, the model can research products across websites, verify prices, and generate comparison plans.
It also supports online RL training on trajectories exceeding 100 steps, with approximately 10,000 concurrent environments for rollout generation.
Safety Integrated End-to-End: Knowing When to Stop
Qwen-UI-Agent integrates safety judgments throughout task execution:
High-risk requests: Directly rejected with no action execution;
Sensitive scenarios: Pauses at critical steps (payments, data deletion, privacy authorization) to seek user confirmation before proceeding.
Real-World Scenarios: From Coffee Booking to Business Travel
The model demonstrated cross-app, cross-device execution capabilities:
Travel planning: Search addresses, find nearby cafes, summarize reviews.
Business travel: Check train schedules, calculate commute times, schedule meetings.
File management: Find receipts, move to remote desktop, rename, create Excel summary.
It offers proactive service, seamless cross-device task handover, and integration with DeepSearch.
Qwen-UI-Agent marks the evolution of GUI agents from "simulation toys" to "production tools." Behind the numbers—92.2% success on real devices, 97.5% on AndroidDaily—lies AI's first true capability to "understand screens and operate software." Its significance: countless legacy applications without open APIs can now be automated without modification, as AI operates them just like a human. When AI can use every screen like a human, the boundaries of digital automation are fundamentally redrawn.
The relevant links are as follows:
Technical Report:https://arxiv.org/abs/2607.28227
Project Homepage:https://tongyi-mai.github.io/Qwen-UI-Agent
GitHub:https://github.com/Tongyi-MAI/MAI-UI