Alibaba Releases Qwen-UI-Agent: Enabling AI to Truly "Use" Every Screen, Surpassing GPT and Claude Across Mobile and Desktop

8.20 Alibaba officially released Qwen-UI-Agent, a real-world-centric foundation GUI agent model covering mobile devices, desktops, web browsers, and DeepSearch environments. The model matches or surpasses flagship models like GPT-5.6 Sol and Claude Opus 4.8 across multiple core benchmarks, transforming AI into a "digital executor" capable of understanding screens and operating software.

On August 20, Alibaba officially released Qwen-UI-Agent, a real-world-centric foundation GUI agent model covering mobile devices, desktops, web browsers, and DeepSearch environments. The model matches or surpasses flagship models like GPT-5.6 Sol and Claude Opus 4.8 across multiple core benchmarks, transforming AI into a "digital executor" capable of understanding screens and operating software.

Real-Device Training: Bridging Simulation and Reality

Unlike most GUI agents trained in simulated environments, Qwen-UI-Agent was built on a real-device training infrastructure covering over 100 real smartphones and 150+ applications for task construction, trajectory collection, and model training and evaluation.

The team established the MobileWorld-Real benchmark (400+ tasks across 100+ apps), fully bridging the gap from simulation to real-world deployment. MobileWorld, a mobile agent benchmark jointly released by Alibaba and multiple institutions at ACL 2026, features tasks averaging 27.8 steps with 62.2% cross-app tasks, significantly harder than previous AndroidWorld.

Performance Leadership: Comprehensive SOTA Achievements

Qwen-UI-Agent achieved new SOTA records across multiple benchmarks:

BenchmarkQwen-UI-AgentComparative Performance
MobileWorld82.1%Outperforms GPT-5.6 Sol by 12.0%, Claude Opus 4.8 by 14.6%, Seed 2.1 Pro by 8.9%
MobileWorld-Real (Real Device)92.2%Exceeds Gemini 3.1 Pro, Claude Opus 4.8, GPT-5.6 Sol, Seed 2.1 Pro
AndroidDaily97.5%Near-perfect score
OSWorld-Verified79.5%Surpasses GPT-5.5, Gemini 3.1 Pro, Seed 2.1 Pro
OSWorld-v2 Partial40.0%Saves 58% execution steps vs baseline
WebArena73.6%Ranked #1 among all compared models
ScreenSpot-Pro81.5%Sets new SOTA on 4 additional grounding benchmarks

2026-08-21_110421_470

Hybrid Actions and Batch Execution: Efficient "Work" Like a Human

Qwen-UI-Agent supports standard GUI clicks, command-line (CLI) operations, and batch action output in a single decision step.

In desktop tasks, CLI actions account for nearly half of all operations, with approximately 40% output in batches, significantly reducing execution trajectories. For example, the model can research products across websites, verify prices, and generate comparison plans.

It also supports online RL training on trajectories exceeding 100 steps, with approximately 10,000 concurrent environments for rollout generation.

Safety Integrated End-to-End: Knowing When to Stop

Qwen-UI-Agent integrates safety judgments throughout task execution:

  • High-risk requests: Directly rejected with no action execution;

  • Sensitive scenarios: Pauses at critical steps (payments, data deletion, privacy authorization) to seek user confirmation before proceeding.

Real-World Scenarios: From Coffee Booking to Business Travel

The model demonstrated cross-app, cross-device execution capabilities:

  • Travel planning: Search addresses, find nearby cafes, summarize reviews.

  • Business travel: Check train schedules, calculate commute times, schedule meetings.

  • File management: Find receipts, move to remote desktop, rename, create Excel summary.

It offers proactive service, seamless cross-device task handover, and integration with DeepSearch.

Qwen-UI-Agent marks the evolution of GUI agents from "simulation toys" to "production tools." Behind the numbers—92.2% success on real devices, 97.5% on AndroidDaily—lies AI's first true capability to "understand screens and operate software." Its significance: countless legacy applications without open APIs can now be automated without modification, as AI operates them just like a human. When AI can use every screen like a human, the boundaries of digital automation are fundamentally redrawn.

The relevant links are as follows: 

Technical Report:https://arxiv.org/abs/2607.28227

Project Homepage:https://tongyi-mai.github.io/Qwen-UI-Agent

GitHub:https://github.com/Tongyi-MAI/MAI-UI