Amap Releases World's First 3D-Native City World Model ABot-Earth 0.7: 10-Minute Kilometer-Scale City Generation, 1,000x Efficiency Gain

9.10 On September 10, Amap (Gaode), a subsidiary of Alibaba Group, officially released ABot-Earth 0.7, the world's first 3D-native city world model. The model supports integrated full-scale AI generation from planetary view to street level, covering over 196 countries and regions—making it the most extensive digital Earth to date. Its official experience portal is now live, with capabilities already deployed in Flight Street View 2.0.

On September 10, Amap (Gaode), a subsidiary of Alibaba Group, officially released ABot-Earth 0.7, the world's first 3D-native city world model. The model supports integrated full-scale AI generation from planetary view to street level, covering over 196 countries and regions—making it the most extensive digital Earth to date. Its official experience portal is now live, with capabilities already deployed in Flight Street View 2.0.

20260910173554

A Technical Path Revolution: From "Capture and Paste" to 3D-Native Generation

For two decades, digital Earth has been built primarily on satellite imagery, aerial photography, and point cloud reconstruction—essentially capturing reality, stitching it together, and presenting it on screens. What users can see and from where remains limited by existing materials and capture paths.

ABot-Earth 0.7 pioneers a 3D-native technical path, training the model on spatiotemporal data including spatial and temporal information to build native understanding of 3D space. It then generates city scenes end-to-end in 3DGS (3D Gaussian Splatting) format.

"Our large models understand language, and Amap's spatial intelligence understands the world," said Guo Ning, CEO of Amap. "The world doesn't just exist in language—it exists in 2D network topologies, 3D spatial structures, real-world 'flows,' and the changes of time."

Efficiency Leap: Kilometer-Scale City in 10 Minutes, 1,000x Improvement

Given a satellite image or a text description, the model can generate a kilometer-scale 3D city scene in just 10 minutes on a consumer-grade GPU—a 1,000x efficiency gain over traditional methods.

More importantly, ABot-Earth 0.7 autonomously completes and generates full, continuous 3D space. Users are no longer constrained by predefined routes; they can continuously enter, freely explore, and receive real-time interactive feedback, giving city exploration and destination browsing a genuine sense of presence.

20260910173624

Full-Scale Consistency: From Planet to Street Level in a Single Model

Unlike 3D models focused on small-scale scenes or game graphics, ABot-Earth 0.7 achieves continuous full-scale generation—from planet to city to street-level landmarks—within a single model.

At the city scale, the model must maintain the overall relationships of road networks, blocks, and building clusters. As the view advances to landmark close-ups, it must further reproduce individual buildings' geometric contours, facade layers, and material textures. Powered by 3D-native representations, spatial structure and visual information remain consistent as the view moves from city overview to landmark detail, achieving photo-realistic visual quality. This "cross-scale consistency" is a capability traditional 3D reconstruction struggles to match.

From Digital Earth to a Spatial Intelligence Gateway

ABot-Earth 0.7 is positioned as the world's first full-modal, predictive, 3D-native city world model. Its core value lies in transforming digital Earth from a static product you can only browse into Amap's gateway for spatial intelligence to understand the real world.

Its capabilities are already deployed in Flight Street View 2.0, where users can enter complex buildings and large scenic areas, and navigate 3D scenes with a joystick. This marks Amap's effort to convert its accumulated spatiotemporal data advantages in navigation and mapping into a core competitive edge in spatial intelligence.

The core breakthrough of ABot-Earth 0.7 lies in the words "3D-native"—it no longer overlays a 3D shell onto 2D images but builds a model that understands the world in 3D from day one of training. This path choice yields dual leaps in efficiency and consistency: a city generated in 10 minutes, seamless cross-scale perspective shifts. While large models compete over "who understands language better," Amap has chosen a differentiated track—making AI truly "understand space." This is not just a technical product but a strategic move to redefine Amap's role in the AI era.