SenseTime Releases and Open-Sources SenseNova-Vision: One Model Unifies Four Major Vision Tasks, Surpassing Vision Banana

7.13 SenseTime has officially released and fully open-sourced the SenseNova-Vision unified vision model. Unlike previous approaches that packaged multiple expert models for detection, segmentation, and depth prediction

SenseTime has officially released and fully open-sourced the SenseNova-Vision unified vision model. Unlike previous approaches that packaged multiple expert models for detection, segmentation, and depth prediction, SenseNova-Vision makes vision a native capability of the foundation model, achieving true unification of all classic vision tasks.

61564417532

I. Core Concept: From "Patchwork Unification" to "Native Unification"

Traditional vision AI followed a "one task, one model" path, with different expert models operating in isolation. SenseNova-Vision is the first to formulate all vision tasks as a unified multimodal generation problem solvable by a single foundation model—eliminating the need for task-specific prediction heads and modeling text, pixels, semantic information, and geometric features within one shared representation space.

This "native integration" creates a two-way synergy: decades of high-quality vision data directly enhance the foundation model's visual understanding, while the LLM's reasoning capabilities enable vision tasks to achieve cross-domain mastery.

II. One Model Dominates Four Core Vision Tasks

In multiple权威 evaluations, SenseNova-Vision leads across four core vision domains with a single model, matching or surpassing specialized expert models:

1. Structured Visual Understanding: Outperforms general-purpose models in object detection, referring expression comprehension, OCR, and keypoint localization—particularly excelling in dense small-object detection and long-tail recognition.

2. Dense Geometric Prediction: Achieves depth estimation and surface normal estimation accuracy comparable to specialized geometric models, with high stability across indoor and outdoor scenes.

3. Segmentation: Covers general, reasoning, and interactive segmentation, with standout performance in reasoning segmentation and GCG segmentation.

4. Multi-view 3D Geometry: Delivers high-quality multi-view point cloud reconstruction and camera pose estimation with a single model.

3001967668649

III. Comprehensive Surpass of Vision Banana

Against semantic-oriented models like Youtu-VL, SenseNova-Vision achieves comprehensive leadership. Against generation-oriented models like Vision Banana, it demonstrates clear generational advantages:

  • Core Metric Leadership: Outperforms Vision Banana on the vast majority of authoritative benchmark indicators.

  • Broader Task Coverage: Vision Banana handles only "two" of four core task categories, while SenseNova-Vision covers all four—structured understanding, dense geometry, panoptic segmentation, and multi-view 3D.

On depth estimation, SenseNova-Vision achieves 98.1 δ₁ on NYUv2, surpassing Vision Banana's 94.8; on surface normal estimation, it reaches 14.4 mean error on NYUv2, outperforming 17.8.

IV. Complex Scene Performance: See Through Mirrors, Defy Visual Illusions

SenseNova-Vision demonstrates remarkable generalization:

  • Zero-shot Generalization: On unseen game scenes, simultaneously handles surface normal, instance segmentation, and keypoint detection without retraining.

  • Ultra-dense Segmentation: Precisely segments overlapping objects like fish schools, sheep flocks, or cluttered shelf items.

  • See Through Reflections: In environments with mirrors and glass, accurately estimates object orientation and depth despite reflections.

  • Defy Visual Illusions: Correctly segments occluded objects and outputs accurate surface normals even with perspective tricks and forced perspective.


V. Full Open-Source to Drive Multimodal AI Accessibility

SenseTime is simultaneously open-sourcing the SenseNova-Vision Corpus-50M—a vision instruction dataset containing 50 million high-quality samples. The model is available on GitHub, Hugging Face, and ModelScope.

SenseTime has led China's visual AI market for ten consecutive years and in 2025 achieved global No.1 market share in video analytics. SenseNova-Vision evolves the model from an "execution tool" to a "world understanding model".

SenseNova-Vision's core technology will be fully integrated into the SenseNova U-series foundation models.

SenseNova-Vision's significance lies not in topping individual benchmarks, but in proving that "one model unifying all vision tasks" is achievable. When LLM reasoning capabilities deeply integrate with natively unified vision capabilities, vision AI's transition from "execution tool" to "world understanding model" is becoming reality. For a company that has led China's visual AI market for a decade, this open-source release carries a broader meaning: it's not just running fast itself, but paving a new path for the entire industry.

https://github.com/OpenSenseNova/SenseNova-Vision