ByteDance Releases SeedRealtime: Native Audio-Video Full-Duplex Foundation Model, Ushering in the Era of "See, Listen, and Speak" Multimodal Interaction

8.5 On August 5, ByteDance Seed officially launched the native audio-video full-duplex foundation model SeedRealtime, announcing that it has been fully rolled out in the Doubao App, achieving the industry's first large-scale commercialization of audio-video full-duplex technology.

On August 5, ByteDance Seed officially launched the native audio-video full-duplex foundation model SeedRealtime, announcing that it has been fully rolled out in the Doubao App, achieving the industry's first large-scale commercialization of audio-video full-duplex technology.

2026-08-05_135916_719

Technical Breakthrough: From "Q&A" to "Full-Duplex Synchronization"

Audio-video real-time interaction had long been trapped between two paradigms. Cascade systems relied on串联 ASR, visual models, TTS, and other modules, with latency stacking up at each layer and information degrading step by step. End-to-end models were smoother, but many still depended on external VAD for turn-taking, essentially remaining half-duplex interaction.

SeedRealtime's core breakthrough lies in unifying sound, vision, timing, and expression into a single end-to-end model. Instead of listening first, then watching, then responding, it performs perception, understanding, decision-making, and expression simultaneously on continuous audio-video streams, allowing "heard" and "seen" information to participate together in every real-time judgment.

Three Core Capabilities: Understanding, Proactivity, Rhythm

The model achieves three key breakthroughs:

Audio-Video Joint Understanding: Natively supports deep integration of sound, visuals, and timing information. The model can resolve homophone ambiguity through visual context and accurately understand temporal references in vision. When a user says "how do I do this," the model must combine the current visual scene, gestures, gaze, and historical actions to determine what "this" refers to.

Proactive Interaction: Possesses continuous environmental perception and proactive expression. The model can proactively alert when it detects changes in the visual scene (e.g., a key target appears) and can invoke tools to enhance its expression, upgrading interaction from passive response to proactive collaboration.

Smooth Interaction Rhythm: Can perceive the user's conversational state and pace in real time, naturally interjecting, pausing, and responding at appropriate moments; it also has strong anti-interference capabilities, distinguishing between casual chatter and background noise, avoiding false triggers, and maintaining smooth communication.

Real-World Scenario Testing: From Group Gatherings to Museum Tours

In multiple official demonstrations, SeedRealtime showcased robust scenario adaptability:

Group Gathering: During a noisy dinner with four friends, the model matched names to faces based on physical traits, consistently associating each voice with the correct identity throughout the conversation, distinguishing who said what, and even providing a travel plan balancing everyone's needs.

Foreign Tourist Ordering: When a foreign diner faced a Chinese-only menu, the model recognized dishes directly from the visual feed, recommended them in English, and explained cultural context, such as why "fish-flavored pork" contains no fish.

Proactive Museum Guide: When the user said "remind me when you see the tiger-eating-deer bronze screen stand," the model continuously monitored the moving camera feed and proactively alerted when the artifact appeared, then provided detailed explanations based on the visual details.

Coffee Machine Operation Assistance: When it observed the user pouring whole coffee beans directly into the portafilter, the model quickly pointed out the error; after extraction, it proactively offered adjustment suggestions based on visual analysis of the crema and liquid level.

Academic Reading Assistant: When the user said "watch for me and alert me when you get to the training parameters section," the model continuously viewed the rapidly flipping pages, accurately recognized the "3.4 Implementation" section, and proactively called out.

The model also performs reliably in complex environments. At a busy airport, the model ignored a casual mention of a flight, but when the user formally asked, even though the flight information had moved off-screen, it still combined previously seen display information to answer the real-time arrival time. When a mother asked the model to help her daughter learn English, it remained focused on the child's interaction despite background phone noise.

Performance and Commercial Deployment

End-to-end human evaluations showed that, compared to cascade systems, SeedRealtime reduced rhythm issues in audio-video conversations by half—the model's timing for speaking became more natural, with significant reductions in interruptions, delayed responses, and false triggers from background noise. The probability of completing a smooth, uninterrupted conversation also notably increased.

SeedRealtime is now fully live in the Doubao App. Users can update to the latest version, select "Make a Call" from the dialog box, and enter the video call interface to experience it. Doubao had previously introduced video calling, screen sharing, and photo Q&A features in its HarmonyOS version; the launch of SeedRealtime elevates this experience to a new level of full multimodal real-time interaction.

SeedRealtime's significance lies in enabling AI, for the first time, to "see, listen, and speak" like a real person—synchronously perceiving, understanding, and responding in continuously changing real-world scenarios. From accurately identifying individuals in noisy group settings to proactively alerting in museums and maintaining anti-interference capabilities in airports, these capabilities are pushing AI from a "passive answering tool" toward a "proactive collaborative partner." ByteDance's decision to fully deploy this cutting-edge technology directly in the Doubao App, rather than keeping it a lab demo, signals that the AI model competition has moved from "parameter racing" to "real-world experience supremacy." As the model continues to improve in end-to-end latency, complex multi-person scene understanding, and tool invocation, AI is poised to become an intelligent agent capable of "assisting with real-world tasks".

Project homepage address:
https://seed.bytedance.com/seedrealtime