A proactive full-duplex interaction system from Ant Group's Venus Team and Tsinghua University: two separately trained 9B models that keep listening (and, for one, watching) while they speak, plus a harness that runs slower work in the background. Realtime-Venus-Omni takes aligned video and audio (SigLIP2 and Whisper-Medium encoders) and can speak up unprompted when an event warrants it; Realtime-Venus-Audio drops the visual branch for spoken dialogue. Both are adapted from OpenBMB's MiniCPM-o 4.5, keeping its Omni-Flow architecture and Qwen3-8B language backbone (the Audio checkpoint's config still declares MiniCPM-o's architecture, version 4.5). User inputs, model outputs and delegation events share one causal timeline: each second the model decides whether to listen or speak, and it emits text, S3 speech tokens for a streaming flow-matching decoder, and hidden <delegate> requests. The new piece is asynchronous delegation. In a dual-loop runtime the conversation continues while Realtime-Venus-Harness records each request with a snapshot of the evidence available when it began, routes it to a registered backend capability, and returns a polished reply as private context for the frontend to speak when the moment fits. One post-training recipe covers both models, about 2.8 million samples (roughly 56% offline understanding, 37% proactive duplex interaction, 6% delegation), and updates only the Thinker while the acoustic decoder stays frozen.

Among the online models compared, Realtime-Venus-Omni scores highest on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%) and Daily-Omni (81.3%), 2.3 and 4.0 points above MiniCPM-o 4.5 on the first two, though it trails MiniCPM-o 4.5 on ProactiveVideoQA and WorldSense. Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%) and Speech CMMLU (67.8%) and ties the best VoiceBench AlpacaEval score (4.81). On Full-Duplex-Bench v1.5 it responds to 75% of user interruptions and keeps talking through 97% of backchannels, 88% of other-directed speech and 86% of background speech, higher than Gemini 3.1 Live and GPT-4o on all three. Tool use through the harness is weaker: on Full-Duplex-Bench v3 the Omni frontend reaches 86.0% tool-selection F1 but 43.0% Pass@1, against 87.6% and 60.0% for GPT-Realtime.

The technical report went to arXiv on September 12, 2026 (revised September 18 and 23); the date above marks the September 16 release of both checkpoints on Hugging Face (BF16, 40,960-token context) and of the code on GitHub, all under Apache 2.0. The code covers the Omni model integration, the Harness package and a browser demo that uses a Codex task backend, and an Android demo followed as a beta (119 GitHub stars at filing). The corresponding authors and project leaders are Jian Liu and Yuge Huang of Ant Group and Junliang Xing and Yuntao Wang of Tsinghua University.

Model Details

Architecture DENSE
Context window 40,960
License Apache 2.0
Base model minicpm-o4.5

Benchmark Scores

Benchmark Score Mode
StreamingBench 70.2% Realtime-Venus-Omni
OVO-Bench 64.7% Realtime-Venus-Omni
Daily-Omni 81.3% Realtime-Venus-Omni
MMAU 78.0% Realtime-Venus-Audio
MMAU-Pro 63.2% Realtime-Venus-Audio
Full-Duplex-Bench v3 (tool selection F1) 86.0% Realtime-Venus-Omni

Variants

Name Parameters Notes
Realtime-Venus-Omni — 9B; audio-visual full-duplex frontend; adapted from MiniCPM-o 4.5 (Qwen3-8B language backbone)
Realtime-Venus-Audio — 9B; audio-only full-duplex frontend; adapted from MiniCPM-o 4.5 (Qwen3-8B language backbone)

Paper

Authors: Ruixiang Zhao · Hualei Wang · Renhe Sun · Enzhi Zhou · Zihang Liu
multimodalaudiospeechvideoreal-timeagentsagent-harnessopen-weight

Related