NVIDIA Unveils NemotronLabs VoiceChat 11B Open Full-Duplex Speech Model with 450 ms Turn-Taking and Live Tool Calling
NVIDIA
NVIDIA has released NemotronLabs VoiceChat 11B, an open full-duplex speech-to-speech model that performs streaming speech understanding and generation in a single network, achieving 448 ms turn-taking latency. It is the first open full-duplex model to support tool calling while conversing, using a side channel and operator-defined on-hold messages. The model is available for research pilot projects only, requires an 80 GB GPU, and has no hosted API yet.
NVIDIA released NemotronLabs VoiceChat 11B, an open 11B end-to-end speech-to-speech model for real-time, full-duplex conversation. Unlike cascaded systems, it unifies speech understanding and generation in one network, eliminating multi-model orchestration and reducing latency to 448 ms on Full-Duplex-Bench 1.0. The model allows barge-in, with a take-over rate of 1.00 at 480 ms. It also pioneers tool calling in an open full-duplex model, emitting <TOOLCALL> scripts on a separate channel and using operator-defined on-hold lines to avoid silence while APIs run. The model combines a Fast Conformer speech encoder from Nemotron-Speech-Streaming-En-0.6b, the Nemotron Nano v2 LLM backbone, and an NVIDIA TTS decoder, trained on ~550k hours of audio. However, it is for research purposes only, with documented failure modes like a two-minute audio context limit and degradation after several turns. It requires one 80 GB GPU (e.g., A100, H100, RTX 6000 Pro, or B200) and has no hosted API. Constraints include a maximum of five tools per session and ASCII-only system prompts. Performance scores include 58.5% simple tool calling on AU Harness BFCL-v3 and 82.5% tool selection on Full-Duplex-Bench v3. The model ranks #2 among open full-duplex models on VoiceBench and Full-Duplex-Bench 1.0.
- Abbreviations
- ASR = Automatic Speech Recognition — автоматическое распознавание речи
- LLM = Large Language Model — большая языковая модель
- TTS = Text-to-Speech — синтез речи из текста
- GPU = Graphics Processing Unit — графический процессор
- VRAM = Video Random Access Memory — видеопамять
- API = Application Programming Interface — программный интерфейс приложения
Source: MarkTechPost —
original
