NVIDIA Open-Sources TwoTower AI Model: 2.42x Inference Speed, 98.7% Quality Retention

7.3 On July 2, NVIDIA released and open-sourced Nemotron-Labs-TwoTower, a discrete diffusion language model based on a pre-trained autoregressive backbone network, designed to address LLM token generation speed bottlenecks.

On July 2, NVIDIA released and open-sourced Nemotron-Labs-TwoTower, a discrete diffusion language model based on a pre-trained autoregressive backbone network, designed to address LLM token generation speed bottlenecks.

Dual-Tower Architecture: Separating Context and Denoising

The model has 60B total parameters, featuring a dual-tower architecture comprising a 30B autoregressive model and a 30B diffusion/denoising tower, with each tower activating 3B models and 128 routable experts.

TwoTower's key innovation separates text generation tasks into two independent neural network "towers": one tower remains frozen, maintaining the text's autoregressive context; the other handles denoising of noisy blocks, with the two collaborating through layer-by-layer cross-attention connections.

Performance Metrics: Balancing Speed and Quality

NVIDIA reports that in terms of comprehensive benchmark quality, the dual-tower architecture retains 98.7% of quality performance while achieving a 2.42x improvement in runtime throughput.

On specific tasks, MMLU scored 78.24 (compared to 78.56 for the AR baseline), and HumanEval scored 75.58 (compared to 79.27), with minimal quality loss.

2026-07-03_162655_463

The model is released as open weights on the HuggingFace platform under the NVIDIA Nemotron Open Model License.

Editor's Note: In the high-compute-cost environment of 2026, "inference speed" and "token cost" are key determinants of a model's commercial value. NVIDIA's trade-off — 98.7% quality for 2.42x speed improvement — suggests a promising direction for inference optimization. For developers, this means being able to deliver near-top-tier inference results with fewer GPUs and lower costs.