On July 2, NVIDIA released and open-sourced Nemotron-Labs-TwoTower, a discrete diffusion language model based on a pre-trained autoregressive backbone network, designed to address LLM token generation speed bottlenecks.
Dual-Tower Architecture: Separating Context and Denoising
The model has 60B total parameters, featuring a dual-tower architecture comprising a 30B autoregressive model and a 30B diffusion/denoising tower, with each tower activating 3B models and 128 routable experts.
TwoTower's key innovation separates text generation tasks into two independent neural network "towers": one tower remains frozen, maintaining the text's autoregressive context; the other handles denoising of noisy blocks, with the two collaborating through layer-by-layer cross-attention connections.
Performance Metrics: Balancing Speed and Quality
NVIDIA reports that in terms of comprehensive benchmark quality, the dual-tower architecture retains 98.7% of quality performance while achieving a 2.42x improvement in runtime throughput.
On specific tasks, MMLU scored 78.24 (compared to 78.56 for the AR baseline), and HumanEval scored 75.58 (compared to 79.27), with minimal quality loss.

The model is released as open weights on the HuggingFace platform under the NVIDIA Nemotron Open Model License.
Editor's Note: In the high-compute-cost environment of 2026, "inference speed" and "token cost" are key determinants of a model's commercial value. NVIDIA's trade-off — 98.7% quality for 2.42x speed improvement — suggests a promising direction for inference optimization. For developers, this means being able to deliver near-top-tier inference results with fewer GPUs and lower costs.