NVIDIA Nemotron-Labs-Diffusion — tri-mode LLM, 6× faster than Qwen3-8B
NVIDIA released Nemotron-Labs-Diffusion, a language model family that unifies three decoding modes in a single architecture: autoregressive, diffusion-based parallel decoding, and self-speculation.
• Available in 3B, 8B, and 14B parameter sizes
• Base, instruct, and vision-language variants included
• Achieves 6× token throughput over Qwen3-8B on comparable tasks