Nemotron 3 Ultra: NVIDIA's 550B Open-Weights MoE
NVIDIA released Nemotron 3 Ultra on June 1 2026: a 550B mixture-of-experts model with 55B active parameters, open weights on Hugging Face, with 5x faster inference and 30% lower cost than Nemotron 2.
NVIDIA released Nemotron 3 Ultra on June 1 2026: a 550B mixture-of-experts model with 55B active parameters, open weights on Hugging Face, with 5x faster inference and 30% lower cost than Nemotron 2.
NVIDIA announced DGX Station for Windows at Computex 2026, putting a GB300 Grace Blackwell Ultra chip and 748GB of memory into a desktop that runs trillion-parameter models locally.
NVIDIA officially launched the RTX Spark Superchip at GTC Taipei on June 1, 2026, putting a 6144-CUDA-core GPU, a 20-core Grace Arm CPU, and 128GB of unified memory into a single thin-and-light PC chip.
Step-by-step prep for Adobe Photoshop, Premiere Pro, Substance 3D, and Blender on the NVIDIA RTX Spark Blackwell superchip. Audit, benchmark, and plan procurement before the 2x speedup ships later in 2026.
NVIDIA Cosmos 3 ranks first among open-source models on the Artificial Analysis Text-to-Image leaderboard. The 64B-parameter model ships with image-to-video capability and open commercial-use weights.
Developer Oscar Molnar installed a secondhand Tesla V100 SXM2 into his gaming PC alongside an RTX 4080, building a 32GB dual-GPU setup for under £200 total.
Nvidia is bringing Cosmos, Nemotron, GR00T, and Ising under OpenMDW-1.1. Here is what the unified AI model license means for creative AI developers.
NVIDIA ships an NVFP4 4-bit quantized build of Qwen3.6-35B-A3B, cutting GPU memory 3x with under 1% accuracy loss on eight benchmarks.
Ship a local multilingual NPC on RTX hardware in 30 minutes with NVIGI SDK 1.6 plus DLSS 4.5 for Unreal Engine 5. Step-by-step, troubleshooting included.
Nvidia released LocateAnything-3B on May 27, 2026, a vision-language model 10x faster than Qwen3-VL at object detection with SOTA accuracy on GUI grounding and document understanding.
CUDA 13.3 lands Python 1.0 stable, CompileIQ for 15% LLM inference speedups, and PyTorch/JAX zero-copy. What creative AI builders gain.
NVIDIA officially retired its GeForce Control Panel on May 26, 2026 after 20 years, completing the migration of all GPU settings to the NVIDIA app.
NVIDIA released the Nemotron-Labs-Diffusion family on Hugging Face, an open-weights LLM that switches between autoregressive, diffusion, and self-speculation decoding for 2.7x to 3.3x throughput gains.
NVIDIA delivered the first Vera CPUs, its first chip purpose-built for agentic AI, to Anthropic, OpenAI, SpaceXAI, and Oracle on May 18, 2026. The chip targets the orchestration layer Claude, ChatGPT, and Grok run on.
NVIDIA and HuggingFace published a full fine-tuning guide for Cosmos Predict 2.5 today, showing how to adapt the 2B-parameter video world model to any domain using LoRA or DoRA on a single GPU.
NVIDIA's SANA-WM is a 2.6 billion-parameter open-source world model that generates 720p, 60-second video with 6-degree-of-freedom camera control on a single GPU.
NVIDIA released Nemotron 3 Nano Omni on April 28: a 30B open-weight model that handles text, image, video, and audio in one architecture, with joint audio-visual reasoning and multi-hour context.
NVIDIA unveiled Cosmos 3 at GTC, the first world foundation model that unifies synthetic world generation, physical AI reasoning, and action simulation.
NVIDIA CloudXR 6.0 now streams RTX-powered 4K graphics directly to Apple Vision Pro, enabling professional 3D design reviews and immersive workflows from any RTX workstation.
NVIDIA unveiled DLSS 5 at GTC 2026, introducing generative AI-powered neural rendering that transforms game visuals with photoreal lighting and materials.
NVIDIA announced DGX Spark at GTC 2026, a desktop workstation powered by the Grace Blackwell Superchip that runs AI models up to 120B parameters locally.
NVIDIA DLSS 5 replaces traditional upscaling with generative AI that creates photoreal lighting and materials in real time. Here is why the gaming community is divided and what it means for creators.
NVIDIA released Dynamo 1.0 on March 16, graduating its distributed AI inference framework to production-ready status with up to 7x higher throughput on Blackwell GPUs.
For years, the most powerful AI creative tools lived in the cloud. You uploaded your prompts, waited for a remote GPU cluster to process them, and downloaded the results