NVIDIA Releases Qwen3.6 NVFP4: 35B Multimodal MoE in 4-bit
NVIDIA ships an NVFP4 4-bit quantized build of Qwen3.6-35B-A3B, cutting GPU memory 3x with under 1% accuracy loss on eight benchmarks.
NVIDIA ships an NVFP4 4-bit quantized build of Qwen3.6-35B-A3B, cutting GPU memory 3x with under 1% accuracy loss on eight benchmarks.
Ship a local multilingual NPC on RTX hardware in 30 minutes with NVIGI SDK 1.6 plus DLSS 4.5 for Unreal Engine 5. Step-by-step, troubleshooting included.
Nvidia released LocateAnything-3B on May 27, 2026, a vision-language model 10x faster than Qwen3-VL at object detection with SOTA accuracy on GUI grounding and document understanding.
CUDA 13.3 lands Python 1.0 stable, CompileIQ for 15% LLM inference speedups, and PyTorch/JAX zero-copy. What creative AI builders gain.
NVIDIA officially retired its GeForce Control Panel on May 26, 2026 after 20 years, completing the migration of all GPU settings to the NVIDIA app.
NVIDIA released the Nemotron-Labs-Diffusion family on Hugging Face, an open-weights LLM that switches between autoregressive, diffusion, and self-speculation decoding for 2.7x to 3.3x throughput gains.
NVIDIA delivered the first Vera CPUs, its first chip purpose-built for agentic AI, to Anthropic, OpenAI, SpaceXAI, and Oracle on May 18, 2026. The chip targets the orchestration layer Claude, ChatGPT, and Grok run on.
NVIDIA and HuggingFace published a full fine-tuning guide for Cosmos Predict 2.5 today, showing how to adapt the 2B-parameter video world model to any domain using LoRA or DoRA on a single GPU.
NVIDIA's SANA-WM is a 2.6 billion-parameter open-source world model that generates 720p, 60-second video with 6-degree-of-freedom camera control on a single GPU.
NVIDIA released Nemotron 3 Nano Omni on April 28: a 30B open-weight model that handles text, image, video, and audio in one architecture, with joint audio-visual reasoning and multi-hour context.
NVIDIA unveiled Cosmos 3 at GTC, the first world foundation model that unifies synthetic world generation, physical AI reasoning, and action simulation.
NVIDIA CloudXR 6.0 now streams RTX-powered 4K graphics directly to Apple Vision Pro, enabling professional 3D design reviews and immersive workflows from any RTX workstation.
NVIDIA unveiled DLSS 5 at GTC 2026, introducing generative AI-powered neural rendering that transforms game visuals with photoreal lighting and materials.
NVIDIA announced DGX Spark at GTC 2026, a desktop workstation powered by the Grace Blackwell Superchip that runs AI models up to 120B parameters locally.
NVIDIA DLSS 5 replaces traditional upscaling with generative AI that creates photoreal lighting and materials in real time. Here is why the gaming community is divided and what it means for creators.
NVIDIA released Dynamo 1.0 on March 16, graduating its distributed AI inference framework to production-ready status with up to 7x higher throughput on Blackwell GPUs.
For years, the most powerful AI creative tools lived in the cloud. You uploaded your prompts, waited for a remote GPU cluster to process them, and downloaded the results
NVIDIA releases Cosmos 2.5 world foundation models for synthetic data generation and physical AI reasoning, now available on Hugging Face.
Nvidia invested $2 billion in Nebius, an Amsterdam-based AI cloud company, in a strategic partnership announced March 11, 2026. The deal commits to deploying more than 5 gigawatts of Nvidia systems through Nebius by 2030, adding a major new full-stack AI cloud alternative to the existing hypersca...
NVIDIA released Nemotron 3 Super on March 11, 2026, an open-source 120B-parameter hybrid Mamba-Transformer model that delivers 5x higher throughput than its predecessor for agentic AI workloads
NVIDIA is launching NemoClaw, an open-source enterprise AI agent platform, ahead of GTC 2026 (March 16-19). The platform integrates with the NeMo framework and Nemotron models, backed by partnerships with Salesforce, Google, Adobe, Cisco, and CrowdStrike
NVIDIA announced DLSS 4.5 at GDC 2026, introducing Dynamic Multi Frame Generation that automatically adjusts generated frames in real time to hit a target frame rate