Reve 2.0: 4K Image Generation With Code-Based Layouts
Reve 2.0 generates 4K images using code-based layout controls, giving designers precise composition without prompt engineering.
Reve 2.0 generates 4K images using code-based layout controls, giving designers precise composition without prompt engineering.
Ideogram released its first open-weight text-to-image model on June 3, 2026: a 9.3B parameter Diffusion Transformer with JSON-structured prompting, in-image text rendering, and day-zero ComfyUI support.
Microsoft used its Build 2026 keynote on June 2 to ship a refreshed MAI media stack: MAI-Image-2.5 with editing, MAI-Voice-2 for TTS, and MAI-Transcribe-1.5 for ASR.
Fizgig v1.2.4 makes full Flux 2 Klein 9B LoRA training possible on 16GB GPUs using fp8 Base DiT at 9.6GB VRAM. The free, open-source studio includes training presets, repair tools, and profiler.
NVIDIA pushed new PiD checkpoints June 2 with a FLUX.2 color-fix variant plus Qwen-Image support, all on Apache 2.0 for direct 4K decode in ComfyUI.
Microsoft shipped three creator-focused AI models on June 2: MAI-Image-2.5 beats Gemini on Arena benchmarks, MAI-Transcribe-1.5 runs 5x faster with 43-language support, and MAI-Voice-2 clones voices from short audio samples across 15 languages.
WorkInProcess is a free browser-based image studio with AI upscaling and object removal. No uploads, no accounts, everything runs locally.
Microsoft's MAI-Image-2.5 hit No. 3 on Arena. We compare it against Imagen 4 Ultra, GPT-Image-2, FLUX.2 max, and Recraft V3 on five axes that matter for production creative work.
PrismML released Bonsai Image 4B with 1-bit and ternary checkpoints under Apache 2.0. The model retains 95% of FLUX.2 Klein 4B quality at 6.4x smaller size and runs directly on iPhone.
ComfyUI merged native multi-GPU support on May 26, 2026, giving creators with dual or multi-GPU rigs the ability to split image and video generation workloads across all their hardware for the first time.
Microsoft open-sourced Lens, a 3.8-billion-parameter text-to-image diffusion model, on May 25, 2026. It rivals FLUX and SD3, runs in diffusers and ComfyUI, under MIT license.
MooshieUI is a new ComfyUI frontend that replaces the node graph with a clean three-panel interface. It includes one-click install, 13 model family presets including Anima and Illustrious, and built-in tiled upscaling.
A new ComfyUI custom node released May 23, 2026 brings the Untwisting RoPE technique to Z-Image Turbo, enabling training-free style transfer without any model fine-tuning.
NVIDIA Toronto AI Lab open-sources PiD, a plug-in pixel diffusion decoder that replaces VAE in FLUX, SD3, and Z-Image to output 2K-4K in one distilled pass.
Black Forest Labs shipped FLUX Erase, removing objects and shadows in a single pass without inpainting artifacts.
Midjourney, FLUX Pro, and GPT Image 1.5 lead AI image generation in 2026, but the gap between the top tools has narrowed to the point where price, speed, and workflow fit matter more than raw quality.
OpenAI is adopting C2PA and SynthID watermarking to make AI-generated images verifiable.
ByteDance Research released Lance, a 3B Apache 2.0 unified multimodal model that handles image and video generation, editing, and understanding in a single framework. Strong VBench and GenEval scores.
OpenAI says Indian users created 1B ChatGPT Images 2.0 in under a month. The 5 visual templates driving the volume.
Three significant open-source models arrived in ComfyUI on May 14, 2026: VOID for video object deletion, BiRefNet for image segmentation, and Gemma 4 for multimodal reasoning.
ComfyUI v0.21.1, released May 13, 2026, adds six major partner node integrations including Claude LLM, Grok Image Edit, and OpenAI Image generation inside your workflow graph.
OpenAI shuts down the DALL-E 2 and DALL-E 3 APIs on May 12, 2026. Step-by-step migration to gpt-image-1 with code snippets, prompt recalibration, pricing changes, and a fallback provider list.
Krea launched its first foundation image model on May 12, 2026. Built from scratch for aesthetics and reference-based style transfer, with adjustable multi-source style mixing and a free tier.
OpenBMB shipped MiniCPM-V 4.6, a 1.3B Apache 2.0 vision-language model that runs on iPhone, Android, and HarmonyOS and leads the sub-2B open-weights pack.