First Open-Source Diffusion ASR Model: How It Works
Interfaze's diffusion-gemma-asr-small is billed as the first open-source diffusion-based speech recognition model, refining random tokens into a transcript instead of decoding left to right.
Interfaze's diffusion-gemma-asr-small is billed as the first open-source diffusion-based speech recognition model, refining random tokens into a transcript instead of decoding left to right.
ZeroLabs stitches six open-source models into one free browser studio for voice cloning, design, cleanup, sound effects, and transcription, running on Hugging Face ZeroGPU.
claude-real-video is a local CLI that pulls scene-change frames, dedupes repeats, and transcribes audio so Claude, ChatGPT, or Gemini can actually watch a video, not just read its transcript.
Zhipu has shipped ZCode, an official desktop coding agent built on its open-weight GLM-5.2 model, and it lands with a feature no other coding agent ships today: you can trigger and advance tasks from a chat app.
Cline, the open-source AI coding agent trusted by 8M+ developers, launched ClinePass, a flat $9.99/month subscription bundling curated open-weight coding models directly into its IDE, CLI, and SDK.
ComfyUI-Angelo v2.2.1 adds one-click outpainting and 2x super-resolution to its click-to-refine sampler for FLUX 2 Klein, Qwen-Image-Edit, and SDXL.
Kage, a new open-source verification layer built on Google's Open Knowledge Format, stops AI coding agents from recalling hallucinated or out-of-date memories.
Ploof is a new open-source command-line tool that lets AI coding agents like Claude Code generate images, video, and audio without leaving the terminal.
DeepSeek shipped DSpark and open-sourced DeepSpec, a speculative decoding stack that speeds up V4 inference by up to 80 percent with identical output.
audio.cpp packs more than a dozen audio AI models into one ggml-powered C++ runtime, running TTS, voice cloning, and music up to 5x faster than Python.
OpenKnowledge is an open-source, local-first markdown editor and LLM wiki where Claude, Codex, and Cursor edit your notes directly through MCP.
DeepReinforce has released Ornith 1.0, an open-weights agentic coding model that runs on a single GPU yet posts frontier-level scores on agent benchmarks.
Liquid AI's LFM2.5-230M is a 230M open-weight model small enough to run agentic tool-use directly on a phone or a Raspberry Pi.
Hugging Face now ships huggingface_hub every week using a GitHub Actions pipeline that drafts release notes with an open-weight AI model and keeps a human reviewer in the loop.
Shumai is a new open-source, self-hosted Frame.io alternative for creative teams, with AI semantic video search, frame-by-frame review, and Gemini metadata tagging.
Oak is a new version control system built specifically for AI agents, now in public beta with a Rust core tuned for agent workflows.
Baidu has open-sourced Unlimited-OCR, a compact document-parsing model that transcribes dozens of pages in a single forward pass.
If you are building anything that reads documents, OCR is where most projects quietly break. This guide compares 2026 five best AI OCR tools, Mistral OCR 4, Baidu Unlimited-OCR, DeepSeek-OCR, Google Document AI, and AWS Textract, on accuracy, languages, self-hosting, and cost.
ComfyUI shipped v0.26.0 with native nodes for Kling V3-Turbo, Luma Rays 3.2, Krea2, Qwen3-VL, Boogu-Image, SCAIL-2, HappyHorse 1.1, and a new advanced 3D loader, all in one release.
An open-source 0.2B inpainting model that rivals FLUX.1-Fill at under 2% the size, now running entirely in your browser via a WebGPU port built with Claude Code.
Krea AI has released the open weights for Krea 2, its from-scratch text-to-image foundation model, on Hugging Face.
VNCCS Utils 0.5.3 brings UniCanvas, an infinite-canvas image generation and editing workspace, to ComfyUI with SAM masking and layer-based editing.
LTX Director 2.0 turns Lightricks LTX 2.3 into a full timeline-based video editor inside ComfyUI, adding IC-LoRA, Retake Mode, and audio inpainting.
ComfyUI shipped v0.25.0 and v0.25.1 in three days, adding Kling V3-Turbo, Depth Anything 3, Tripo3D, and native 3D preview nodes.