Shumai: Open-Source AI Frame.io Alternative
Shumai is a new open-source, self-hosted Frame.io alternative for creative teams, with AI semantic video search, frame-by-frame review, and Gemini metadata tagging.
Shumai is a new open-source, self-hosted Frame.io alternative for creative teams, with AI semantic video search, frame-by-frame review, and Gemini metadata tagging.
Hugging Face now ships huggingface_hub every week using a GitHub Actions pipeline that drafts release notes with an open-weight AI model and keeps a human reviewer in the loop.
ComfyUI shipped v0.26.0 with native nodes for Kling V3-Turbo, Luma Rays 3.2, Krea2, Qwen3-VL, Boogu-Image, SCAIL-2, HappyHorse 1.1, and a new advanced 3D loader, all in one release.
If you are building anything that reads documents, OCR is where most projects quietly break. This guide compares 2026 five best AI OCR tools, Mistral OCR 4, Baidu Unlimited-OCR, DeepSeek-OCR, Google Document AI, and AWS Textract, on accuracy, languages, self-hosting, and cost.
Oak is a new version control system built specifically for AI agents, now in public beta with a Rust core tuned for agent workflows.
Baidu has open-sourced Unlimited-OCR, a compact document-parsing model that transcribes dozens of pages in a single forward pass.
An open-source 0.2B inpainting model that rivals FLUX.1-Fill at under 2% the size, now running entirely in your browser via a WebGPU port built with Claude Code.
Krea AI has released the open weights for Krea 2, its from-scratch text-to-image foundation model, on Hugging Face.
VNCCS Utils 0.5.3 brings UniCanvas, an infinite-canvas image generation and editing workspace, to ComfyUI with SAM masking and layer-based editing.
LTX Director 2.0 turns Lightricks LTX 2.3 into a full timeline-based video editor inside ComfyUI, adding IC-LoRA, Retake Mode, and audio inpainting.
ComfyUI shipped v0.25.0 and v0.25.1 in three days, adding Kling V3-Turbo, Depth Anything 3, Tripo3D, and native 3D preview nodes.
Yes, AI can turn a sentence into a real 3D model you can print, but only some tools give you a model you can also edit. The split that decides everything is parametric versus mesh.
Hugging Face benchmarked LoRA against newer PEFT methods like OFT, which beat it on image generation while using less memory. The default is not always the best fit.
A new open-source tool, UCP-Local, grounds Claude Desktop, Cursor, and LM Studio in your own files for fully offline retrieval. Released June 16, 2026 under Apache-2.0.
PixlStash, the open-source self-hosted image manager for AI creators, now ships as a native desktop app for Windows, macOS, and Linux with one-click install.
SlipMate is a free, open-source generative DJ instrument that runs two AI music models locally and lets you mix their output in real time like vinyl.
Zhipu has shipped GLM 5.2, a coding-first model now live across every tier of its Z.ai Coding Plan, led by a 1-million-token context window.
A new open-source Audio-Reactive LoRA for LTX 2.3 turns a still image and an audio track into a music-driven clip, with motion synced to the beat.
VibeClip is a new open-source, self-hosted tool that turns long videos into vertical captioned shorts you direct by chatting.
Moonshot released Kimi K2.7-Code on June 12, an open-weights coding model that beats Claude Opus 4.8 on MCPMark tool use at far lower cost.
Google DeepMind released DiffusionGemma on June 10, 2026, an Apache 2.0 open model that generates text up to 4x faster by denoising blocks of tokens in parallel, and it runs on a single RTX GPU.
Xiaomi open-sourced MiMo Code, a free MIT-licensed terminal coding agent with persistent memory that runs on its MiMo-V2.5-Pro model and rivals Claude Code.
OpenCV 5.0 now runs LLMs, vision-language models, diffusion, and inpainting natively, turning the most-used vision library into a local runtime for an entire creative pipeline.
NVIDIA released Nemotron 3.5 ASR on June 4: an open 600M streaming speech model covering 40 language-locales with sub-100ms latency for voice agents.