Mistral Shieldstral: Open 3B Multimodal Safety Model
Mistral has released Shieldstral, a 3-billion-parameter multimodal safety classifier with open weights under Apache 2.0 that scores text and images against policies you write in plain language.
Mistral has released Shieldstral, a 3-billion-parameter multimodal safety classifier with open weights under Apache 2.0 that scores text and images against policies you write in plain language.
Simon Willison shipped LLM 0.32, adding provider-side tools, streaming reasoning traces, OpenAI Responses support, and a Git-style logging store to the popular command-line tool.
NVIDIA released open weights for VoiceChat-11B, the first open full-duplex voice model that calls tools mid-conversation. Here is how it works and how to self-host it.
Cloudflare shipped three inference optimizations that make Kimi K2.6 and GLM 5.2 roughly 30% cheaper to run on Workers AI, and its benchmark tables show the answers do not change.
Alibaba announced Qwen3.8-Max on August 3, 2026: a 2.4 trillion-parameter coding flagship with a 1 million-token context and a promised open-weights release.
LoRA Dataset Studio is a free, self-hosted web app that runs the entire LoRA training lifecycle inside a single browser tab, from one reference photo to a trained, ranked model.
audio.cpp 0.5 turns a pure C++ engine into a full local audio studio: expressive DramaBox TTS, Confucius4 cross-lingual voice transfer, music, and AMD ROCm support, all with no Python.
Y Combinator open-sourced QM, an MIT-licensed multiplayer agent harness that gives every teammate an isolated workspace and drives Claude Code, Codex, OpenCode, and Pi from one core.
DeepSeek shipped the official production build of V4 Flash on July 31, 2026: an open-weight, MIT-licensed 284B-parameter MoE model with a 1M token context, native OpenAI Responses API and Codex support, and API output at $0.28 per million tokens.
MiniMax released H3 on July 31, 2026: native 2K video with built-in stereo sound, all-modal input, and a price it says runs at a third of rival models.
Thinking Machines released Inkling-Small on July 30, a 276B open-weight MoE with 12B active params that matches or beats the full Inkling on reasoning and coding at a third the token price.
TurboFieldfare is an open-source Swift and Metal engine that runs Google Gemma 4 26B in about 2GB of RAM on any M-series Mac by streaming experts from SSD.
Audio8 TTS Preview 0.6B is an open-weights, Apache 2.0 text-to-speech model that clones a voice from a few seconds of audio and speaks 11 languages.
ComfyUI tagged v0.29.0 on July 28, 2026, bundling new HeyGen avatar and voice nodes, OpenAI GPT-5.6, Google Gemini 3.5 Flash, Gemma4 12B, ByteDance seed-audio, and a JoyImageEdit editor in one release.
Rescript is a free, open-source alternative to Descript that lets you edit video and audio by editing the transcript text, running entirely in your browser with local Whisper transcription.
Hubo is an open-source plugin that adds a two-agent implement-and-review loop to AI coding tools like Claude Code and Codex, so every change is critiqued before it reaches you.
An open-source project called Codex Slides turns a prompt, a folder, or an entire code repository into a finished presentation deck, running entirely inside OpenAI's Codex.
audio.cpp 0.4 adds Higgs Audio v3 4B, Fish Audio S2 Pro, and Voxtral Realtime ASR to one C++/ggml engine, running flagship text-to-speech locally on CUDA with no Python.
Echo, a new public-alpha endpoint from Tracer, pools open-weight models like GLM-5.2 and Kimi behind one OpenAI-compatible API, promising Claude-class output at roughly a third of the cost.
NvChat is a free, open-source Windows app that gives you a native desktop chat interface for NVIDIA free hosted LLM API, with access to more than 100 models including vision and reasoning models.
SynthCut is a free, GPL-3.0 video editor built to be operated by an AI agent, exposing itself as a Model Context Protocol server so a client like Claude Desktop can run real, frame-accurate edits locally.
Andrew Ng released OpenWorker, an open-source desktop agent that hands you finished work instead of a chat transcript. It runs locally on your Mac using your own model API keys.
NVIDIA shows how to customize open-weight Nemotron 3 Nano on Prime Intellect Lab in minutes for under $5, lifting a task from 21.9% to 90.6% accuracy.
NVIDIA released Qwen-Image-Flash, a four-step distilled version of Alibaba's 20B Qwen-Image that keeps about 96 percent of the quality while cutting denoising steps more than 12 times.