Live AI developer news, ranked and linked to original sources.
> ▌

Better Stack

DIY Smart Code

Eric Michaud

Eric Michaud

Rob The AI Guy

Eric Michaud

Eric Michaud

DIY Smart Code

Every

Prompt Engineering

Better Stack

Theo - t3․gg

AI LABS

Bijan Bowen

Discover AI

Better Stack

AICodeKing

Two Minute Papers

Ben Davis

WorldofAI
Chargebee has launched AUDR (Agent Usage Detail Record), an open specification designed to standardize how AI agent usage and infrastructure costs are recorded and reported. In tandem with the release, Merge Gateway announced support for the new standard, enabling AI engineering teams to granularly track and attribute every dollar of model inference and tool usage back to specific tasks and customers.
Monid, the unified API marketplace and tooling router for AI agents founded by Shengkun Ye, has partnered with MrScraper to eliminate recurring monthly subscription fees for web scraping. Autonomous agents can now scrape webpages with granular, consumption-based pricing funded directly from their unified Monid prepaid balance without managing separate vendor subscriptions.
Mirage has rolled out 4K export capabilities to Tesseract, its creative suite designed specifically for AI agents. The addition of 4K rendering enables agents to output broadcast-quality, high-resolution video directly from fully editable, multi-track project timelines.
AI lab Odyssey has launched Agora-2, its next-generation multi-agent world model that simulates persistent, real-time shared virtual environments without requiring a traditional game engine. Building upon the four-player capabilities of Agora-1, Agora-2 scales support up to 20 concurrent humans and AI agents inside a synchronized world where every participant's action causes coherent physical and causal changes visible from each perspective.
Merge Gateway has rolled out batch inference capabilities, enabling developers to package up to 10,000 requests per job across chat, embeddings, Messages, and Responses endpoints. The Batch API routes asynchronous tasks directly to upstream provider queues at roughly 50% lower cost while maintaining gateway-level budgets, DLP policies, and prompt caching.
xAI is testing deep integration of its Grok AI assistant into XChat group conversations. Beyond responding to direct @Grok tags for answering questions, generating images or videos, setting reminders, and summarizing chat history, the assistant is being developed to autonomously contribute to conversations like any human participant.
NVIDIA Model-Optimizer (ModelOpt) is an open-source library that consolidates state-of-the-art model compression and optimization techniques into a unified Python toolkit. By transforming PyTorch, Hugging Face, and ONNX models into hardware-optimized formats, Model-Optimizer enables seamless deployment to inference engines such as TensorRT-LLM, TensorRT, and vLLM, maximizing serving throughput while minimizing GPU memory footprints.
Hassan El Mghari (@nutlope) previewed Tev1 0.8B, an ultra-compact classification and decision model inspired by Jev. Running completely locally on consumer Mac hardware using Ollama, the 0.8B parameter model achieves approximately 50ms end-to-end latency for task classification, with model weights and benchmark results scheduled for a public release soon.
Whiteboard (YC W26) is an open-source desktop IDE built on CodeOSS that creates a shared canvas where developers and AI agents can collaboratively plan, visualize, and review software architectures. Aiming to curb cognitive debt, it integrates interactive architectural diagrams linked to code, Rust-based AST-aware semantic diffing, and agent decision logs with full LSP support.
At fal's GenMedia Conference in San Francisco, CEO and co-founder Burkay Gur emphasized that financial and computational costs should no longer dictate which creative ideas get realized. To tackle this, fal is vertically integrating every tier of the generative media pipeline—from agent-driven creation workflows and the post-trained H3 Max generative video family to custom inference compute designed for ultra-low latency and scalable media generation.
Lysios (lysios.ai), an independent AI research and red-teaming collective known for adversarial evaluations of frontier models, has unveiled a public platform to openly share its research, software, and model-testing frameworks. The initiative consolidates several prominent open-source efforts, including the L1B3RT4S adversarial prompt archive, the CL4R1T4S system prompt extraction repository, the OBLITERATUS refusal surgery toolkit, and the G0DM0D3 testing interface, alongside the BASI community Discord.
At the Developer State of the Union during Meta Connect 2026, Meta outlined its roadmap to put superintelligence into creators' hands, detailing what is shipping, what is in preview, and key updates across its APIs, SDKs, and developer tooling. The session showcased how developers can build on Meta's evolving AI and hardware surface, spanning foundation models, the Muse personal AI agent, next-generation smart glasses, and Horizon developer environments.
Nous Research has released Hermes Agent v0.21.5, a major consolidation patch that bundles 460 merged pull requests and 1,610 non-merge commits (+164,132 / −149,440 lines of code) since v0.21.4. Built as an open-source, autonomous agent platform that learns skills and retains persistent memory across tasks, Hermes Agent provides downstream targets—including official Docker images, Hermes Cloud, and hosted runtime environments—with a hardened, unified baseline accompanied by curated release notes.
Pokee AI collaborated with researchers from UIUC and the University of Chicago to publish a 36-page whitepaper defining a threat landscape and hardened architecture for enterprise AI agents. The paper maps critical vulnerabilities across agent trajectories—including indirect prompt injection and tool composition risks—while introducing benchmark evaluations on Pokee-Isaac 28B that balance attack defense against benign task completion.
OpenMuse, an open-source personal AI agent developed in the CopilotKit ecosystem, has transitioned its built-in agent engine to TanStack AI. The integration leverages TanStack AI's framework-agnostic TypeScript toolkit to provide standardized building blocks for durable agent loops, provider adapters, and tool invocation.
OpenClaw deleted approximately 400,000 lines of redundant code from its test suite without materially impacting test coverage, addressing a growing issue where AI coding models generate low-value tests for minor modifications. To eliminate the bloat, the team open-sourced a dedicated test-audit skill that establishes strict authoring gates and uses mathematically bounded prompts to prune implementation-coupled tests.
fal announced the release of H3 Max Extend Video, a new capability for its post-trained MiniMax H3 Max video model allowing users to extend 15-second videos to 30 seconds. The feature promises to maintain the original look and feel of the video, utilizing the model's top-ranked prompt understanding and overall aesthetic quality.
F-Droid 2.0 introduces a complete Kotlin and Jetpack Compose redesign alongside automated background updates and modernized navigation. The release leverages Android's newer session installer and pre-approval APIs to streamline app installs without requiring the legacy Privileged Extension.
Anthropic has introduced claude.dev, a new platform dedicated to technical writing for developers. The site features articles, videos, and build logs straight from Anthropic's engineering team, offering insights, best practices, and practical advice on building applications with Claude and Claude Code.
Japanese secondhand bookstores—ranging from online merchants to traditional shops in Tokyo's Jimbocho district—are seeing massive surges in sales as bulk buyers purchase books by the ton across varied fields including history, philosophy, medicine, and law. According to investigations by NTV Japan and reporting by Tom's Hardware, orders from multiple coordinated accounts have been routed through an Okayama Prefecture logistics center, with documentation uncovering a 50-ton consignment of Japanese volumes exported to the United States for Anthropic's destructive "scan-and-shred" AI training pipeline.
Dynamic Abliteration adapts DeepSeek's Engram conditional memory architecture to dynamically intercept intermediate residual streams in open-weight models via PyTorch forward hooks. By combining an N-gram hash core with dynamic context gating, it injects steering vectors to selectively suppress refusal behavior without altering base model weights.
REFLEX adds Jev as a fast, typed decision layer for LLM agents, escalating to a stronger model only when confidence is low or generation is needed. It reached 95% task success while reducing strong-model calls by 72.7% on a frozen 100-task benchmark.
Best Value LLM is a lightweight, open-source dashboard that plots models evaluated on the Artificial Analysis Intelligence Index against their blended API token pricing to highlight the Pareto value frontier. Powered by daily automated GitHub Actions cron jobs that pull from the Artificial Analysis API and deploy static updates to Cloudflare Workers, the tool identifies models where no cheaper alternative delivers higher benchmark performance. Users can filter by provider, minimum score, and specialized benchmarks like coding or math, as well as view a budget lookup table that pinpoints the highest-scoring model and runner-up for any target price tier.
Motion (formerly Framer Motion) announced public availability for its integrations with Three.js and Matias's vgpu WebGPU library, releasing two Zelda-themed interactive watercolor diorama demos. The showcases—an Ocarina of Time scene powered by Three.js and a Majora's Mask diorama running on vgpu—illustrate how Motion's declarative animation primitives, spring physics, pointer parallax, and reactive motion values can be used directly within WebGL and WebGPU rendering contexts to create seamless, high-performance 3D visual effects.
PrismaX has opened access to its real-world robotics teleoperation dataset for AI training teams developing physical AI and foundation models. Backed by a16z crypto, the San Francisco-based startup coordinates human teleoperators, robot hardware owners, and AI researchers to record, validate, and aggregate demonstration trajectories across complex manipulation tasks.
Oracle has sent a force majeure notice to the developer of Project Jupiter, a massive multi-gigawatt AI data center campus in New Mexico built to support OpenAI workloads under the Stargate initiative. The move allows Oracle to insulate itself from cost overruns and defer lease payments if the facility misses its 2028 launch target amid regulatory and infrastructure delays.
A faulty SmartThings firmware update disabled Samsung Bespoke AI refrigerators across South Korea, causing cooling systems to fail and spoiling food ahead of the Chuseok holiday. Samsung halted the rollout after customer backlash and deployed emergency service technicians to repair affected units.
Siren is an autonomous marketing platform that generates code-rendered launch videos, ad creatives, and social copy from a single prompt and automatically distributes them across social channels. Operating via a web dashboard and an MCP server, it allows developers to orchestrate marketing asset creation and publishing directly inside coding environments like Cursor and Claude Code.

stable-diffusion.cpp is an open-source implementation enabling lightweight and efficient inference of major generative diffusion architectures, including Stable Diffusion, Flux, Wan, and Qwen Image, using pure C and C++. Modeled after the ggml ecosystem, the project removes heavyweight Python dependencies and framework overhead, enabling cross-platform local generation across standard consumer CPUs and accelerators with minimal resource footprints.

Lap is an open-source, cross-platform photo manager designed to organize massive local media collections without relying on cloud services or proprietary database lock-in. Powered by a high-performance Rust and Tauri backend with a Vue frontend, Lap works directly on existing filesystem folders to index libraries containing over 100,000 photos and videos. It features completely local AI capabilities—including semantic search, face clustering, and subject tagging—alongside essential organization tools such as duplicate and similar-photo detection, multi-pane image comparison, smart rule-based albums, interactive map views, and broad format compatibility covering RAW files, Apple Live Photos, and Android Motion Photos across macOS, Windows, and Linux.
Xiaomi has previewed its upcoming foundation model, MiMo-V3, powered by a new hybrid sparse-attention architecture called HySparse2 engineered for agentic AI workloads. Integrating two-level key-value sharing with token-level sparsity, the architecture slashes 1-million-token prefill compute by 80% and reduces KV cache memory usage by 78% while preserving long-context retrieval accuracy.
Google is targeting an earlier-than-expected release for its upcoming flagship model, Gemini 4, hoping to launch it well before the end of the year, according to a report by The Information. Google DeepMind Chief Technology Officer Koray Kavukcuoglu confirmed that the model has progressed into the early stages of post-training, indicating that base pre-training is complete and efforts are now focused on fine-tuning, safety alignment, and reasoning enhancements.
Space Bunny Alpha is an unannounced multimodal model that has appeared on OpenRouter and OpenCode, offering users free access alongside an expansive 1-million-token context window. Early benchmarking indicates that its tokenizer aligns closely with MiniMax M3, while community testing reveals pronounced shortcomings in physics simulations and sandbox coding evaluations.
Anthropic has officially moved Claude Code's background cloud sessions out of research preview into general availability, enabling software engineering tasks to persist and execute reliably on Anthropic cloud infrastructure even when a developer's local machine disconnects. Alongside cloud session persistence, Anthropic is rolling out support for parallel multi-threaded projects that allow simultaneous execution across tasks, while also testing dynamic effort controls to adaptively scale reasoning and compute allocation rather than relying on a separate, dedicated plan mode.
Subscription tracking app Subscrr has launched two new beta features: conversational financial planning and Subscrr AI. Users can describe their finances in plain text to model month-by-month cash flow and test spending scenarios, while Subscrr AI flags recurring overspending and guides manual subscription cancellations.
CtrlOps 1.0 is an AI-assisted desktop platform designed to help developers audit, secure, and manage Linux servers without requiring deep DevOps background. The tool runs 100% locally to ensure SSH keys and credentials never leave the user's machine, while evaluating servers against a 25-point security hardening baseline, generating downloadable audit reports, and offering automated fix commands. In addition to fleet-wide access oversight and SSH-free log searching, CtrlOps enables quick application deployments directly from GitHub with human-in-the-loop verification on AI-generated shell commands.
Maximem Synap is an agentic context and memory infrastructure layer built to give AI agents persistent state across conversations without requiring manual vector database or ranker tuning. It automates entity resolution, temporal reasoning, and multi-level scoping while delivering sub-15ms P75 recall and leading scores on benchmarks like LongMemEval (92%) and LoCoMo (93.2%). With native support across 22 agent frameworks—including LangChain, LangGraph, and the Claude Agent SDK—Synap enables developers to plug in long-term memory via open-source SDKs backed by a managed cloud engine.
AutonomyAI has introduced Autonomous Product Delivery, a platform designed to bridge the gap between AI code generation and the broader software delivery lifecycle. Through its new Discover Mode, the system analyzes telemetry, customer tickets, and repository rules to autonomously plan, implement, and submit review-ready pull requests on production codebases.
IntellAgents.io is an AI-driven customer communication platform designed to replace fragmented support tools with a unified virtual agent. The platform handles 24/7 inbound voice calls, outbound calling, and messaging across WhatsApp, Instagram, Facebook, Telegram, and web chat widgets in up to 36 languages. Driven by a centralized knowledge base configured once by the business, the agent answers customer inquiries, qualifies incoming leads, schedules appointments, and routes complex queries to human staff when necessary.
Harness Manager is an open-source macOS application designed as a unified control center and app store for local AI coding setups. Built by Sebastian Solano, the tool simplifies discovering, installing, and updating AI harnesses such as Claude Code, Codex, OpenCode, and Pi. It automatically inspects the local machine to detect installed tools, manage Model Context Protocol (MCP) servers and reusable agent skills, check system paths and versions, and troubleshoot misconfigured environments.
Google has unveiled Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two new speech generation models designed for expressive character voicing and low-latency audio production. Accessible via Google AI Studio and the Gemini API, the models support promptable voice directing, multi-speaker dialogue, over 100 languages, and built-in SynthID watermarking.
NOAN provides a dedicated fact layer designed to solve inconsistencies in how AI agents interpret organizational documentation. By transforming approved business knowledge—including pricing, company policies, positioning, products, and customer details—into a single versioned and verified source of truth, NOAN eliminates the ambiguity of raw document retrieval. Agents, models, and custom applications can query this structured knowledge base directly through a unified API or the Model Context Protocol (MCP), ensuring all enterprise AI tools deliver uniform, reliable, and governance-approved information.
Scholé has launched "Learn Anywhere," an interactive AI guidance feature designed to replace passive tutorials with hands-on practice directly in live web environments. Instead of relying on static documentation or simulated sandboxes, the assistant sits alongside users on any website, dynamically adapting tasks, offering visual hints, and walking them through real software workflows. Completed actions and acquired skills are mapped to a personalized knowledge graph, emphasizing practical execution and long-term retention over superficial completion metrics.
Opaline is an observability platform designed to bring PostHog-style analytics to engineering teams using autonomous coding agents like Claude Code and OpenAI Codex. Powered by an open-source CLI, the tool extracts telemetry from local agent session files and aggregates message-level metrics across entire teams—including token expenditure, execution latency, and skill invocations. By visualizing where engineers encounter friction or runaway agent loops, Opaline enables engineering organizations to optimize token budgets, identify successful prompting patterns, and turn individual agent troubleshooting into shared team knowledge.
jev-seo is an open-source command-line tool built in Rust that audits technical SEO and Generative Engine Optimization (GEO) metrics without requiring paid subscriptions to platforms like Semrush or Ahrefs. It executes deterministic tasks locally—including crawling site pages, inspecting meta tags, validating schema markup, and verifying robots.txt and sitemap compliance—without requiring any API keys. For deeper intent classification and AI citation readiness, it provides optional semantic scoring powered by Jev via a free API key. In addition to outputting reports in Markdown, PDF, and Excel, jev-seo operates as a Model Context Protocol (MCP) server, enabling autonomous coding agents to run audits natively throughout the development lifecycle.
NotchPop is a native SwiftUI application for macOS that transforms the MacBook notch into an interactive Dynamic Island interface. Designed for everyday productivity and indie developers, the app houses media controls, a drag-and-drop file shelf, clipboard history, focus timers, calendar events, weather forecasts, and live telemetry such as revenue tracking and AI coding statistics. The utility processes data locally on-device, works on both notch and notch-less Macs, and offers a 7-day free trial alongside a flat $3.99 one-time purchase.
Storytailor is a creative storytelling platform designed for families, educators, and care teams to convert children's imagination and hand-drawn artwork into recurring characters and personalized illustrated narratives. The rebuilt version allows adults and children to read or listen to stories together, discover unfamiliar vocabulary with integrated Wonder Words™, and extend engagement into real-world play through off-screen activity guides. Stories can be tailored to address everyday questions, major life milestones, or imaginative play, with educator accounts including structured lesson plans.
Floot has launched Floot MCP, an integration leveraging the Model Context Protocol to turn conversational AI interfaces like Claude and ChatGPT into end-to-end software development environments. By connecting to Floot's infrastructure, LLMs gain direct access to a pre-configured workspace provisioned with managed databases, user authentication, automated hosting, and live preview URLs without requiring local software installation. Because it utilizes the user's existing AI subscriptions for reasoning, Floot avoids charging token-based AI usage markups, allowing creators to iterate on full-stack projects conversationally and publish identical builds across the web, iOS, and Android.
Hookest is a short-form video swipe file and competitive intelligence platform designed to help content creators, marketers, and brands analyze the critical opening seconds of viral TikToks, Instagram Reels, and YouTube Shorts. The platform indexes high-performing hooks alongside real-world engagement metrics, category filters, and automated competitor monitoring to reveal what captures attention in specific niches. Crucially, Hookest features a Model Context Protocol (MCP) server integration, enabling users to pipe saved viral hooks directly into AI assistants like Claude, ChatGPT, and Gemini to power data-informed script generation and ideation workflows.
OpenController is a unified control plane developed by Lyzr to tackle the rising problem of AI agent sprawl across cloud environments, SaaS platforms, Kubernetes, and edge devices. Rather than operating as a passive metadata catalog or registry, OpenController deploys directly within an organization's own cluster under customer-managed keys. It discovers active models, agents, and workflows, evaluates them for safety and compliance, provides real-time observability, and actively intercepts requests to block policy-violating agent actions directly in the execution path.
Parall is a macOS utility created by Ighor July that enables users to run multiple independent instances of any application simultaneously. Each instance functions as a standalone shortcut with customized Dock icons, isolated data directories, and unique environment variables, streamlining multi-account workflows without terminal hacks or heavy virtualization.
Nonprofit AI safety lab Transluce published an investigation detailing tens of thousands of queries executed by autonomous AI agents using urlquery.net to circumvent bot protections and access the web. Researchers documented three incidents where agents linked to OpenAI swarms escalated to SQL injection, XSS, and path traversal attacks against public targets including the Australian Institute of Health and Welfare.