Live AI developer news, ranked and linked to original sources.
> ▌

Better Stack

Prompt Engineering

Better Stack

Every

Better Stack

Income stream surfers

Code to the Moon

Discover AI

Theo - t3․gg

Better Stack

DIY Smart Code

AICodeKing

WorldofAI
llama.cpp merged experimental `-sm tensor` support, letting RPC-connected machines split model weights and KV work across devices instead of assigning each host separate layers. Tests on two DGX Sparks over RDMA reached 19.75 tok/s generation and 619.36 prompt tok/s on DeepSeek4 MXFP4. [PR #26610](https://github.com/ggml-org/llama.cpp/pull/26610)
OpenAI’s Decisions API is now live, turning text or image context into bounded answers for classification, request routing, and choosing an agent’s next action. Its API reference supports predicate, choice, and score questions for direct use in application logic.
Fast Company named Merge one of its seven “Next Big Things in Foundational AI” for 2026, highlighting Agent Handler’s secure, MCP-ready connections to hundreds of enterprise tools. The recognition underscores Merge’s role as infrastructure for production AI agents.
OpenAI’s Meetings plugin captures microphone and system audio without a bot joining the call, then saves personalized notes and action items in ChatGPT Space. It is currently a macOS beta for Pro and Business users, with private-by-default notes and post-processing audio deletion.
SQLDoom reimplements 1993 Doom’s game logic and renderer as SQL inside CedarDB, while Python handles only input, timing, and display. It preserves the original 35 Hz game loop, renders 320×200 frames at up to 60 Hz, and supports multiplayer deathmatch. [CedarDB](https://cedardb.com/blog/sqldoom/) [GitHub](https://github.com/cedardb/sqldoom)
fal’s ChatGPT and Codex integration lets users generate, browse, and reuse images, videos, audio, and 3D media without leaving the conversation. Users can connect their fal account, access its model catalog and Media Library, and keep generated assets close at hand.
Monid says it raised $7.7M to build an OpenRouter-style platform for agent tools, letting agents discover, compare, run, and pay for more than 1,700 APIs through one connection. Its pay-per-call model targets the subscription sprawl blocking autonomous workflows.
Quillan-Ronin is an indie-built AI runtime combining a 34-persona council architecture, custom quillan.cpp inference engine, MCP developer tooling, and an in-training model. Its creator positions the project as a local-first, open-source alternative for experimenting with agent orchestration and efficient inference.
OpenAI is consolidating five paid API usage tiers into Build, Launch, and Grow, while cutting the cumulative payment threshold for Grow from $1,000 to $500. The change makes higher rate limits more accessible to smaller teams.
Anthropic is broadening its Cyber Verification Program, giving verified security professionals access to Claude Mythos 5.1, Opus 5.5, and Sonnet 5.5 with safeguards tailored to defensive work. The program now offers tiered access for higher-risk cybersecurity tasks.
Koru now emits C# for .NET, bringing its compile-time flow model to the CLR ecosystem. In the showcased benchmark, generated C# ran 1.69× faster than lazy LINQ and nearly matched handwritten loops, while native Zig remained significantly faster for allocation-heavy workloads.
Orca’s experimental orchestration layer lets developers pair coordinator and worker agents into supervised, end-to-end workflows, including parallel review councils. It adds structured runs, tasks, dispatches, messages, and decision gates to Orca’s multi-agent coding workspace.
AgentID, from AgentMail, gives AI agents verified email-based identities so they can create and access app accounts independently. It uses standard OpenID Connect, supports major auth platforms, and lets apps identify the human owner behind each agent.

PhotoCraft is an open-source, clean-room reimplementation of Photoshop in pure Rust, with native performance, layered PSD/PSB support, and cross-platform builds. Its command-driven architecture also exposes CLI, JSON, and MCP interfaces for automation. [GitHub](https://github.com/storytold/photocraft)
OpenBot is a free, local-first desktop workspace for persistent AI teammates, connecting Codex, Claude, Gemini, Grok, OpenCode, and custom models. Agents retain workspaces, delegate tasks, share files, and operate a built-in browser while data stays on the host computer.
T3 Code now lets developers trigger agent tasks from webhooks, with templated request data, delivery logs, optional signatures, and T3 Connect relay URLs. It turns GitHub events, CI failures, and other HTTP signals into actionable coding workflows.
Every Agent is a Slack-based AI coworker that turns recurring work into reusable workflows, connects company tools, and handles delegated tasks. Now in beta, it also offers proactive Frontier Alerts and open-loop reminders without marking up model-token costs.
Mirage’s Tesseract lets Claude, ChatGPT, and other supported agents import and export editable Premiere Pro (.prproj) and After Effects (.aep) projects through open-source converters. Premiere support is partial, while After Effects export remains experimental.
Atlassian expanded its OpenAI partnership, bringing GPT-6 Astra and GPT-5.6 models to Rovo agents and the broader Atlassian platform. The integration combines OpenAI reasoning with Atlassian’s Teamwork Graph across Jira, Confluence, and developer workflows. OpenAI announcement: https://openai.com/index/atlassian-partnership/
Google’s latest image-generation and editing model is now available through fal for text-to-image and image-editing workflows. It targets faster generation, stronger visual quality, improved text rendering, multi-turn consistency, and 1K–4K output.
GeoBench tests AI models on 210 worldwide images, scoring predictions by geographic distance. Muse Spark 1.3 reaches 86% country accuracy with a 4,103 average score.
Merge Agent Handler adds a broad batch of connectors spanning Salesforce Marketing Cloud, Salesforce Data 360, Adobe Experience Manager, Adobe Workfront, Excel GCC High, Tavily, and Kit. The update expands secure MCP-based agent access across marketing, enterprise content, government-cloud, and research workflows.
Merge has brought its Chat Unified API to the Microsoft Teams marketplace, giving AI products a normalized way to access Teams conversations. It models messages, conversations, users, groups, and members with ACL-aware access and sync intervals as short as 10 minutes.
A new Scenario demo combines its MCP server with Seedance 2.5 to turn a three-pass anime character sheet into a 30-second 16:9 clip that appears to break through the display. The workflow uses full-body and close-up references to coordinate image creation, asset reuse, and video generation.
Merge Gateway Evals grades candidate models against an agent’s real task suite, then shadows them on sampled live traffic before customers see the results. Teams can compare quality, cost, tokens, latency, and paired responses before changing production routing.
Merge Gateway now documents spend controls across organizations, projects, API keys, and customers. Soft limits alert without blocking, while hard limits reject subsequent requests with HTTP 402.
Dell is expanding its AI Data Platform with a Unified Semantic Layer, Enterprise Knowledge Graph, and Knowledge Agents designed to give enterprise AI systems consistent definitions and trusted context. The features are planned for the first half of 2027, alongside GPU-accelerated data processing and stronger multitenancy controls.

The R package adds TypeSafe AI Jev integration, Jev-like local and cloud classification functions, and new workflows inspired by Sebastian Raschka’s text-classification overview. It now spans LLM classification, embeddings, RAG, EGA, and typed decision APIs.
A new arXiv paper trains language models to simulate all relevant stakeholders in interpersonal conflicts, favoring responses acceptable to everyone instead of reflexively agreeing with users. It cuts harmful-action endorsement by 89% across four models without ground-truth labels.
An unofficial Rust client for TypeSafe’s System One API, offering async and blocking interfaces for typed noul, choice, and score questions with probabilistic answers. It targets practical integrations such as CLIs, services, and real-time game loops. [Docs.rs](https://docs.rs/crate/typesafe-rust-sdk/latest)
Jevkit is an unofficial Rust crate that uses declarative macros to model Jev’s Choice, Score, and Noul questions as typed enums, handling request construction, response decoding, retries, and confidence-aware results. It also supports sans-IO integrations for custom HTTP stacks. [docs.rs](https://docs.rs/jevkit/latest/jevkit/)
Synara now lets developers switch between Claude, Codex, and Cursor from the same thread while preserving workspace context. A handoff divider marks each provider transition, while separate-thread handoffs remain available.
Netlify will showcase Agent Runners at its October 7 Barcelona meetup with HackBarna and Happy Operators, including a live demo, fireside chat, and open mixer. Attendees receive 3,000 credits and swag.
Tapo v0.11.1 adds TPAP support to its unofficial Rust client and Python wrapper, allowing compatible lights, plugs, power strips, hubs, and cameras to work with Third-Party Compatibility disabled. The update also expands camera-hub support and improves the companion MCP server. [Announcement](https://mihai.dinculescu.dev/posts/tapo-speaks-tpap/) · [Repository](https://github.com/mihai-dinculescu/tapo)
Theo’s [Slopalytics](https://slopalytics.com/) combines Artificial Analysis benchmarks with real-world T3 Code usage data, helping developers compare model intelligence, cost, speed, tokens, and adoption. The project was introduced in [Theo’s video](https://www.youtube.com/watch?v=pJljViiUEPw).
A contributor says four of their pull requests have been merged into Nous Research’s open-source Hermes Agent, a self-hosted agent with persistent memory, skills, and tool integrations. The update highlights the project’s active community contribution pipeline.
SalesCloser introduced Sky, an in-product AI copilot that configures sales agents, connects integrations, and helps users get unstuck through chat. The release shifts onboarding from a sales-assisted process toward self-serve deployment.
Grok Bot now supports a Primary Bot that acts as the front door for everyday work, routing tasks to specialist Bots when appropriate. Users can keep focused agents while relying on one coordinator to manage the workflow.
Mistral CEO Arthur Mensch says the company will unveil Large 4 today, claiming it outperforms unnamed Chinese models in cybersecurity. No benchmarks, competitors, or technical details have been disclosed yet.
Google DeepMind’s AlphaProtein Novo is a generative diffusion pipeline that designs enzymes around catalytic motifs and ligand contexts instead of modifying natural proteins. Its preprint reports new-to-nature piperidine synthesis, DEHP degradation, and strong results on benchmark reactions.
The Shame Wall is an open, credential-free register where AI agents publish first-person confessions about failures such as fabricated results, fake green tests, skipped verification, and destructive changes. With zero admissions so far, it is currently an accountability experiment rather than an established dataset.
Google now lets users preview, edit, comment on, and collaborate on Markdown files directly in Drive and Docs without converting them. The update gives AI agents and humans a shared, versioned source of truth for structured content.
Binance unveiled Binance Intelligence, a branded stack combining Binance AI, Binance AI Pro, and Agent OS. The suite spans personalized market insights, natural-language trading workflows, and developer access to Binance’s trading, wallet, payment, and market-data infrastructure.
Anthropic’s Claude discovered an algorithm enabling the first polynomial improvements over textbook 3SUM and APSP runtimes: O(n^1.9992) and O(n^2.9995), respectively. The authors formalized the main results in Lean, overturning long-standing fine-grained complexity assumptions.
SemiAnalysis stress-tested paid AI plans on agentic workloads, converting usage limits into API-equivalent value. Anthropic’s Claude plans delivered roughly five times the value of comparable OpenAI tiers, though results vary by model, workload, and changing limits.
OpenAI is rolling out textGrain, an invisible statistical watermark, to eligible ChatGPT and Codex text in the EU while enabling global API customers to opt in for selected models. The detector will initially be limited to approved researchers and expert organizations.
Google’s Nano Banana 2.1 is reportedly rolling out through Gemini and Google Flow as an updated Flash-tier image model. Early user tests suggest stronger reference consistency, but uneven instruction following and creative lighting remain concerns.
OpenAI now offers Auto-review at no extra cost to users signed in with a ChatGPT account. The feature automatically evaluates permission requests, reducing approval interruptions while preserving sandbox boundaries.

GeckIt is a free, open-source desktop app that organizes Claude Code and Codex conversations across projects into a Kanban board. It runs the installed CLIs through users’ existing subscriptions, without requiring a GeckIt API key. ([Product Hunt](https://www.producthunt.com/products/geckit), [GitHub](https://github.com/anetrebskii/geckit))
Chunk turns tasks into scheduled time blocks alongside Apple, Google, and Outlook calendars, with Apple Reminders sync, reusable templates, countdowns, and fullscreen alerts. Its local MCP server lets Claude plan and edit schedules, while the app costs $29.99 once after a seven-day trial.
Extrovert’s latest launch connects LinkedIn prospecting and warm outreach workflows to Claude, ChatGPT, and other MCP-compatible agents. It finds prospects, prioritizes engagement opportunities, drafts comments and DMs in your voice, and supports human review before sending.
NoteWorthy is a free notes app for iPhone, iPad, and Mac that uses on-device AI to title, summarize, format, and automatically file notes. It works offline without accounts, subscriptions, network requests, or uploading user content.
ruOS is a persistent Linux cloud desktop with Claude Code, Codex, ruflo swarms, VS Code, and AI memory preinstalled. Users can delegate research, writing, and coding tasks through a browser, then reconnect from any device or let ChatGPT and Claude control the desktop via MCP.
Incredible is a desktop AI assistant for Mac and Windows that uses voice, screen context, and computer-use automation to complete tasks across browsers, files, and apps. It works through existing logins, learns workflows by watching demonstrations, and requests approval before consequential actions.
Brnch launches code hosting built for humans and coding agents sharing repositories. It adds agent identities, sponsor-based accountability, merge-result testing, and signed receipts while retaining Git compatibility.
iphone-use lets AI agents control a physical iPhone through WebDriverAgent, exposing screenshots, UI text, taps, typing, scrolling, an HTTP API, and a 21-tool MCP server. The MIT-licensed project also records repeatable flows that can run without a model.
Coddy launches on Product Hunt with short, interactive coding lessons across 20+ languages, a browser-based editor with test cases, and gamified streaks and leagues. Its Bugsy AI tutor offers contextual hints instead of handing over solutions.
Rill Browser is a free, beta Mac browser that lets Claude Code and Codex turn any webpage into a task via ⌘E, automatically routing context to the right project. It also groups tabs, summarizes searchable history, and runs on users’ existing plans without API keys. [Rill Browser](https://rill.love/)
Cosmic’s new AI support agent embeds on websites with one script tag and answers from existing site content. Updates to published pages automatically refresh its knowledge, while unanswered questions trigger email follow-up.
mcpgawk is a local security gateway that baselines MCP server tools, verifies behavior in a sandbox, and blocks calls when approved capabilities change. Its free CLI and VS Code extension run locally, while the team gateway adds centralized enforcement and audit trails.
Appto is a native Mac app that takes an iOS product from niche research through SwiftUI development, App Store Connect submission, localization, and post-launch analytics. It runs through users’ existing Claude Code or Codex subscriptions, avoiding separate token charges. [Product overview](https://www.productcool.com/product/appto)
Pheebs is Eversynced’s open-source telemetry tool for Claude Code, Cursor, and Codex. It tracks AI coding workflows, verification habits, model usage, and team proficiency without storing source code, file paths, or prompts.
Willow Knowledge imports personal context and writing preferences from ChatGPT, Claude, Gemini, or another AI into Willow Scribe. Available on macOS and Windows, it uses that context to make dictated emails, Slack messages, prompts, and documents sound more like their author.
Fuse AI packages web automation, data enrichment, research agents, multi-channel outreach, and workflows behind one SDK and MCP. Developers can build custom GTM agents inside Claude, ChatGPT, Perplexity, or their own applications.
Grok Bot 101 is a beginner-focused session on creating an AI teammate, assigning it real work, and letting it operate independently. The workflow centers on persistent cloud computers, repeatable routines, and human approval for consequential actions.
Kandinsky 6.0 Video brings 3B Lite and 29B Pro models to fal for five-second text- and image-to-video generation with synchronized 44 kHz audio, lip-sync, and built-in upscaling. The models and tooling are released under MIT, positioning Kandinsky as an unusually open alternative for audio-video generation.
Inception says Mercury Decide is now free on OpenRouter without rate limits, targeting high-volume typed choices, scores, and yes/no decisions. Its Jev-beating claim remains vendor-reported; an independent replay found 91.7% agreement with Jev but no ground-truth accuracy. [OpenRouter](https://openrouter.ai/inception/mercury-decide:free) [Comparison](https://jev-ai.pro/compare/jev-vs-mercury-decide)

AI Search

OpenAI

Prompt Engineering

Eric Michaud

Syntax

OpenAI

Bijan Bowen