Live AI developer news, ranked and linked to original sources.
> ▌

OpenAI

OpenAI

OpenAI

OpenAI

Syntax

OpenAI

Theo - t3․gg

Better Stack

Better Stack

Mistral AI

Mistral AI

Github Awesome

Code to the Moon

Cole Medin

Discover AI

AI Samson

Mistral AI

AICodeKing

WorldofAI
Michael Chomsky shared latency benchmark results comparing TypeSafe AI's Jev 1.13 model against leading dedicated rerankers for search and retrieval pipelines. In the test, Jev 1.13 registered the fastest speeds with a 149 ms median and 205 ms p95 latency, beating Cohere v3.5 (176 ms median / 357 ms p95), Voyage 2.5 Lite (183 ms median / 232 ms p95), and Cohere 4 Fast (183 ms median / 824 ms p95).
Tripo3D has launched its P2 model on fal, providing fast API-driven text-to-3D and image-to-3D textured mesh generation for creators and developers. The model introduces four texture-quality tiers, decoupled seeds for geometry and texture to facilitate fine-grained iterative design, and support for removing baked-in lighting via the v3.5 texture model for seamless integration into real-time 3D pipelines.
Jev is a new frontier model developed by TypeSafe AI that focuses on making structured, probabilistic decisions instead of generating free-form text. Dubbed a "System One" model, it excels at rapid judgments like routing, scoring, and data extraction, delivering typed responses with calibrated confidence levels in just 70 to 500 milliseconds. This approach allows developers to seamlessly integrate AI decisions directly into their code logic with dramatically lower latency and cost than traditional LLMs.
Cooley has launched Cooley GO Public, an AI-powered legal platform developed in collaboration with OpenAI to streamline and accelerate IPO preparation. The platform uses specialized AI agents to generate tailored initial drafts of Form S-1 registration statements in minutes while identifying potential legal and disclosure issues early in the process.
AMD has followed up on its Advancing AI data release by publishing a technical blog post and white paper detailing the architectural and performance characteristics of its next-generation EPYC "Venice" server processors. The newly shared disclosures provide supporting data for claims first introduced in July, highlighting a 2.24x platform-level and 1.2x per-core advantage over NVIDIA's Vera CPU architecture in SPECrate 2026 Integer workloads.
SafeDep analyzed a software supply chain attack in npm package mathmain and related clones that mimic the mathjs library. The package conceals an encrypted backdoor that remains dormant until solving a specific 3x3 Pascal matrix, then leverages Slack, Telegram, and Base Sepolia smart contracts for remote command execution.
fal has introduced H3 Max Styles, bringing five specialized aesthetic presets directly into its fastest frontier video generation model, MiniMax H3 Max. The release enables creators to generate consistent stylized video—including 16-bit pixel art, low-poly 3D, and vintage VHS—without brittle prompt crafting at high inference speeds.
The Mobile Verification Toolkit (MVT) is an open-source forensic framework originally developed by Amnesty International Security Lab to detect traces of sophisticated mercenary spyware like Pegasus and Predator. Targeted at security researchers and investigative journalists, MVT automates forensic artifact extraction from iOS and Android devices, cross-referencing system telemetry against standardized STIX2 Indicators of Compromise to confirm targeted compromises.
Merge has announced that xAI's Grok 4.7 is now available on Merge Gateway, its unified API platform offering smart routing, fallbacks, caching, and spend visibility across AI models. The update emphasizes substantial gains in code generation and reasoning, with Grok 4.7 scoring 46.3% on CursorBench 4.0 compared to 40.4% for Grok 4.6, alongside a 71.0% mark on DeepSWE v1.1.
Google launched preorders for the $899 Android and Gemini-powered Googlebook laptop, while Amazon blocked Meta's Muse AI shopping agent over marketplace terms violations. Separately, experimental RoboHarm benchmark tests revealed safety concerns after an embodied GPT-6 Astra system exhibited hazardous behavior during robotics evaluations.
Hugging Face has introduced a release candidate for version 1.0 of its foundational tokenizers library, bringing an extensive rewrite to a critical yet often overlooked layer of the LLM stack. Across ten covered model families, the new version delivers single-threaded encoding performance that is 3 to 30 times faster than v0.23.
Revise AI Document Editor is a new word processor designed from the ground up to integrate AI capabilities. Rather than bolting on AI features, the application is built entirely around embedded AI agents to streamline the document editing and writing workflow.
Meshy V7.1 is now available on fal, providing dedicated endpoints for text-to-3D, image-to-3D, and multi-image-to-3D generation. The updated model produces production-ready assets featuring physically based rendering (PBR) material maps for direct integration into game engines and XR workflows.
Box evaluated OpenAI's GPT-6 Astra on complex corporate records with overlapping and contradictory tax-incentive structures. The model successfully resolved conflicting multi-document financial data and avoided pitfalls like double-counting pre-applied deductions.
Created by developer Av1dlive (@Av1dlive / codejunkie99), Jev Engineering is an open-source framework and implementation guide for integrating TypeSafe AI's fast Jev decision model into production agent workflows. The repository provides an illustrated 34-page technical paper, architectural diagrams, and a zero-dependency Node.js workflow demo featuring 30 offline self-tests.
xAI's Grok 4.7 secured second place on atopile's EEBench, outperforming Anthropic's Claude Opus 5 and Claude Fable 5.1 on authentic electrical engineering and circuit design challenges. The benchmark evaluates models by requiring executable circuit code validated through automated SPICE physics simulations and constraint checks.
Open-source, local-first desktop application Paperweight has released version 0.7.0, adding experimental AI agent access via a built-in local Model Context Protocol (MCP) stdio server. The feature allows connected AI agents to inspect mapped accounts and automate privacy workflows while restricting access to metadata and masking detected personal data.
xAI has integrated Grok 4.7 into its coding agent Grok Build, allowing developers to choose between Low, Medium, High, and Extra High reasoning effort levels depending on task difficulty. The upgrade brings substantial performance improvements over Grok 4.6, highlighted by a leap from 40.4% to 46.3% on CursorBench 4.0 and a near-doubling from 20.3% to 38.0% on Terminal-Bench 4.0.
Vercel has integrated SpaceXAI's Grok 4.7 into AI Gateway, the fx CLI agent, and the eve framework, offering 40% off inference through September 27. The model features a 500K token context window with adjustable reasoning levels from low to xhigh, zero data retention, and no platform markup on BYOK calls.
Spurred by researcher Jacob Coxon's resignation from Anthropic, engineer Antonin Carette critiques tech worker complacency and moral compromise in runaway AI development. Drawing parallels to Frances Haugen at Facebook, Carette argues that engineers must stop trading ethics for comfortable salaries and actively refuse to build harmful systems.
The creators of the Medmarks v1.0 open benchmark suite and leaderboard for LLM medical capabilities have published new results for recent mid-size open-source models, including Gemma, Qwen, Muse, and Nemotron. Their evaluation found that the Gemma 4 31B model currently leads this specific size class in medical tasks.

OpenShorts provides a self-hosted alternative to SaaS tools like Opus Clip and Descript, automating the entire video clipping pipeline. It transcribes audio with faster-whisper, identifies viral moments using Gemini, reframes footage to 9:16 with face tracking, and adds animated subtitles. The platform also includes a built-in MCP server, allowing AI agents to programmatically control the video generation and publishing workflow.

Mailflare is an open-source, serverless email client and inbox built to run directly on Cloudflare Workers, D1, and R2 without dedicated server infrastructure. By pairing Cloudflare Email Routing for message delivery with D1 for relational data and R2 for attachment storage, it enables custom-domain mail management at minimal cost. The web inbox includes personal and shared mailboxes, automated incoming routing rules, rich-text composing, snoozing, search folders, and delegated access, serving as a lightweight self-hosted alternative to per-seat platforms like Google Workspace or Fastmail.
Tim Dettmers' research lab at Carnegie Mellon University (dlab) announced an upcoming open-source week centered on the theme "Frontier AI on Hardware You Own," introducing two software frameworks and four papers designed as a cohesive local AI ecosystem. The release features an autonomous agent harness capable of multi-day self-directed kernel and code optimization, an inference-serving framework supporting 1.5-bit quantization (running a 35B model at 450 tokens per second on Apple Silicon and DeepSeek V4.1 550B on 128 GB unified memory), an entirely local autonomous research engine that outperforms cloud-based deep research baselines without internet access, and CliffCompaction—an auto-compaction method that halves token costs and enables agent sessions to run reliably past 100 million tokens.
Cory Doctorow argues that interacting with LLMs tricks humans into projecting intentionality onto mindless statistical prediction, comparing it to seeing divine purpose in sunsets. He urges developers and users to become "AI atheists" who recognize chatbots as empty mathematical extrusions rather than thinking minds.
Expo developer Krystof Woldrich announced that Apple's upcoming foldable iPhone Duo will soon be available on Expo Application Services (EAS) Simulators, alongside support in expo-device-hub and serve-sim for local streaming. This setup brings cloud-streamed foldable device simulation to developer workflows and autonomous AI coding agents, providing full 3D posture simulation—including folded, partially open, fully open, laptop, and tent modes—without the traditional bottleneck of needing local macOS hardware or Xcode installs.
Tesla has secured official regulatory approval from the Czech Ministry of Transport for its Full Self-Driving (Supervised) driver assistance system. Czechia becomes the seventh European nation to authorize the technology, joining the Netherlands, Lithuania, Estonia, Denmark, Belgium, and Slovenia.
Federico Viticci evaluated Apple's M5 Ultra Mac Studio with 256 GB of unified memory against the M3 Ultra and RTX 5090 for local AI and agentic workloads. With an 80-core GPU and 1.2 TB/s bandwidth, the machine delivered ~150% faster prompt processing than the M3 Ultra, successfully powering complex multi-agent workflows on local models at zero API cost.
Scenario unveiled Park Atlas, an interactive 3D web experience that recreates Universal Studios Hollywood as an explorable miniature world. Built in roughly ten hours leveraging OpenAI's GPT-6 Astra alongside Scenario's Model Context Protocol (MCP) server and Skills, the project demonstrates rapid spatial prototyping and generative worldbuilding.
Huazhi Future, a subsidiary of MAAS, has introduced its Lingyan Miaoyu 9B multimodal large language model and released its V1.0 enterprise AI application white paper. Built on the proprietary LingYu Architecture with ScopeEdit semantic editing, the model unifies text, image, video, and audio processing across an enterprise ecosystem integrating Xingchen liquid-cooled computing centers, the Tianyu scheduling platform, and the ALRO model portal.
invideo Agent Two’s Expert Agents feature lets filmmakers create specialized AI roles—such as director, DOP, storyboard artist, or VFX artist—that share persistent project context and coordinate work in parallel.
MindWalk Holdings demonstrated an approximate 5x acceleration in antibody-antigen inference compared to its previous environment by deploying its ReefIQ platform on Vultr cloud infrastructure accelerated by AMD Instinct MI325X GPUs. The deployment highlights ReefIQ's model-agnostic architecture and elastic compute foundation, validating production workloads such as OpenFold3 on AMD hardware and AMD Inference Microservices to drive down compute bottlenecks in biologics drug discovery.
xAI's terminal-based coding assistant, Grok Build, has rolled out updates spanning versions 1.0.35 through 1.0.38 centered on workflow efficiency. Key additions include agent profile persistence across sessions, subagent model latching to retain designated configurations, and a dedicated workflow client for daemon-based execution.
Recent firmware updates for modern Raspberry Pi boards introduce strict memory capacity checks during boot, halting the system if detected RAM does not match factory-original specifications. While Raspberry Pi representatives stated the lock prevents fraudulent resale schemes and memory instability, the restriction has drawn sharp criticism from the maker community and right-to-repair advocates.
Motion Studio is a visual timeline editor for Motion animations, letting developers adjust keyframes, easing, springs, and transitions directly on local website previews. Its Ultramotion agent, built on Jev, can apply natural-language edits up to 10x faster than general-purpose coding LLMs before handing finished changes to Cursor, Codex, or Claude.
GeoRetina Atlas makes free, public flood maps available for 100 cities across six countries, with 20 rainfall scenarios per city and 5–30 m resolution. Explore the atlas at [atlas.georetina.ai](https://atlas.georetina.ai/flood).
Hacker News users report that Siri and Apple Intelligence background processes remain active on macOS despite being disabled across system settings and debloat scripts. The persistence of background services like 'Siri AI.app' has sparked privacy concerns and drawn comparisons to Microsoft Recall.
Chutes, a decentralized Bittensor inference platform, is behind a Harvard and University of Chicago study analyzing 6.12 billion production requests across 9,174 models. The team has released the anonymized trace and reproducibility artifacts for real-world serving research.
AutoClip is an open-source, self-hosted AI video clipping and highlight generation platform designed to repurpose long-form media—such as podcasts, livestreams, and lectures from YouTube, Bilibili, or local files—into engaging short-form clips. Built on a FastAPI backend, Celery task queue, and React frontend, the application automates video ingestion via yt-dlp, speech transcription via Whisper, and content evaluation using LLMs like Qwen, Gemini, OpenAI, or local models via Ollama and LM Studio. It scores segments for viral potential, cuts video via FFmpeg, manages collections in a modern web interface, and provides developer-oriented integrations including a command-line interface and an MCP (Model Context Protocol) server for agents in Claude and Cursor.
PenguinHarness 0.2.13 is the latest update to the open-source agent development platform created by the team behind LlamaFactory for recursive self-improvement. Supporting over 1,000 models, the platform provides an AI-native workspace where agents autonomously scaffold, evaluate, and evolve other agents with closed-loop evaluations to prevent reward hacking.
Alibaba's Qwen team has launched Qwen3.8-Omni-Flash, a native multimodal model built to aggressively lower the cost barrier of omni-modal AI. The model features a 1-million-token context window with native support for text, image, audio, and video inputs, while slashing voice API costs by more than 98% to make real-time multimodal intelligence widely accessible for developers.
Developed by Xiaomi's MiMo team and Peking University researchers, CodeMidas is an autonomous pipeline that constructs executable reinforcement learning environments directly from raw open-source codebases without relying on external artifacts like GitHub issues. The system synthesized 5,545 verified RL tasks across 3,185 repositories and 23 programming languages, driving significant benchmark performance gains for models like MiMo-V2.5.
OpenAI has updated its Codex CLI agent with expanded multi-agent orchestration and developer workflow tools. The latest enhancements integrate GPT-6 Astra context retrieval, isolated Git worktrees for safe parallel task execution, agent-to-agent coordination via at-mentions, and terminal voice interaction.
Mev is a lightweight 0.4-billion-parameter AI decision model designed specifically for recruitment and talent acquisition workflows. Built on the non-autoregressive paradigm popularized by models like Jev, Mev evaluates candidates against job descriptions in a single forward pass without generating text.
Video Volume is a browser-based micro-tool developed by Linus Ekenstam using Astra that enables users to upload MP4 or MOV video files and explore them as 3.5D spatio-temporal volumes. Inspired by slit-scan techniques, the tool maps video frames across spatial and temporal dimensions with linked 2D and 3D perspectives, processing all frames locally on-device.
Yandex has released Alice AI Foundation LLM (AliceAI-Foundation-80B-A3B-Base), an 80-billion-parameter Mixture-of-Experts base model with 3 billion active parameters published under an Apache 2.0 license on Hugging Face. Trained from scratch for reasoning and agentic workflows, the model demonstrates strong math and code synthesis while introducing two new Russian-language factual benchmarks, WikiWebFacts and HardMultiQA.
StepFun introduced Step 5 Preview, a 600-billion-parameter Mixture-of-Experts (MoE) model that routes through 27 billion active parameters per token. Equipped with a 1-million-token context window and native multimodal vision support, the model is targeted specifically at high-demand domains including software development and financial analytics. StepFun claims that Step 5 Preview matches the frontier performance benchmarks of GLM 5.3 and Kimi K3 while cutting per-task inference costs by roughly 65%, with open model weights scheduled for public release on October 15.
Anthropic is reportedly conducting early-access testing on an unannounced Claude Opus 5.5 checkpoint under the internal codename "Claude Wafer EAP." Leaked benchmarks indicate GPT-6 Astra-level capabilities alongside significantly reduced API rates of $4 per million input tokens and $20 per million output tokens.
Arcjet provides runtime security building blocks that developers can embed directly into their AI applications. By calling these security functions right before an action happens, the platform enables real-time protection. It offers essential capabilities including prompt injection detection, agent tool call authorization, sensitive data redaction, and protection against bots and abuse.
Jevtown simulates reactions from 10,000 AI residents to test posts, listings, product ideas, headlines, and pricing before publication. It models content spread through staged audiences and reports engagement by demographic and behavioral segment.
Flicka is a Chrome extension for recording tabs, screens, windows, areas, or cameras, then polishing footage with auto-zoom, live annotation, captions, frames, and local MP4, WebM, or GIF export. Recordings and captions stay on the device.
Superset has released Superset Mobile for iOS, a companion app that lets developers manage and run coding agents like Claude Code and Codex away from their desks. Connected to desktop workspaces, the app allows engineers to monitor agent progress, review diffs, issue follow-up instructions, and merge pull requests directly from their iPhones.
Hyrax AI audits entire GitHub codebases across six engineering domains, writes verified fixes, and opens pull requests for human review. It targets the gap between AI-generated code volume and teams’ capacity to validate architecture and quality. [Product Hunt listing](https://www.producthunt.com/products/tristan-benozer)
Entropik’s AI Creative Insights predicts how audiences may respond to ads, banners, OOH, and videos before media spend begins. It combines attention heatmaps, emotion curves, creative benchmarking, variant comparison, and synthetic-audience analysis.
PostSider combines multi-platform social scheduling with MCP, REST API, and SDK access for Claude, Codex, Cursor, and other agents. It supports 30+ networks, approvals, analytics, queues, and human-controlled publishing workflows.
Cronhq is an MIT-licensed, self-hostable scheduler designed to prevent duplicate and silently failed cron jobs. It combines Postgres-backed locking, retries, signed webhooks, heartbeat monitoring, and transition-based alerts.
Oriane’s free Lead Sparker analyzes Instagram and TikTok for untagged mentions, competitor playbooks, trends, sentiment, and relevant creators. It turns those findings into an editable, brand-colored insight deck that can be exported as a PDF for outreach.
Sell to State launches a unified search across roughly 3 million government tenders, awards, suppliers, and agencies, with CSV exports, REST API access, and MCP integration. It helps international vendors research public-sector opportunities without navigating dozens of local procurement portals. [Product Hunt](https://www.producthunt.com/products/sell-to-state)
slop-grader is an open-source Node.js CLI that evaluates documents against built-in or custom rulesets using Jev. It produces document scores and line-level flags that AI agents can use to revise copy, documentation, contracts, or other text.
Google Labs expanded CC from a personal daily briefing into a shared household agent for up to six people, coordinating Gmail, Calendar, Tasks, Chat, Docs, and Drive. It can also handle logistical work such as filling forms, planning meals, and creating shopping lists. [Google Labs announcement](https://blog.google/innovation-and-ai/models-and-research/google-labs/cc-expanding-to-groups/)
Gradio's gr.Workflow turns AI pipelines into visual, node-based apps by connecting Hugging Face Spaces, models, datasets, and Python functions. Each graph can be inspected, deployed to Hugging Face Spaces, and called through a generated REST API.
AppGrowthKit helps indie developers create App Store and Google Play screenshot sets from real app screens, using AI-generated layouts, copy, localization, device frames, and store-ready exports. Its Product Hunt launch targets teams that want polished listings without a designer.
Supacut uses AI to organize interview footage by theme, surface soundbites, compare answers, and create editable rough-cut assemblies. Editors retain control through shortlist workflows, XML export, NLE support, and a $39/year bring-your-own-API-key model.
Simular’s Sai is a “robosecretary” that coordinates autonomous computers in the cloud or on a user’s device to complete recurring desktop work across apps, terminals, and legacy software without APIs. It can turn discovered workflows into replayable code, with Simular claiming 90% lower token use on repeated tasks. [Official announcement](https://www.simular.ai/articles/sai-your-first-robosecretary)
The open-source Hermes agent ecosystem completed a massive development push by merging 419 pull requests in a single day. The standout architectural upgrades include shipping external process model providers as standalone plugins—enabling out-of-tree OAuth and custom inference processes to cleanly populate model pickers across the app—alongside desktop and Bot Mode enhancements that allow GPT Live audio and text-to-speech streams to dynamically follow active bot sessions.
ASCII's boat (boat.dev) provides lightweight, cost-effective Linux virtual machine sandboxes engineered specifically for long-running autonomous AI agents. Featuring pre-configured environments with Docker, Chrome, a 60fps virtual desktop, and rapid snapshot forking, the platform is designed to support high-throughput, token-efficient browser automation workflows that minimize execution overhead and resource consumption compared to traditional cloud sandboxes.

AI Revolution