AI Benchmarks

What is AICrier?

Live AI developer news, ranked and linked to original sources.

> ▌

BENCHMARK×
⌘K
FRIDAY // 2026-09-04
17 items
SEP 04
GPT-6 Just Did the Impossible... 99% AGI
PT16M14S
// WATCH
YouTube

GPT-6 Just Did the Impossible... 99% AGI

+3

AI Revolution

Introducing GPT-6 Astra for developers
PT3M24S
// WATCH
YouTube

Introducing GPT-6 Astra for developers

+3

OpenAI

GPT-6 Astra in Codex: How I Got Access EARLY Today (AND YOU CAN TOO)
PT2M32S
// WATCH
YouTube

GPT-6 Astra in Codex: How I Got Access EARLY Today (AND YOU CAN TOO)

+1

Income stream surfers

Google Gemini Released Gemini 3.8 & More Updates! (New Google Gemini Models & Use Cases)
PT10M44S
// WATCH
YouTube

Google Gemini Released Gemini 3.8 & More Updates! (New Google Gemini Models & Use Cases)

+7

Rob The AI Guy

🚨🚨 Working on Omarchy Automation Framework 🚨🚨
PT2H12M20S
// WATCH
YouTube

🚨🚨 Working on Omarchy Automation Framework 🚨🚨

The PrimeTime

Claude Code Improves Massively w/ 2nd Harness
PT35M40S
// WATCH
YouTube

Claude Code Improves Massively w/ 2nd Harness

+1

Discover AI

GPT-6 Astra: The harness matters more than you think
PT16M47S
// WATCH
YouTube

GPT-6 Astra: The harness matters more than you think

+5

Prompt Engineering

Use ChatGPT Work to analyze ad performance and refine creative
PT2M35S
// WATCH
YouTube

Use ChatGPT Work to analyze ad performance and refine creative

+2

OpenAI

Use ChatGPT Work to build custom creative tools
PT2M47S
// WATCH
YouTube

Use ChatGPT Work to build custom creative tools

OpenAI

Use ChatGPT Images to explore campaign concepts
PT2M1S
// WATCH
YouTube

Use ChatGPT Images to explore campaign concepts

OpenAI

Use ChatGPT Work to create campaign emails
PT2M48S
// WATCH
YouTube

Use ChatGPT Work to create campaign emails

+1

OpenAI

Use ChatGPT Work to pressure-test marketing campaign briefs
PT2M27S
// WATCH
YouTube

Use ChatGPT Work to pressure-test marketing campaign briefs

OpenAI

Use ChatGPT Work to turn marketing ideas into finished work
PT2M35S
// WATCH
YouTube

Use ChatGPT Work to turn marketing ideas into finished work

+1

OpenAI

I rebuilt 𝕏 from memory
PT16M11S
// WATCH
YouTube

I rebuilt 𝕏 from memory

Syntax

Fish Audio S2.1 Pro Full Test – Building A Video Translation Pipeline!
PT32M8S
// WATCH
YouTube

Fish Audio S2.1 Pro Full Test – Building A Video Translation Pipeline!

Bijan Bowen

GPT-6 Astra: OpenAI's Most Dangerous Model Yet
PT23M15S
// WATCH
YouTube

GPT-6 Astra: OpenAI's Most Dangerous Model Yet

+2

AI Samson

GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model anymore? - Worse than Fable?
PT14M22S
// WATCH
YouTube

GPT-6 Astra (Benchmarks Deep-dive): This is not a good coding model anymore? - Worse than Fable?

+2

AICodeKing

THURSDAY // 2026-09-03
6 items
SEP 03
Try Omarchy: Run Omarchy on MacOS without any setup
PT23S
// WATCH
YouTube

Try Omarchy: Run Omarchy on MacOS without any setup

Github Awesome

Google Just Dropped Its Most Powerful Cyber AI Yet
PT14M19S
// WATCH
YouTube

Google Just Dropped Its Most Powerful Cyber AI Yet

+4

AI Revolution

Is GPT-6 Astra Too Dangerous to Deploy? #openai #gpt6 #safety
PT2M44S
// WATCH
YouTube

Is GPT-6 Astra Too Dangerous to Deploy? #openai #gpt6 #safety

+2

DIY Smart Code

GPT-6 Astra Hits 99.9% on ARC-AGI-3
BENCHMARK// 1d ago

GPT-6 Astra Hits 99.9% on ARC-AGI-3

ARC Prize reports GPT-6 Astra scoring 99.9% on ARC-AGI-3 Semi-Private with a Provider Adapter harness, versus 62.7% under the standard harness. Astra also used fewer actions than median human participants on 96% of levels, marking a major interactive-reasoning milestone without proving AGI.

arc-agi-3benchmarkevaluationagentreasoningtool-use+6+5+4+3+2+1
Armature benchmarks coding agents’ tool picks
BENCHMARK// 1d ago

Armature benchmarks coding agents’ tool picks

Armature analyzed 16,893 coding-agent sessions across 75 repositories, 1,163 task variations, and Claude Code, Codex, and Cursor to see which third-party services they actually install. Only 42% of agent decisions matched, with programming language and repository context often changing the winner.

armaturecoding-agentai-codingagentbenchmarkevaluationtool-usedevtool+8+7+6+5+4+3+2+1
Qwen3.8-27B Makes 64K Local Context Practical
BENCHMARK// 1d ago

Qwen3.8-27B Makes 64K Local Context Practical

A hands-on comparison found Qwen3.8-27B scoring 98.7 across 21 tests while running a 64K context fully on a 24GB GPU. The dense open-weight model targets coding, research, professional work, and long-horizon agents.

qwen3.8-27bllmopen-weightslong-contextai-codingagentinferencebenchmark+8+7+6+5+4+3+2+1