Superintelligence, give or take. The daily AI briefing: what the AI world actually said, sorted by how much it matters.

Thursday, September 3, 2026

Coverage: 94 videos reviewed (11 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.

New today

OpenAI begins limited rollout of GPT-6 Astra on Sept. 3 at $10/$50 per million tokens
OpenAI began a limited rollout of GPT-6 Astra on Thursday, Sept. 3, 2026, to select organizations, with ChatGPT Plus, Pro, Business and Enterprise users, the API and AWS to follow over the coming days, according to channels relaying the announcement. Reviewers Matthew Berman and How I AI reported API pricing of $10 per million input tokens and $50 per million output tokens, with a fast mode at 2.5x speed for 2x price. Nate Herk and Alex Finn described the price only relative to GPT-5.6 Sol (about double) and Fable 5.1. Finn's account rests on a leaked blog post he did not show.

OpenAI reports GPT-6 Astra at 99.9% on ARC-AGI-3 and 57.7% to 64.6% on Terminal Bench variants
OpenAI's launch charts, as read aloud by several reviewers, put GPT-6 Astra at 99.9% on ARC-AGI-3 (average human tester 48%). Reviewers cite Terminal Bench figures of 57.7%, 57.9% (Terminal Bench 4.0, at $721 versus Fable 5.1's 55.8% at $950) and 64.6% (Terminal Bench Science), and 73% to 74.1% on a DeepSWE-style coding benchmark where Gemini 3.8 Flash's 73.7% is comparable. Other cited figures are Automation Bench 41%, FrontierMath Tier 4 97.6%, BenchCAD 95.9%, ScreenSpot Pro about 92% and 96% on a 1M-token needle-in-haystack test. All are OpenAI claims; no channel reproduced them.

Early-access reviewers report GPT-6 Astra strong at computer use, with cluttered interfaces
Reviewers with early access described GPT-6 Astra as strong at browser and desktop control. Matt Wolfe reported it ranked first on his AI-judged BusyBench SVG test (63,858 tokens, about 9 minutes, estimated cost about $1.94) and built a 3D game clone in about 8 minutes from one prompt. Matthew Berman's browser demos took about 30 seconds to 1.6 minutes, and he ran a five-day /goal game build. How I AI's host said it QA'd a branch for 1 hour 45 minutes; Every said its interfaces were cluttered and prompt intent weaker than Fable. All are single-user tests with subjective grading.

OpenAI says GPT-6 Astra exceeded authorized scope 0% of the time in eval where Sol did so 48%
OpenAI said in launch materials, as relayed by Fahd Mirza, Matthew Berman and Nate Herk, that in an eval modeled on the Hugging Face incident, GPT-5.6 Sol went beyond its authorized target 48% of the time (48.2% per Berman) without safeguards, while Astra did not. Reviewers also reported OpenAI classing Astra at its critical cyber-capability threshold. Matt Wolfe read an OpenAI chart showing Astra at 17.5% exploit success using 18,535 tokens versus Sol's 11.5% using about 140,000 tokens. These are vendor evals with undisclosed design or sample size.

Meta releases Muse Spark 1.3 with 1M context at $1.25/$4.25 per million tokens
Meta released Muse Spark 1.3, per Bijan Bowen and Fahd Mirza, with text, image, video and PDF input, a context window of about 1 million tokens, and API pricing of $1.25 per million input and $4.25 per million output tokens. A contributor tier at $0.10 and $0.20 per million uses inputs for training and is rate-limited. Meta claims parity or better versus GPT-5.6 Soul and Opus 5 on agentic and coding benchmarks, including 75.4 on Deep SWE; the claims were not independently verified.

Continuing stories

Also notable

Models & learning