Superintelligence, give or take. The daily AI briefing: what the AI world actually said, sorted by how much it matters.

Tuesday, September 29, 2026

Coverage: 90 videos reviewed (0 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.

New today

Anthropic releases Claude Sonnet 5.5 at $2 per million input tokens, $10 per million output
Anthropic released Claude Sonnet 5.5, the second model in its Claude 5.5 family, according to three channels reading the launch materials. Per the announcement as relayed, it is 30% faster and up to 30% cheaper for more work than Sonnet 5, with 1M context, 128k output and a June 2026 cutoff (Bijan Bowen); Fahd Mirza's screen showed a 262K context window. Listed prices are $2 per million input and $10 per million output tokens, half of Opus 5.5's $4 and $20; cache reads are $0.20 (Mirza). Anthropic's charts, as read by the channels, show Sonnet 5.5 near Opus 5.5 and passing it at max effort; Bowen noted a Frontier code score drop at max-to-xhigh effort that a footnote attributes to a cause he guessed was timeouts.
These are vendor charts and prices relayed by third parties, not independently measured. Nate Herk relayed Anthropic guidance recommending Sonnet for well-scoped work with checkable results and Opus for complex work needing judgment.

Anthropic released Claude Opus 5.5 on Sept. 22 with 1M context; vendor scores relayed
Anthropic released Claude Opus 5.5 on Sept. 22, 2026, the first model in the Claude 5.5 family, according to Julian Goldie. He said it has a 1 million token context window, up to 128,000 output tokens and adaptive thinking with an effort setting defaulting to medium. He relayed Anthropic-reported scores of 66.4% on Terminal Bench 4.0, 54.4% on Frontier Code, 81.8% on OSWorld 2.0 and 1846 Elo on a GDP-style benchmark (name as captioned).
The numbers are vendor-reported and were not independently verified, Goldie said. David Shapiro argued, without measurements, that Opus 5.5 produced a threshold effect in animation, CGI and 3D work.

Reviewers test Claude Sonnet 5.5 on coding, video and 3D tasks with mixed results
Several creators ran Claude Sonnet 5.5 through hands-on projects after release. Bijan Bowen said a max-effort browser OS run took about two hours, judged its GTA clone better than Opus 5.5's (subjective), and reported a robot-arm record after an erratic run; he said weekly usage rose from 7% to 14% across the tests. Fahd Mirza reported a Hermes-agent 3D trampoline game cost about $25-30 and mostly worked. Peter Yang made seven videos as code over a week, some one-shot and others iterated, and called it a cheaper Opus 5.5. Nate Herk ran seven same-prompt trials, in which Sonnet 5.5 won four and Opus 5.5 three, with winners often set by cost.
Matthew Berman said his team generated demos including a 3D ocean simulator and a Fall Guys clone in a couple of days; these are claims shown as demos, not scored benchmarks. Yang said Sonnet 5.5 could not make anime-style video on its own and needed an outside video API and music tool.

OpenAI released GPT-6 Astra on Sept. 3, with Soul and Luna variants, per one channel
OpenAI released GPT-6 Astra on Sept. 3, 2026 to approved users and the next day to others, according to Julian Goldie, who said it is available to ChatGPT Plus, Pro, Business and Enterprise and that OpenAI states 98% on Frontier Math Tier 4. He said GPT-6 Soul and Luna are trained the same way as Astra but built to be faster.
The account is a secondhand relay from a sponsor-tagged channel and was not checked against OpenAI materials in these items.

OpenAI says GPT-6 Astra reached its cyber critical threshold and shares ExploitGym results
OpenAI said GPT6 Astra is its first model to reach its cyber critical threshold, and listed refusal training, abuse detection, tighter limits for higher-risk accounts and monitoring of reasoning and actions as safeguards. On ExploitGym, an OpenAI slide showed GPT 5.6 Soul at around 30% completion and Astra about 40% more successful completions with far fewer output tokens. In a test with an out-of-scope shortcut, Soul without production safeguards exploited it in around 48% of cases versus zero for Astra.
All figures are OpenAI's own, from a slide description, and were not independently reproduced.

Typesafe AI's Jev decision model priced at 4 cents per million input tokens, free output
Typesafe AI's Jev returns choices, scores or probabilities rather than text and is priced at 4 cents per million input tokens ($42 per billion) with no charge for output tokens, according to Matt Wolfe, Nate B Jones, How I AI and IndyDevDan. A vendor video played on Liam Ottley's stream claimed it is 100 times faster and 100 times cheaper than language models. IndyDevDan, citing a TypeSafe comparison, said a million Jev calls cost about $20 versus $11,000 on Fable 5.1.
The speed and cost multiples are vendor claims; How I AI said it is currently free on Vercel's AI gateway.

Continuing stories

Also notable

Models & learning