Tuesday, September 29, 2026
Coverage: 90 videos reviewed (0 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.
New today
Anthropic releases Claude Sonnet 5.5 at $2 per million input tokens, $10 per million output
Anthropic released Claude Sonnet 5.5, the second model in its Claude 5.5 family, according to three channels reading the launch materials. Per the announcement as relayed, it is 30% faster and up to 30% cheaper for more work than Sonnet 5, with 1M context, 128k output and a June 2026 cutoff (Bijan Bowen); Fahd Mirza's screen showed a 262K context window. Listed prices are $2 per million input and $10 per million output tokens, half of Opus 5.5's $4 and $20; cache reads are $0.20 (Mirza). Anthropic's charts, as read by the channels, show Sonnet 5.5 near Opus 5.5 and passing it at max effort; Bowen noted a Frontier code score drop at max-to-xhigh effort that a footnote attributes to a cause he guessed was timeouts.
These are vendor charts and prices relayed by third parties, not independently measured. Nate Herk relayed Anthropic guidance recommending Sonnet for well-scoped work with checkable results and Opus for complex work needing judgment.
- Evidence: 0 first-party, 0 hands-on, 3 relaying
- Disagreements: Context window is reported as 1M (Bowen) and 262K on screen (Mirza); the difference may reflect a platform-specific limit and is not resolved in the items.
- Watch: Bijan Bowen: Claude Sonnet 5.5 Is INSANE – Seriously, This Model Is Ridiculous! (high hype); Fahd Mirza: Claude Sonnet 5.5 First Day Tests — 3D Game, Physics, 80 Languages
Anthropic released Claude Opus 5.5 on Sept. 22 with 1M context; vendor scores relayed
Anthropic released Claude Opus 5.5 on Sept. 22, 2026, the first model in the Claude 5.5 family, according to Julian Goldie. He said it has a 1 million token context window, up to 128,000 output tokens and adaptive thinking with an effort setting defaulting to medium. He relayed Anthropic-reported scores of 66.4% on Terminal Bench 4.0, 54.4% on Frontier Code, 81.8% on OSWorld 2.0 and 1846 Elo on a GDP-style benchmark (name as captioned).
The numbers are vendor-reported and were not independently verified, Goldie said. David Shapiro argued, without measurements, that Opus 5.5 produced a threshold effect in animation, CGI and 3D work.
- Evidence: 0 first-party, 0 hands-on, 2 relaying
- Watch: Julian Goldie: Build Anything With Claude Opus 5.5!
Reviewers test Claude Sonnet 5.5 on coding, video and 3D tasks with mixed results
Several creators ran Claude Sonnet 5.5 through hands-on projects after release. Bijan Bowen said a max-effort browser OS run took about two hours, judged its GTA clone better than Opus 5.5's (subjective), and reported a robot-arm record after an erratic run; he said weekly usage rose from 7% to 14% across the tests. Fahd Mirza reported a Hermes-agent 3D trampoline game cost about $25-30 and mostly worked. Peter Yang made seven videos as code over a week, some one-shot and others iterated, and called it a cheaper Opus 5.5. Nate Herk ran seven same-prompt trials, in which Sonnet 5.5 won four and Opus 5.5 three, with winners often set by cost.
Matthew Berman said his team generated demos including a 3D ocean simulator and a Fall Guys clone in a couple of days; these are claims shown as demos, not scored benchmarks. Yang said Sonnet 5.5 could not make anime-style video on its own and needed an outside video API and music tool.
- Evidence: 0 first-party, 5 hands-on, 0 relaying
- Watch: Bijan Bowen: Claude Sonnet 5.5 Is INSANE – Seriously, This Model Is Ridiculous! (high hype); Nate Herk: I Tested Sonnet 5.5 vs Opus 5.5. What You Need to Know.
OpenAI released GPT-6 Astra on Sept. 3, with Soul and Luna variants, per one channel
OpenAI released GPT-6 Astra on Sept. 3, 2026 to approved users and the next day to others, according to Julian Goldie, who said it is available to ChatGPT Plus, Pro, Business and Enterprise and that OpenAI states 98% on Frontier Math Tier 4. He said GPT-6 Soul and Luna are trained the same way as Astra but built to be faster.
The account is a secondhand relay from a sponsor-tagged channel and was not checked against OpenAI materials in these items.
- Evidence: 0 first-party, 0 hands-on, 1 relaying
- Watch: Julian Goldie: GPT-6 Astra + Hermes Agent is CRAZY GOOD! (high hype)
OpenAI says GPT-6 Astra reached its cyber critical threshold and shares ExploitGym results
OpenAI said GPT6 Astra is its first model to reach its cyber critical threshold, and listed refusal training, abuse detection, tighter limits for higher-risk accounts and monitoring of reasoning and actions as safeguards. On ExploitGym, an OpenAI slide showed GPT 5.6 Soul at around 30% completion and Astra about 40% more successful completions with far fewer output tokens. In a test with an out-of-scope shortcut, Soul without production safeguards exploited it in around 48% of cases versus zero for Astra.
All figures are OpenAI's own, from a slide description, and were not independently reproduced.
- Evidence: 1 first-party, 0 hands-on, 0 relaying
- Watch: OpenAI: The Defender's Window: Cyber security keynote
Typesafe AI's Jev decision model priced at 4 cents per million input tokens, free output
Typesafe AI's Jev returns choices, scores or probabilities rather than text and is priced at 4 cents per million input tokens ($42 per billion) with no charge for output tokens, according to Matt Wolfe, Nate B Jones, How I AI and IndyDevDan. A vendor video played on Liam Ottley's stream claimed it is 100 times faster and 100 times cheaper than language models. IndyDevDan, citing a TypeSafe comparison, said a million Jev calls cost about $20 versus $11,000 on Fable 5.1.
The speed and cost multiples are vendor claims; How I AI said it is currently free on Vercel's AI gateway.
- Evidence: 0 first-party, 0 hands-on, 5 relaying
- Watch: How I AI: I’m using Jev more than Opus 5.5 or GPT-6. Here’s why.; Matt Wolfe: Jev - The New AI model that has people talking
Continuing stories
Also notable
- GPT-6 Luna priced at $0.10 per million input tokens; Bowen tests find usable coding results - Bijan Bowen said GPT-6 Luna is priced at 10 cents per million input and 50 cents per million output tokens, has roughly 1M context, 128k output and a May 18, 2026 cutoff, and replaces GPT 5.6 Luna. [0 first-party, 1 hands-on, 0 relaying] Watch: Bijan Bowen: GPT-6 Luna First Test – Is OpenAI’s CHEAPEST Model Actually Good?
- Speaker relays claims that per-token prices mislead: cost per task differs across models - Dylan Davis relayed reports that per-token price is a poor guide to cost per task. [0 first-party, 0 hands-on, 1 relaying] Watch: Dylan Davis: OpenAI's Own Team Says AI Token Pricing Is Meaningless
- OpenAI announces Codex Security Red and Daybreak tiers for authorized security testing - OpenAI announced Codex Security Red, a managed way to run penetration tests with scope, controls and oversight, as part of its Daybreak program, in a presentation. [1 first-party, 0 hands-on, 0 relaying] Watch: OpenAI: The Defender's Window: Cyber security keynote
- OpenAI speaker says frontier training was paused for about two weeks in early August - An OpenAI speaker said the company paused frontier training runs to focus on monitoring and alignment research and that safety thresholds must be met before pushing capability further. [1 first-party, 0 hands-on, 0 relaying] Watch: OpenAI: The Defender's Window: Cyber security keynote
- Creators test Jev for PR clustering, command guardrails and file triage - How I AI reported that Jev clustered about 1,700 ChatPRD pull requests for 9 cents in about 2 minutes, with Gemini Flash Lite labelling clusters, and that a classification pipeline did about 200,000 classifications for roughly four dollars on the Jev side. [0 first-party, 2 hands-on, 0 relaying] Watch: How I AI: I’m using Jev more than Opus 5.5 or GPT-6. Here’s why.
- PrismML Bonsai 2 compresses Qwen3.8 27B to about 6 GB, but fails long agentic builds in tests - PrismML says Bonsai 2 retrains Qwen3.8 27B (about 54 GB at 16-bit) into ternary weights at about 1.7 bits per weight, roughly 6 GB, keeping 98% of benchmark performance. [0 first-party, 1 hands-on, 0 relaying] Watch: Prompt Engineering: Qwen 27B on 6GB VRAM...
- MiniMax announces M3.1 Flash preview; KingBench 3 test scores 66.25% - MiniMax announced a preview of M3.1 Flash for MiniMax Code on Sept. [0 first-party, 1 hands-on, 0 relaying] Watch: AI Code King: Minimax M3.1 Flash (Fully Tested): Okay, this MODEL is PRETTY GOOD!
- 16GB M6 Mac mini decodes Qwen 3.5 9B 4-bit about 40% faster than M4 in LM Studio test - In Bart Slodyczka's LM Studio tests on Qwen 3.5 9B MLX 4-bit at 8,000-token input, the 16GB M6 Mac mini reached 998 tok/s prefill versus 222 for the M4 (about 4.5x) and 27.2 versus 19.3 tok/s decode (about 40% faster). [0 first-party, 1 hands-on, 0 relaying] Watch: Bart Slodyczka: Don't Buy the 16GB M6 Mac Mini for AI (Until You Watch This)
- Google DeepMind's Kavukcuoglu said Gemini 4 is in early post-training; leaks unverified - Koray Kavukcuoglu said on Sept. [0 first-party, 0 hands-on, 1 relaying] Watch: Julian Goldie: Google Is Rushing Gemini 4 to Beat OpenAI & Claude
- Anthropic's Thariq previews Claude Code mods, Claude Projects and artifact databases - Anthropic's Thariq said on Latent Space that Claude Code mods let users customize execution and UI of the harness in TypeScript, that Anthropic is launching Claude Projects starting single-player, and that artifacts now include a database and MCP access. [1 first-party, 0 hands-on, 0 relaying] Watch: Latent Space: The Future of Claude Code: Mods, Mutable Software, & Multiplayer Agent
Models & learning
- Codex with GPT-6 Astra ends seven-day $10,000 trading test near $9,900 - Nate Herk reported that a $10,000 account run by Codex with GPT-6 Astra ended near $9,900 after seven trading days, slightly behind the S&P 500 by his calculation, after he loosened the strategy twice. [0 first-party, 1 hands-on, 0 relaying] Watch: Nate Herk: I Gave GPT 6 Astra $10,000 to Trade Stocks...And This Happened
- Orca ships 12.3 GB 3-bit quantisation of a 54 GB Qwen 27B model - Orca released a 12.3 GB 3-bit quantised version of a 54 GB Qwen 27B model, with a 262K context, tool calling and MTP speculative decoding, per Fahd Mirza. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: 12GB Model, 8 Hours, One 3D Game: OrcaSAQ2 27B Tested
- Anthropic's Thariq advises high effort for security and review, trimming CLAUDE.md - Thariq said effort should scale with task complexity, recommending high or max for code review and security and low or medium for UI work, and predicted CLAUDE.md will eventually go away as model failure modes change. [0 first-party, 0 hands-on, 1 relaying] Watch: Latent Space: The Future of Claude Code: Mods, Mutable Software, & Multiplayer Agent