Tuesday, September 29, 2026
Coverage: 51 videos reviewed (0 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.
New today
Anthropic releases Claude Sonnet 5.5 at $2 input, $10 output per million tokens
Anthropic released Claude Sonnet 5.5, the second model in the Claude 5.5 family, according to two channels reading Anthropic's announcement. Per that announcement as relayed, it is 30% faster and up to 30% less costly than Sonnet 5, with a 1M-token context and 128k output (Bijan Bowen) and a June 2026 cutoff; Fahd Mirza's screen showed a 262K context window. Both speakers said the price is $2 per million input and $10 per million output tokens, half of Opus 5.5; Mirza added $0.20 cache reads and availability on Amazon Bedrock. Anthropic's charts, as read by the speakers, show Sonnet 5.5 near Opus 5.5 and ahead of it at max effort, with a drop on Frontier code at max-to-xhigh effort that a footnote attributes to timeouts.
- Evidence: 0 first-party, 0 hands-on, 2 relaying
- Disagreements: Context window: Bijan Bowen relays roughly 1M tokens; Fahd Mirza's on-screen figure was 262K.
- Watch: Bijan Bowen: Claude Sonnet 5.5 Is INSANE – Seriously, This Model Is Ridiculous! (high hype)
OpenAI says GPT6 Astra reached its cyber critical threshold, citing ExploitGym and scope tests
OpenAI said GPT6 Astra is its first model to reach the cyber critical threshold, and listed safeguards including refusal training, abuse detection and blocking, tighter restrictions for higher-risk accounts and monitoring of reasoning and actions. In a slide described by the channel, GPT 5.6 Soul reached around 30 percent completion on ExploitGym while Astra achieved around 40 percent more successful completions with far fewer output tokens. OpenAI also said Soul without production safeguards exploited an out-of-scope target in about 48 percent of cases, versus zero for Astra. These are OpenAI's own results.
- Evidence: 1 first-party, 0 hands-on, 0 relaying
- Watch: OpenAI: The Defender's Window: Cyber security keynote
OpenAI announces Codex Security Red and adds GPT6 Soul and Luna to Daybreak Blue
OpenAI announced Codex Security Red, a managed penetration-testing offering with scope controls, isolated sandboxes for investigation agents and a guardian agent reviewing outgoing traffic, described as part of Daybreak. OpenAI said Daybreak Blue, its tier of general-purpose frontier models with safeguards for authorized security work, now includes GPT6 Soul and Luna, with GPT6 Astra to follow; Daybreak Red is the highest tier for approved red teams. No availability or pricing details were given in the items.
- Evidence: 1 first-party, 0 hands-on, 0 relaying
- Watch: OpenAI: The Defender's Window: Cyber security keynote
OpenAI says it paused frontier training runs to focus on monitoring and alignment
An OpenAI speaker said the company paused frontier training runs while doubling down on monitoring and alignment research, and that safety thresholds must be met before pushing capability further. An earlier speaker in the same presentation put the pause at a couple of weeks in early August. The statement is OpenAI's own and no independent confirmation was given.
- Evidence: 1 first-party, 0 hands-on, 0 relaying
- Watch: OpenAI: The Defender's Window: Cyber security keynote
GPT-6 Luna priced at 10 cents per million input tokens, 50 cents per million output
Bijan Bowen said GPT-6 Luna, described as the smallest and cheapest GPT-6 model, costs 10 cents per million input and 50 cents per million output tokens and replaces GPT 5.6 Luna. He said it has roughly 1M context, 128k output, text and image input and a May 18, 2026 cutoff. Vendor charts as he read them show a modest gain over its predecessor, for example 66.6% versus 62.2% on one benchmark at max effort (name garbled in captions).
- Evidence: 0 first-party, 0 hands-on, 1 relaying
- Watch: Bijan Bowen: GPT-6 Luna First Test – Is OpenAI’s CHEAPEST Model Actually Good?
Continuing stories
Also notable
- Bijan Bowen tests Sonnet 5.5 on game builds and robot arm; max effort takes about two hours - In hands-on tests, Bijan Bowen reported Sonnet 5.5 at max effort took about two hours on a browser OS build, and he concluded he would not run it at max. [0 first-party, 1 hands-on, 0 relaying] Watch: Bijan Bowen: Claude Sonnet 5.5 Is INSANE – Seriously, This Model Is Ridiculous! (high hype)
- OpenAI and Trail of Bits report 37 patches merged in first week of Patch the Planet - A speaker in OpenAI's presentation said the Patch the Planet initiative with Trail of Bits, covering projects such as Python, curl and Go, saw 37 patches merged in its first week. [1 first-party, 0 hands-on, 0 relaying] Watch: OpenAI: The Defender's Window: Cyber security keynote
- OpenAI reports under 1 percent false positives in its internal vulnerability defense factory - A field CTO in OpenAI's presentation said that in OpenAI's internal defense factory, dynamic validation cut false positives below 1 percent, ownership assignment reached about 90 percent and fix rollbacks were under 1 percent. [1 first-party, 0 hands-on, 0 relaying] Watch: OpenAI: The Defender's Window: Cyber security keynote
- Jev, a typed-output model, priced at 4 cents per million input tokens with free output - Three channels said Jev, from a company they name Type-Safe AI (as spoken), returns a choice, score or probability rather than text and costs 4 cents per million input tokens with no charge for output. [0 first-party, 0 hands-on, 3 relaying] Watch: How I AI: I’m using Jev more than Opus 5.5 or GPT-6. Here’s why.
- Hosts test Jev for PR clustering, command guardrails and file triage at low cost - How I AI reported that clustering about 1,700 ChatPRD pull requests with Jev cost 9 cents and took about 2 minutes, and that a product-insights pipeline made about 200,000 classifications for roughly four dollars on the Jev side. [0 first-party, 2 hands-on, 0 relaying] Watch: How I AI: I’m using Jev more than Opus 5.5 or GPT-6. Here’s why.
- Dylan Davis relays claims that per-token price misleads on per-task cost - Dylan Davis relayed several secondhand cost comparisons: an Anthropic developer's test in which Sonnet 5 needed about 30 to 40 turns versus 4 to 5 for Opus 4.8 and billed about twice as much, and a claim that GPT6 Astra runs 30 to 40 percent cheaper per task than Fable 5.1 at equal token prices. [0 first-party, 0 hands-on, 1 relaying] Watch: Dylan Davis: OpenAI's Own Team Says AI Token Pricing Is Meaningless
- MiniMax announces M3.1 Flash preview; KingBench 3 score of 66.25% in AI Code King test - MiniMax announced an M3.1 Flash preview for MiniMax Code on September 27, per AI Code King, which could not verify a benchmark table, API pricing or weights. [0 first-party, 1 hands-on, 0 relaying] Watch: AI Code King: Minimax M3.1 Flash (Fully Tested): Okay, this MODEL is PRETTY GOOD!
- PrismML releases Bonsai 2, a roughly 6 GB ternary compression of Qwen3.8 27B - PrismML released Bonsai 2, a ternary-weight (about 1.7 bits per weight) compression of Qwen3.8 27B from about 54 GB to about 6 GB, and claims it keeps 98% of benchmark performance. [0 first-party, 1 hands-on, 0 relaying] Watch: Prompt Engineering: Qwen 27B on 6GB VRAM...
- Orca team ships 12.3GB 3-bit compression of 54GB Qwen 27B model - Fahd Mirza reported the Orca team released a 12.3GB 3-bit compression of a 54GB Qwen 27B model, with a model card citing 262K context, thinking mode, tool calling and MTP speculative decoding. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: 12GB Model, 8 Hours, One 3D Game: OrcaSAQ2 27B Tested
- Bart Slodyczka measures 16GB M6 Mac mini at 998 tok/s prefill on Qwen 3.5 9B - In LM Studio tests with an 8,000-token input on Qwen 3.5 9B MLX 4-bit, Bart Slodyczka measured prefill of 222 tok/s on a 16GB M4 versus 998 tok/s on a 16GB M6 (about 4.5x), and decode of 19.3 versus 27.2 tok/s (about 40% faster). [0 first-party, 1 hands-on, 0 relaying] Watch: Bart Slodyczka: Don't Buy the 16GB M6 Mac Mini for AI (Until You Watch This)
Models & learning
- Peter Yang uses Sonnet 5.5 in Claude Code to make seven videos over a week - Peter Yang said he tested Sonnet 5.5 in Claude Code for a week and produced seven videos as code; some were one-shot and others needed iteration. [0 first-party, 1 hands-on, 0 relaying] Watch: Peter Yang: Sonnet 5.5 is Here! It's Insane at Making Videos (7 Incredible Example
- Fahd Mirza tests Sonnet 5.5 on multilingual prompts and a Hermes-agent trampoline game - Fahd Mirza reported that Sonnet 5.5 followed format and produced facts that looked right in a one-line test across many languages, including low-resource Saraiki, based on his spot checks of languages he knows. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: Claude Sonnet 5.5 First Day Tests — 3D Game, Physics, 80 Languages
- Bijan Bowen tests GPT-6 Luna: C++ skate game in 9 minutes, robot-arm task failed - In hands-on tests, Bijan Bowen reported GPT-6 Luna built a C++ skate game in 9 minutes at usable quality after a browser OS test needed a fix under 2 minutes, and a rally game took 23 minutes with rendering problems. [0 first-party, 1 hands-on, 0 relaying] Watch: Bijan Bowen: GPT-6 Luna First Test – Is OpenAI’s CHEAPEST Model Actually Good?
- AWS benchmarks Epic Lore 0.8.6: 105 GB import finishes in about 8m45s via edge pod - AWS presenters benchmarked Epic's Lore 0.8.6 across 21 server configurations with 1 to 29 clients and a 105 GB import. [0 first-party, 1 hands-on, 0 relaying] Watch: JetBrains: JetBrains GameDev Day 2026
- Nate Herk's seven-day Codex trading test with GPT-6 Astra ends about $100 down - Nate Herk reported a seven-day $10,000 trading test using Codex with GPT-6 Astra ended near $9,900, slightly behind the S&P 500 by his calculation, after he loosened the strategy twice. [0 first-party, 1 hands-on, 0 relaying] Watch: Nate Herk: I Gave GPT 6 Astra $10,000 to Trade Stocks...And This Happened
- Anthropic's Thariq advises minimal CLAUDE.md and effort levels by task type - Thariq predicted CLAUDE.md will eventually go away as model failure modes shift, advising users to start a project without it and add only repeated failure notes; he said Anthropic just added eval plugins for skills. [0 first-party, 0 hands-on, 1 relaying] Watch: Latent Space: The Future of Claude Code: Mods, Mutable Software, & Multiplayer Agent