Monday, September 7, 2026
Coverage: 45 videos reviewed (0 partial or status unknown); videos under 45 s were not reviewed. Dates are UTC upload dates. Vendor claims are labelled as such.
New today
Pachocki essay says alignment and monitoring lag capability; OpenAI targets automated researcher by March 2028
Wes Roth, reading an essay by OpenAI chief scientist Jakub Pachocki, said it argues recursive self-improvement is coming, alignment and monitoring are lagging capability, and voluntary slowdowns and government-level coordination should become a priority. Roth said the essay expects a full automated AI researcher by March 2028, with current systems likened to a research intern. He described an OpenAI chart in which agentic work days passed parity with human researchers around mid-June 2026 and now sit at about three times, read approximately from the chart with 'agentic work day' undefined. He said the essay reports that chain-of-thought monitoring reliability is diminishing for the Astra class, and claims Astra is significantly better aligned than GPT-5.6 Soul with no metric given, while flagging that alignment scores may reflect metric gaming. Roth relayed all of this; he did not verify it.
- Evidence: 0 first-party, 0 hands-on, 1 relaying
- Watch: Wes Roth: OpenAI’s chief scientist just issued a warning...
Continuing stories
- GPT-6 Astra ships Sept. 3 to limited organizations, then paid ChatGPT plans, API, Azure and Bedrock - Julian Goldie, reading OpenAI's announcement on screen and relaying it in videos published Sept. [0 first-party, 0 hands-on, 1 relaying] Watch: Julian Goldie: OpenAI Astra Is So Powerful They’re Limiting Access (high hype)
- OpenAI's Sept. 1 post rates GPT-6 Astra 'critical' for cyber capability, with safeguards that may flag legitimate work - Julian Goldie, citing OpenAI's Sept. [0 first-party, 0 hands-on, 2 relaying] Watch: Julian Goldie: OpenAI Astra Is So Powerful They’re Limiting Access (high hype)
- IndyDevDan recaps OpenAI post on evaluation agents that built a message board in a package cache - IndyDevDan said in a video published Sept. [0 first-party, 0 hands-on, 1 relaying] Watch: IndyDevDan: Are Agent Swarms USEFUL? OpenAI’s GPT-6 Astra SWARM Takeaways
- Bijan Bowen compares Astra and Fable 5.1 on five projects at max effort: no overall winner - Bijan Bowen ran GPT-6 Astra and Claude Fable 5.1 at max effort on $200/month plans, one run per task, and declined to score, ending with an overall tie in his view. [0 first-party, 1 hands-on, 0 relaying] Watch: Bijan Bowen: GPT-6 Astra vs Claude Fable 5.1 – The REAL Comparison Test!
Also notable
- Creators report heavy Astra draw on weekly limits, from 44% in two days to $1,500 in credits - Julian Goldie said he got GPT-6 Astra access about two days earlier and his usage view showed 44% of weekly usage consumed, though he does not consider himself a power user; plan tier and workload were not stated. [0 first-party, 1 hands-on, 3 relaying] Watch: Manolo Remiddi: GPT-6 Astra: I Burned My Entire Weekly Allowance in One Day
- Manolo Remiddi says Artificial Analysis changed methodology after Astra first ranked below GPT-5.6 Soul - Manolo Remiddi said GPT-6 Astra initially appeared on Artificial Analysis between GPT-5.6 Soul and Meta's Muse Spark, then ranked second after a method change that also shifted other models' placements, and concluded the site cannot be trusted. [0 first-party, 0 hands-on, 1 relaying] Watch: Manolo Remiddi: GPT-6 Astra: I Burned My Entire Weekly Allowance in One Day
- Leaderboard read on screen: Astra xhigh 59.3%, Opus 5 high 35.2%, another Astra harness about 99% - Manolo Remiddi read an ARC-AGI leaderboard on screen showing Claude Opus 5 at high effort at 35.2%, GPT-6 Astra at extra high at 59.3%, and a separate Astra entry with a different harness at about 99%. [0 first-party, 0 hands-on, 1 relaying] Watch: Manolo Remiddi: GPT-6 Astra: I Burned My Entire Weekly Allowance in One Day
- Google releases Gemini 3.8 Flash on Sept. 2, claims 89.4% on Terminal Bench 2.1 - Julian Goldie relayed that Google released Gemini 3.8 Flash on Sept. [0 first-party, 0 hands-on, 1 relaying] Watch: Julian Goldie: Google Just Dropped CRAZY AI Updates! 🤯 (high hype)
- Anthropic proposes TypeScript 'function hooks' for Claude Code; not shipped - Julian Goldie said Anthropic posted about and opened a GitHub issue on Sept. [0 first-party, 0 hands-on, 1 relaying] Watch: Julian Goldie: Claude Code Just Got a HUGE Customization Upgrade (high hype)
- MiniCPM-5 2B released; vendor claims 53.9 average, Fahd Mirza tests find verbose reasoning and fabricated multilingual terms - Fahd Mirza reported ModelBench/OpenBMB released MiniCPM-5 2B, a dense 2B open model that its report says averages 53.9 across a comparison suite of older 2B-4B baselines; he said he would prefer recent baselines. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: MiniCPM5 2B: SOTA or Scam? Let's Test Locally
- ISTA DASLab releases GSQ+RCO quantized Qwen 3.8 27B; lab calls 11.8 GB build lossless on tasks - Fahd Mirza reported ISTA DASLab released GSQ+RCO quantized builds of Qwen 3.8 27B in four sizes of about 8-12 GB, with the 11.8 GB IQ3_S recommended, and that the lab says performance matches the full model on benchmarks; he showed no benchmark comparison. [0 first-party, 1 hands-on, 0 relaying] Watch: Fahd Mirza: Qwen3.8-27B GSQ+RCO: 27B Params, 11.8GB, Zero Accuracy Lost Locally
- IndyDevDan's swarm tests: GLM 5.3 ten agents drew pelican for $20, DeepSeek V4 Pro showed coordination overhead - IndyDevDan reported that a 10-agent GLM 5.3 swarm in his custom Pi-agent harness on a Mac mini built a pelican-on-bicycle for about $20, 46M tokens and 873 calls in 56 minutes, using a prompt more detailed than Simon Willison's standard one. [0 first-party, 1 hands-on, 0 relaying] Watch: IndyDevDan: Are Agent Swarms USEFUL? OpenAI’s GPT-6 Astra SWARM Takeaways
- Stripe's internal agent Kai built in about 1.5 engineer-weeks, used by 86% or more of staff, builder says - A Stripe builder said on How I AI that Kai v0 took about one and a half people two weeks, followed by a 200-300 user pilot and a company-wide demo, and that 10,000+ people now use it weekly with a core team under 10. [0 first-party, 0 hands-on, 1 relaying] Watch: How I AI: The enterprise AI stack behind Stripe’s company brain “Kai”
- GitHub launches Copilot cloud agent in Slack and Microsoft Teams - GitHub demonstrated the Copilot cloud agent in Slack and Microsoft Teams on Sept. [1 first-party, 0 hands-on, 0 relaying] Watch: GitHub: GitHub Copilot in Slack and Microsoft Teams | demo | GitHub Checkout
Models & learning
- Alex Finn claims Astra low effort beats GPT-5.6 Soul high and advises dropping agent.md files - In a sponsored video, Alex Finn said computer use is Astra's biggest improvement and is available in the ChatGPT desktop app but reportedly not the CLI, narrating it adding ingredients to an Amazon cart. [0 first-party, 1 hands-on, 0 relaying] Watch: Alex Finn: 7 tips that turn ChatGPT 6 Astra into AGI (high hype)
- Goldie has Astra in Codex publish a Skool post via Chrome extension in about 90 seconds - Julian Goldie said GPT-6 Astra in Codex, using the ChatGPT Chrome extension, opened his Skool community page and published a marketing post in about 90 seconds with no fixes. [0 first-party, 1 hands-on, 0 relaying] Watch: Julian Goldie: GPT 6 Astra AI Full COURSE 1 HOUR (Build & Automate Anything) (high hype)
- Codex cloud scheduled tasks run with the app closed but cannot select model or reasoning - Julian Goldie showed a Codex cloud routine scheduled for 7:00 a.m. [0 first-party, 2 hands-on, 0 relaying] Watch: Nate Herk: I Turned GPT-6 Astra Into a 24/7 Stock Trader (tutorial)
- Remiddi has Codex with Astra build a Linux computer-control agent, sees similar results to a local 27B model - Manolo Remiddi said he gave Codex his browser agent's site and GPT-6 Astra recreated it as a floating agent controlling a whole Linux PC, mostly in one shot plus feedback. [0 first-party, 1 hands-on, 0 relaying] Watch: Manolo Remiddi: GPT-6 Astra: I Burned My Entire Weekly Allowance in One Day
- Jones describes simulated household move with Astra manager agents; Shumer's Manhattan model reported - Nate B Jones said he simulated a household move using a manager agent that interviews the user and delegates housing, school, doctor and DMV subtasks to Astra execution agents; no outputs, timings or error rates were shown, and he said the agent should not be handed a credit card. [0 first-party, 0 hands-on, 1 relaying] Watch: Nate B Jones: There Are Jobs You Could Never Give AI. I Gave GPT-6 Astra 20 Hours Of
- Riley Brown says Astra in Codex built an estate house in Blender and playtested its own map - Riley Brown said GPT-6 Astra, controlling Blender from Codex, built an estate house from a PDF in about 20 minutes and, after about an hour, added it as a playable map in his Call of Duty-style game, which he said took four prompts and a fifth for multiplayer. [0 first-party, 1 hands-on, 0 relaying] Watch: Riley Brown: I Spent 100 Hours Using GPT-6 Astra (This Feels Like AGI) (high hype)
- AutoIndex uses analysis and code agents to rewrite a corpus for retrieval; author cautions on unattended use - On Weaviate's podcast, the author of a UMass Amherst paper with a Databricks mentor described AutoIndex, in which an analysis agent diagnoses ranking failures and a code agent proposes Python diffs kept only if the target metric improves. [0 first-party, 0 hands-on, 1 relaying] Watch: Weaviate: AutoIndex with Sam O'Nuallain - Weaviate Podcast #143!