Podcast··41m

Episode 16 — Astra Beat Claude and Found Two Zero-Days

OpenAI's Astra hit 100% on the exploit benchmark and found two zero-days before launch, landing in the same 24 hours as Fable 5.1. The first time OpenAI has felt ahead of Anthropic in public.

Episode notes

First episode back from the summer break, with ten to fifteen model releases to catch up on. Astra launched with a film homage instead of a model card, scored 100% on the exploit benchmark, and landed in the same 24 hours as Fable 5.1.

Chapters

  • 02:00 — Astra's launch video, computer use, and 100% on the exploit benchmark
  • 13:00 — Astra versus Fable 5.1: which model for which job
  • 21:00 — GLM 5.3 and 5.3 Flash: near-frontier intelligence on Chinese chips
  • 31:00 — Gemini 3.8 Flash, Meta's Muse, DeepSeek V4.1 Flash, and the flash-model flood

Episode 16 — Astra Beat Claude and Found Two Zero-Days

Description: OpenAI's Astra hit 100% on the exploit benchmark and found two zero-days before launch, landing in the same 24 hours as Fable 5.1. The first time OpenAI has felt ahead of Anthropic in public.

We took the summer off, rethought the format, and came back to a backlog of ten to fifteen model releases. The headline is not that OpenAI shipped a good model. It is that for the first time in this cycle, what OpenAI released publicly felt stronger than what Anthropic released publicly, and Anthropic has held that lead for a long time.

Astra Launched With a Film Scene, Not a Model Card

Astra is the first model we can remember whose launch was a short film rather than a blog post and a model card. The video rebuilds an old movie scene: someone walks into frame in front of a screen, sits down, and talks to their computer. In the original, slowly, it draws a yellow circle. Someone from OpenAI walks into the same setting, asks for the same yellow circle, and gets it instantly, then extends it into a rocket, then moves into Blender.

  • Emotion over specification: this is car-advertising logic applied to models. You are not selling the architecture, you are selling the hours you give back. Anthropic hired an experience and community lead on a $200-300k salary to run parties and events for the same reason. Fable's own launch had an animated film attached to it.
  • The demo binds the model to a capability: because the video is about computer use, people now associate Astra with computer use specifically. That is a marketing outcome most model launches fail to achieve.
  • 100% on the exploit benchmark: Astra is the first model to score a perfect result there, and it found two zero-days before it shipped. Worth being precise: the 100% score was not the released model. Nobody should read that as an invitation.
  • No tier above it: publicly there is no model above Astra. Tighter guardrails, probably, and likely a flash variant later, but not a hidden family in the way Mephos sat above Fable.

Astra Versus Fable 5.1: Which Model for What

Fable 5.1 and Mephos 5.1 arrived on 1 September, inside the same 24 hours as Astra. Nominal pricing did not change, but cache reads got more efficient, which works out to roughly 25% cheaper in practice. A long run that should have cost $60-70 came in closer to $20-30.

  • Retrained for agentic work: 5.1 feels genuinely optimised for agentic workloads rather than rebadged. Strong on backend, strong on planning.
  • Weaker where it used to be strongest: on creative frontend and UI work it does not match Opus. The original Fable 5, in its roughly six days without hard guardrails, was more impressive than 5.1 is now.
  • Astra's UI has a house style: different YouTubers running completely different prompts produced visually similar interfaces. It reads like a design system baked into the model, the same way a Claude Code frontend is identifiable by the little card strip on the side. Whether that is deliberate differentiation or a single design skill applied to everything, it is consistent.
  • The split that works: Fable 5.1 for backend implementation and as the orchestrator holding the sub-agents together, Astra for pulling the frontend together and for actually driving the computer. Astra is the better answer when the job is operating software rather than writing code.

The strategic read is that computer use turns every human-facing website into an API. That was the awkward gap we talked about before the break: you had to bridge from a site built for a person to something an agent could call. Astra closes that gap by using the site the way a person would, only faster.

Market Share Moved Under Everyone's Feet

A chart from last year shows OpenAI's enterprise API market share roughly halving while Claude overtook it. In 2023, about half of all companies using any model used OpenAI. Google is now on the verge of the same overtake, and Meta started weak and kept declining.

The Cursor episode fits the pattern. After Altman said Cursor could no longer use OpenAI models, one of Cursor's co-founders replied that OpenAI accounted for only about 5% of total API usage in Cursor anyway. For coding specifically, OpenAI's edge is not obvious, though Codex as an application is excellent, and Codex paired with Astra's computer use is close to a cheat code.

GLM 5.3 and 5.3 Flash: Frontier Intelligence You Can Host

GLM 5.3 was open-weighted last week: weights, everything. Put it on your own GPUs and you have frontier-adjacent capability inside your own infrastructure. That matters more than a benchmark, because with Fable 5 and Astra you still have no control over what data lands on someone else's servers, and agents get more useful precisely as you give them more data to act on. Health records and confidential documents are exactly the data you least want sitting in someone's Silicon Valley region.

  • The numbers: GLM 5.3 is 753 billion parameters with 40 billion active. On some benchmarks it is on par with Fable 5, extending the claim GLM 5.2 already made at a fraction of Fable's cost. The charts come from Z.ai, so a grain of salt applies.
  • Flash is the surprise: GLM 5.3 Flash is 320 billion parameters with only 18 billion active. Technically small, and the intelligence-per-parameter is where it stops making sense.
  • Multimodal only on the small one: Flash accepts image input. The full 5.3 is text only. Counterintuitive, and probably a cost or complexity decision on the reasoning side.
  • First model run entirely on Chinese chips: Flash launched quietly as OX Alpha, free for everyone on OpenRouter, with 100 trillion tokens a day of open compute. It ran on new Xiaomi silicon, which is the first time a model of this class has been served end to end without US hardware.
  • Runnable at home: quantised down, GLM 5.3 Flash fits on two DGX Sparks, roughly $8-10k. Not cheap, but not a data center either.
  • Reasoning effort actually matters: on GLM 5.3 the jump from low to high is large, and high to max is still meaningful. On Fable the difference is massive. On 5.3 Flash, high and max are almost indistinguishable, which makes Flash the obvious implementer while the big model does the planning.

The working pattern: GLM 5.3 plans, and 10 to 15 Flash sub-agents scrape the web, scan the codebase, and implement. Of everything released over the break, this is the combination that changed daily output most.

Gemini 3.8 Flash, Meta's Muse, and DeepSeek V4.1 Flash

Gemini 3.8 Flash is excellent for sub-agents. Trust in Google models for coding is still low, mostly as inherited bias, but for orchestrated sub-tasks it performs. Google also has a Flash Cyber model aimed at finding and fixing vulnerabilities, but it is locked inside a trusted-partner program. The real interest is Gemini 4: Google researchers on X keep calling it the largest post-training run they have done, and leaked benchmarks show it outperforming nearly everything except Astra in several categories. Unverified, and possibly fake.

Meta's Muse shipped as an agentic personal assistant app you can text through WhatsApp, US rollout only. It looks like an agent for people who do not want to configure an agent, which is a reasonable market. Muse Spark is the more interesting release: on price per intelligence it is well ahead of Fable or Astra, which means far more experiments inside the same budget. Whether that price survives market-share capture is the open question, and DeepSeek already showed what happens next by making its promotional pricing permanent.

DeepSeek V4.1 Flash landed on 12 September with a completely new architecture, which is why it is smarter and cheaper than V4. It is roughly double the size of GLM Flash, so it should cost more, and V4 Pro sits near two trillion parameters while GLM 5.3 Flash runs 18 billion active and scores above it. That comparison is suspicious enough to want independent numbers.

The Flash Flood, and the One Lab Missing From It

Every lab now has a flash model, and Anthropic does not. Given the pattern, they are next.

There is a marketing read: attach a frontier model's name to a small model and people transfer the intelligence association to something with under a tenth of the parameters. But the stronger read is demand. Agentic workloads need cheap, fast models, and on-premise or home-lab deployments will never have the VRAM for Fable 5.1 or full GLM 5.3. GLM 5.3 Flash is roughly where frontier capability sat six months ago, at a size you can quantise onto two Sparks.

For scale on the other end: Siri handles something like three billion requests a year and the new version, likely Gemini 3.8 Flash underneath, still will not ship in Europe. Google serves between 8.5 and 14 billion searches a day at roughly 90% market share. Everyone on the planet accounts for at least one, bots included.

Conclusion

The gap between frontier and good-enough closed sharply over the summer, and it closed from the cheap end. Astra proved that a launch can sell a capability rather than a spec sheet, and that computer use is the feature that turns the rest of the internet into infrastructure an agent can use. But the models that changed how we work are the small ones. When 18 billion active parameters on Chinese silicon do the implementation while a frontier model does the thinking, the question stops being which lab is ahead and becomes which combination gets the work done.

Listen on Spotify