Podcast··42m

Episode 15 — China Broke the Servers and Mira Murati's First Model

Kimi K3 hit frontier intelligence so cheaply that Moonshot had to stop new sign-ups, while Mira Murati's Thinking Machines shipped Inkling, a natively multimodal model you fine-tune by talking to it.

Episode notes

More than ten models shipped in two weeks, Kimi K3 got so much demand that Moonshot closed new sign-ups entirely, Mira Murati's Thinking Machines released its first open-weight model, and OpenAI's German data center talks quietly collapsed over energy.

Chapters

  • 03:00 — Ten-plus model releases in two weeks and Kimi K3 hitting its own capacity ceiling
  • 13:00 — Moonshot's jump from $2B to $22B and the White House testing framework
  • 22:00 — Mira Murati's Inkling: natively multimodal, fine-tuned by conversation
  • 32:00 — Anthropic's path to profit, xAI's $1B monthly burn, and Europe's energy problem

Episode 15 — China Broke the Servers and Mira Murati's First Model

Description: Kimi K3 hit frontier intelligence so cheaply that Moonshot had to stop new sign-ups, while Mira Murati's Thinking Machines shipped Inkling, a natively multimodal model you fine-tune by talking to it.

Three of us, three countries, one online recording. The two weeks behind this episode produced more than ten model releases, and the pattern in them is hard to miss: the open-weight labs are shipping faster than their own infrastructure can absorb, and the money keeping the closed labs alive is starting to look very different depending on which lab you look at.

Ten-Plus Model Releases in Two Weeks

The release cadence has stopped being trackable as news and started being a job. Scrolling the Hugging Face model list several times a day is now a legitimate way to stay current, which says more about the market than any benchmark does.

  • Fine-tuning comes to AMD: Unsloth now supports AMD hardware, which means the Strix Halo chip with 128 GB can post-train models locally. Not retraining, but post-training, and that puts a fine-tuned Qwen within reach of a desk instead of a data center.
  • Hardware is moving the wrong way: the RTX Pro 6000 sat at $9,999 two months ago. It is now $12,500. That is a better return than most stock portfolios this year, and it is a tax on anyone trying to build local capacity.
  • Qwen 3.8 landed, partly: the max version launched this week at 2.8 trillion parameters, with open weights not yet released. Alibaba has already committed to open-sourcing the smaller variants, 27B and 35B in the mould of Qwen 3.6, which is the part that actually matters for anyone running models at home.

Kimi K3 Broke Moonshot's Own Servers

Kimi K3 is being used so heavily, by Chinese and US companies alike, that Moonshot's inference capacity ran out. Rate limits dropped to almost nothing and new subscriptions were closed entirely.

  • The price gap is the whole story: Kimi K3 is reportedly close to Fable 5 on intelligence at roughly $3 per million input tokens and $15 per million output. Fable is $10 in and $50 out. Same ballpark of capability, a fifth of the cost.
  • Restriction will not hold: Washington advisers are reportedly split on whether to restrict US companies to US models. Even if a rule lands, individuals will route around it. Companies below the top five will take the risk; the top five will not want the fine.
  • Europe is not in this conversation: Germany shipped its first national model a week or two ago. It does not benchmark against Qwen 3.6, let alone the frontier. The first step exists, which is the most that can be said.

Moonshot at $22B and the Valuation Gap

Moonshot was worth $2 billion in May. It is worth $22 billion now, an elevenfold move driven almost entirely by the K series.

That is still small measured against US labs, and the reason is geography rather than output. A US lab shipping the same capability would be valued an order of magnitude higher. The counter-argument is that Moonshot is closer to a model foundry: it produces models, while Anthropic also ships a mature desktop app, Claude Code, and adjacent research. Every major lab now has its own CLI and its own agent-management IDE, and most of them are clones of each other.

The White House Testing Framework

The framework between the administration and the frontier labs went from an idea to near-formalised in about two weeks. New frontier models now pass through a testing and sign-off process before release, explicitly so that the Fable and Mephos situation does not repeat.

It only binds US labs. The administration can control how Chinese models are used inside the US, but not how they are built. That asymmetry is the same one that made export controls awkward in the first place.

Everyone Is Building Their Own Silicon

Google released its own inference chip internally, Frozen V2, built to cut inference cost at scale in its own data centers. The lineage is interesting: the Groq acquisition, the one spelled with a Q, brought in an architecture designed around raw speed, and Frozen V2 looks like it shares that DNA. A chip takes at least six months to surface after an acquisition, so the timing fits.

OpenAI's equivalent is Jalapeño, which went from prototype to deployed in its data centers in under two months. For silicon running trillion-parameter models, that timeline is close to implausible. No benchmarks have been published for either chip yet.

Google's Missing Model

The next Gemini is roughly a month behind schedule. Reuters reported that internal coding-performance goals were not met on the original timeline, so Google is retraining and retuning.

There is a rational case for the delay. Fable 5 is the north star right now, and several new models are comparable to it. Releasing a model that lands below that line invites another round of public criticism, which the previous family already absorbed. The one thing Google has consistently won on is the context layer, the million-token window, and now two million. The models themselves have not kept pace, though the Gemma family aiming at small, scalable, on-device inference is a genuinely smart lane.

Mira Murati's Inkling: Multimodal from the Ground Up

Thinking Machines, founded by former OpenAI CTO Mira Murati and roughly ten to twelve early OpenAI people, raised $2 billion in the largest pre-seed round ever recorded and carries a $12 billion valuation. On 14 July it shipped its first open-weight model, Inkling.

  • Natively multimodal, not bolted together: almost every other lab attaches a vision transformer to a trained text model, which translates images into something the model can read. Inkling was trained from scratch on a run in the tens of trillions of tokens, with a large share of that audio, video, and images. There is no decoder in front of it doing the interpreting.
  • Good, not best, by design: the team was explicit that beating Fable 5 was not the goal. They set out to build a good model, and the benchmarks put it below the frontier.
  • The Tinker API is the real product: you fine-tune Inkling by talking to it. The launch demo told the model to stop using the letters E and A in its responses; the Tinker API post-trained on their hardware, returned new weights, and roughly 29 minutes later the chat was running the updated version. Fine-tuning at that scale is expensive in inference, and the pricing is not clear yet, but a conversational path from "behave differently" to deployed weights is new.

GPT-5.6: Sol, Luna, Terra

OpenAI shipped the 5.6 family, with Sol as the top model. Pricing per million tokens: Sol at $5 in and $30 out, Terra at $2.50 in and $15 out, Luna at $1 in and $6 out. Some people put Sol on par with Fable.

The more interesting property is token efficiency. GLM gets its intelligence by thinking out loud at length and arriving at the right answer through volume. OpenAI went the other direction: the thinking trace on the 5.6 models is short enough that you reach the result noticeably faster. Per-token pricing stops being a useful comparison when two models spend wildly different numbers of tokens on the same task, which is an argument for a new cost scale rather than a new benchmark.

The Money: Anthropic's Path to Profit, xAI's Burn

Anthropic looks like the only frontier lab on a credible path to profitability through inference, possibly within one or two years. The projection comes from Nate Ratner's numbers and covers Q2, so it is still forecast rather than filing. It also depends entirely on how GPUs are booked: counted against Anthropic the picture changes, counted against the partners investing in Anthropic it changes again. The question is not whether the accounting is flattering but by how much.

xAI burned roughly $1 billion per month as of May 2026, before the merge. In a market where every number has an extra zero, a billion a month is still an enormous figure.

And on the supply side, Apple has claimed something like 40% of global memory capacity through 2027 and 2028. That is why a 256 GB DDR5 kit now runs $12,000 and why home data center builds have become genuinely painful to fund.

Why OpenAI's German Data Center Talks Fell Through

OpenAI proposed making Germany a European flagship, with German government investment attached. Germany said no.

  • The location was the problem: the proposed site sat very close to the French border, for cheap and non-fluctuating nuclear power. Stable baseload is exactly what a training cluster wants, and it is exactly what the German grid does not offer.
  • Bureaucracy and price: approval timelines are too long and energy is too expensive. Companies that started European builds and pulled out have said publicly they will not come back.
  • Why Europe was on the list at all: the US is running out of grid. Roughly half of large US data center projects were cancelled or delayed because they could not be connected, which is the physical ceiling underneath the $2.5 trillion investment wave. Musk personally bought a turbine and generator company, which reads less like diversification and more like buying the only available capacity.
  • The nuclear question: a standard reactor runs about $10 billion and takes five to ten years. Against AI capital expenditure that is not expensive. Germany decommissioned its plants roughly five years ago, and the same country is now short exactly the power those plants produced. Meanwhile China added in a single year what Germany produces annually from solar.

Conclusion

The cheap models are winning on volume faster than anyone's infrastructure can serve them, and the expensive models are being defended with accounting rather than capability. Thinking Machines is the outlier worth watching, not because Inkling beats anything but because fine-tuning by conversation removes the last technical barrier between a general model and a specific one. The constraint that decides the next two years is not intelligence. It is electricity, and Europe has not started solving it.

Listen on Spotify