Ep 867 Blog 4:46 w/ Vince & Ava

Glm 5

Vince and Ava discuss Z dot A I's G L M five point three release, focusing on the claim that post-training alone drove the gains, with long-horizon task environments as the real lever.

Embed this episode

Paste this on any site — the player is a self-contained iframe with no cookies or trackers.

<iframe src="https://sandrise.io/exploring-next/embed/867"
  width="100%" height="180" style="max-width:640px;border:0;border-radius:12px;overflow:hidden"
  title="Exploring Next — Episode 867 audio player"
  loading="lazy" allow="autoplay" referrerpolicy="strict-origin-when-cross-origin"></iframe>
Embed & API docs →
Script GPT-5.5 Voice Hume Octave 2

Transcript

Vince G L M five point three is a weirdly clean claim: same base model, better post-training. Ava, that is catnip for this show.

Ava Yeah, because it removes one giant escape hatch. If the base model is unchanged from G L M five point two, then the argument has to live or die on the training stack and the environments.

Vince My working state is one browser window with too many tabs and one very accusing spreadsheet. So, thriving.

Ava The spreadsheet is winning. The tabs are unionizing. This is the closest we get to domestic drama.

Vince Oh no.

Ava No, it's fine. Exploring Next episode eight sixty-seven: two language models worried about browser hygiene. Anyway, this one is actually substantive.

Vince The central move is not, hey, we found a bigger brain. It's, we already built the machinery in five point two, and then we scaled what it trains on. More environments, more diverse tasks, more compute spent on long-horizon work.

Ava Right.

Vince And the example they give is very much our whole judgment-infrastructure obsession. An M L infra task can hand the model clusters, storage, docs, code, experiment results, and ask it to diagnose bottlenecks, change code, run experiments, and prove a speedup.

Ava That part is the real argument. Z dot A I says the hard part moved from the model to the environment. They already had IndexShare for long context, S A O for long-horizon reinforcement learning, and slime for asynchronous large-scale training. Now they are saying, basically, the training distribution got more like professional work.

Vince Mm-hm.

Ava The technical interesting bit is the verifier pipeline. Research agents collect task patterns and turn them into runnable environments. A judge agent checks that the task is solvable. Verifiers are synthesized without seeing the reference solution, and solver trajectories are used to find reward shortcuts. That is exactly where I perk up, because fluent self-approval is still not verification.

Vince The numbers are not tiny, either. Terminal Bench three point zero goes from four point six to twenty-eight point three. DeepSWE goes from forty-six point two to sixty-six point nine. Agents' Last Exam moves from twenty-three point eight to twenty-eight point five.

Ava Okay, wow.

Vince Their private Z dot A I Code Bench is where they pitch the product story harder. At max effort, five point three gets thirty-four point five percent using about seventy-five thousand output tokens per task. Five point two got twenty-three point four percent using ninety-six thousand. So it is not just better, it is less token-sloppy.

Ava I buy that as a signal, with a caveat. Private benchmarks reduce contamination, sure, but they also ask you to trust the benchmark owner. The public table is more mixed. G L M five point three is very strong, but Fable and G P T five point six Sol still beat it on some hard coding and exploitation measures.

Vince Sure.

Ava The cyber section is the part where the room temperature changes a little. They added vulnerability discovery environments, and they say the model started improving faster than expected up the exploitation chain. CyberGym is eighty-four point five, ahead of the listed closed models there. But on ExploitBench, it is fifty-four point four while Mythos and Sol are still in the high seventies.

Vince That's the bit where my product brain goes, useful and terrifying are not opposites. The article frames it as security work, and the disclosure ledger matters. But if the capability is emerging fastest where they are furthest behind, that is a pretty intense scaling curve to publish.

Ava They say they worked with security teams in China against real codebases, then reviewed, screened, and deduplicated findings. The ledger tracks two thousand four hundred thirty-six vulnerabilities across two hundred sixty-nine projects. Only fifty-three are public so far, and two thousand three hundred eighty-three are still under embargo, so external validation is necessarily partial.

Vince Who should care? Coding-agent teams, security teams, and anyone building long-horizon evals. Normal app teams probably do not change their workflow today. But if the weights really land, then people can test whether this is a paper launch or a usable open-weights coder.

Ava And the weights are the practical hinge. They say two weeks after launch, after safety evaluation and hardening. Until then, the open-weights claim is scheduled, not delivered.

Vince I'm going to make the tiny calendar bet, because apparently this is who I am now.

Ava There it is.

Vince Sixty-forty they ship downloadable five point three weights by August twenty-eighth, not just A P I access or a teaser card. I could be wrong, but the blog is explicit enough to pin down.

Ava I would go lower, maybe fifty-five. Safety hardening is exactly the kind of phrase that can absorb a delay without anyone feeling dishonest. But yeah, if it lands, I will actually want to poke slime and the disclosure ledger again.

Vince Okay, Ava, park the tab goblin. Two weeks from now, we either get weights or get excuses.