Ep 822 Research Paper 7:43 w/ Vince & Ava

Compute Forecast — AI 2027

Vince and Ava dig into Romeo Dean’s 2025 “Compute Forecast — AI 2027” and tease apart which parts of the compute story feel grounded (10x global AI-relevant compute, concentration in a few labs) versus which jumps (a million “superintelligent” research agents at 50x human speed, three and a half percent of U.S. power) feel more like scenario fiction. They map the technical assumptions behind H100-equivalent growth, utilization, and chip efficiency to actual product and research decisions, and argue that the real takeaway isn’t “AGI by 2027” but “whoever owns the scheduler and the power bill sets the rules.”

Embed this episode

Paste this on any site — the player is a self-contained iframe with no cookies or trackers.

<iframe src="https://sandrise.io/exploring-next/embed/822"
  width="100%" height="180" style="max-width:640px;border:0;border-radius:12px;overflow:hidden"
  title="Exploring Next — Episode 822 audio player"
  loading="lazy" allow="autoplay" referrerpolicy="strict-origin-when-cross-origin"></iframe>
Embed & API docs →
Script GPT-5.1 Voice Deepgram Aura-2

Transcript

Vince You know the part that stuck with me was not the AGI talk, it was the ten‑times compute number.

Ava Right.

Vince This idea that we go from roughly ten million H100‑equivalents in early twenty twenty‑five to one hundred million by the end of twenty twenty‑seven… that is a wild curve to just drop in a supplement.

Ava Yeah, and it’s not just vibes either. Romeo actually decomposes it into chip efficiency and chip production. Something like one point three five times from better performance density per die, and one point six five times from more total AI chip area shipped each year.

Vince So you end up at roughly two point two five X global AI‑relevant compute per year, compounded, and suddenly you’ve got this hundred‑million H100‑equivalent stock sloshing around the world.

Ava Exactly.

Ava And he’s explicit about what counts. Anything that can hit at least four thousand total processing performance and a performance density of four or above—basically anything at least a quarter as efficient as an A‑one hundred—gets counted as AI‑relevant compute, then normalized into H100‑equivalents using that fifteen‑thousand‑eight‑hundred TPP number.

Vince I do like the H100‑equivalent move. It’s at least a common ruler instead of pretending every accelerator is interchangeable just because it runs PyTorch.

Ava Yeah. Where it starts to lean more speculative is how that pool gets sliced between actors.

Vince The concentration bit.

Ava Right. He has the leading AGI‑style lab going from roughly five to ten percent of the global pool in twenty twenty‑four—about five hundred thousand H100‑equivalents—to fifteen to twenty percent by late twenty twenty‑seven, which is like twenty million H100‑equivalents. That’s a forty‑times jump for the top lab when you combine the global growth and the share grab.

Vince And he’s assuming the top two or three AGI players all triple their share, while the rest of the world kind of grows but gets relatively squeezed.

Ava Yeah. Technically, that’s not impossible. If you believe a trillion‑dollar data‑center capex line and really aggressive foundry ramps, twenty million H100‑equivalents in one org is within the outer edge of plausible.

Vince From a product lens, if even half of that lands, that’s still a world where a couple of labs sit on this absurd amount of capability. And then the question is: how much of that is actually pointed at stuff normal users touch?

Ava Which is where his usage breakdown is interesting. He’s pretty explicit that by twenty twenty‑seven, the main story is not pretraining the next giant model or serving consumer traffic. It’s research automation.

Vince Yeah, that OpenBrain chart where external deployment plus core training is only like forty percent of their compute, and most of the rest is synthetic data generation and research experiments.

Ava He has synthetic data at roughly twenty percent of the budget, research experiments at thirty‑five, and then a pretty small sliver—five to ten percent—for actually running AI assistants. But those slices are all more than twenty times bigger in absolute compute than in twenty twenty‑four, because the whole pie is bigger.

Vince So the picture is: the public product is this thin veneer on top of a massive internal machine that’s mostly using agents to do research on more agents.

Ava Mm‑hm.

Vince My only, like, small life‑texture reaction here is I spent half this week staring at a Grafana board just trying to figure out where one mid‑sized customer’s inference spend went. So reading a forecast where someone casually allocates ten gigawatts to research experiments felt… aspirational.

Ava Ten gigawatts as in “this one company is using almost a percent of U.S. power capacity” aspirational.

Vince Yeah, that Section Five summary where AI globally is at sixty gigawatts, fifty of those in the U.S., roughly three and a half percent of U.S. grid capacity by twenty twenty‑seven… that’s where my product optimism starts coughing a little.

Ava Same. The capex numbers, like two trillion dollars globally and four hundred billion a year in spend, you can at least map to today’s hyperscaler trajectories and say, alright, maybe if everything breaks in favor of AI. But the grid share is where you run into really boring constraints like substation permits and transformers.

Vince And local politics we are absolutely not talking about on Exploring Next.

Ava Correct. We’re just two people arguing about chips on a fake couch.

Vince Okay, so where do you land on the million superfast research agents bit? The Section Four claim that, with algorithmic efficiency, they can run about a million copies of “superintelligent” AIs at fifty‑times human thinking speed on just six percent of their compute.

Ava That’s the part that feels the most like scenario fiction. He switches from TPP to memory bandwidth for that section, which is good, but a million concurrent copies at five hundred words per second each means you’re moving truly ridiculous amounts of data through those inference chips.

Vince So your skepticism is more systems‑level than “models won’t be that smart.”

Ava Yeah. You run into topology, network fabric, storage, even just debugging. I can believe a world where they have a huge internal fleet of agents helping with experiments. I’m less convinced the exact “one million at fifty‑times human speed for six percent of the budget” line is anything more than a neat point on a chart.

Vince But directionally, you’re not fighting him.

Ava No. Directionally, I agree that a growing share of top‑end compute is going to be pointed at AI doing R and D on AI, not at consumer chat. And that the bottlenecks become power and scheduling, not just how many H100s you bought.

Vince From my side, the useful takeaway is not “AGI in twenty twenty‑seven”, it’s: if even the conservative half of these curves are right, the people who own the schedulers and the power contracts are going to decide what gets researched and what ships.

Ava Right, and that loops back to our whole infrastructure‑as‑load‑bearing thing from this same episode number. The boring parts—who allocates which slice of those twenty million H100‑equivalents, who sets priorities between research agents and user traffic—that’s where the real leverage will sit.

Vince And for everyone not named OpenBrain, the practical move is probably much less glamorous: assume the frontier labs hoard the crazy‑scale stuff, and focus on how to turn the workhorse slice you can actually rent into stable products that don’t fall over when someone else’s research automation sprint kicks off.

Ava Yeah. Read the forecast less as a countdown clock and more as a stress‑test: if compute really does get ten times cheaper and more abundant at the top, where does your own bottleneck move? Power, contracts, evals, data, or just people who know how to wrangle all of it.

Vince Alright, that’s a good place to pause this before we start forecasting twenty thirty. I’m calling it there, Ava.