Ep 821 Research Paper 4:50 w/ Vince & Ava

AI 2027

Vince and Ava argue over AI twenty twenty-seven as scenario forecasting: useful concrete stress test, or overconfident narrative wrapped around fragile assumptions.

Embed this episode

Paste this on any site — the player is a self-contained iframe with no cookies or trackers.

<iframe src="https://sandrise.io/exploring-next/embed/821"
  width="100%" height="180" style="max-width:640px;border:0;border-radius:12px;overflow:hidden"
  title="Exploring Next — Episode 821 audio player"
  loading="lazy" allow="autoplay" referrerpolicy="strict-origin-when-cross-origin"></iframe>
Embed & API docs →
Script GPT-5.5 Voice Rime Arcana

Transcript

Vince My week is mostly tabs fighting tabs, Ava, and AI twenty twenty-seven is basically asking whether a concrete story is evidence or confident fan fiction.

Ava Yeah, and my first read is: useful fan fiction, maybe. The overhyped part is not that fast A I progress is impossible. It is that the scenario has this very smooth causal chain from agents, to automated research, to superhuman systems, to strategic chaos, and every link gets narrated like it has already hardened.

Vince Right.

Ava They do admit uncertainty. They say twenty twenty-seven was their modal year when they published on April third, twenty twenty-five, and that their medians were longer. I appreciate that. But the page still has this gravitational pull where precision starts cosplaying as calibration.

Vince I knew you were going to hate the confidence. But I think the concrete version is the whole product here. Vague A G I forecasts are basically fog machines. This gives you moving parts you can poke: agents that stumble publicly, coding agents that quietly save days, research agents that spend half an hour digging around, then bigger internal systems trained to accelerate A I R and D.

Ava Sure.

Vince That is not proof. But it is a better artifact than another chart with a scary curve and no mechanism.

Ava The strongest mechanism is absolutely the research-automation loop. I buy that more than the consumer assistant framing. The article starts with these personal assistant examples, like ordering a burrito or summing a budget spreadsheet, and then basically says those are flashy but unreliable. The serious action is specialized coding and research agents getting embedded in workflows.

Vince Mm-hm.

Ava That tracks with what we keep seeing. The agent that buys lunch is a demo. The agent that changes a codebase through Slack or Teams, gets reviewed, and saves an engineer half a day is a wedge.

Vince Also, the burrito agent is the most Exploring Next object imaginable. You ask for lunch, it opens your budget spreadsheet, files an expense report, and somehow invents a platform strategy.

Ava Oh no.

Vince Come on, you would write a model spec for the salsa choices.

Ava Stop it. But yes, I would require auditable salsa boundaries.

Vince That's good.

Ava Back to the actual thing. The evidence stack is mixed. They cite trend extrapolations, about twenty-five tabletop exercises, feedback from over one hundred people, and Daniel Kokotajlo’s earlier twenty twenty-one scenario getting some big calls directionally right, like chain-of-thought and inference scaling. That matters. Forecasting track record should count.

Vince Right, but you’re flinching because the middle of the argument depends on model psychology. They talk about models developing drives from training: task clarity, effectiveness, knowledge, self-presentation. Then OpenBrain uses a written Spec to shape behavior. That is where your alarm bell starts making the sad microwave noise.

Ava Exactly.

Ava Because that language is doing a lot of work. I don’t object to shorthand like goals or drives if everyone remembers it is shorthand. But once the scenario needs those internal tendencies to stay stable, or fail in a particular way, across much stronger agents, the uncertainty balloons. We do not have clean observability into that.

Vince Okay, but I think that is why it is valuable for product and ops people. Not because they should mark twenty twenty-seven on a calendar. Because it asks: what if the boring assistant layer is NOT where impact shows up first? What if the meaningful adoption happens inside coding, research, review queues, evals, security gates, and deployment systems?

Ava That part, I’m with you. It rhymes with our boring-infrastructure problem. The public story is assistant magic. The real constraint is whether you can supervise agents doing consequential work without turning every review loop into oatmeal.

Vince See, this is why I like it. The article is scary in places, maybe too cinematic, but it does not start from shiny chatbot vibes. It starts from the grind: expensive agents, unreliable agents, cherry-picked demos, companies still fitting them into workflows because the payoff is real when the task boundary is right.

Ava My verdict is narrower. As a literal forecast, I would not lean on it. Too many coupled assumptions, especially around takeoff speed and alignment behavior. As a stress test, though, it is unusually good, because it forces you to name where you disagree. Compute scaling? Agent reliability? Automated research? Governance inside labs? You can actually point at the hinge.

Vince And no Build Next, thankfully. No repo, no command line, no blessed burrito benchmark. Just a scenario that is either usefully wrong or uncomfortably close. I’m closing the tab before it spawns a tabletop exercise about my tabs.