Deepseek Model Cheap AI Price War
Jessica and Cathy dig into DeepSeek V4 Flash, the AI price war, and whether “intelligence as a commodity” really holds up once you look at routing, safety, and who actually pays the bills.
Transcript
Jessica So DeepSeek drops this V4 Flash thing and suddenly everybody’s yelling that intelligence is a commodity now.
Cathy Yeah, that Axios piece basically says, congrats, you poured hundreds of billions into compute and now the brains on top cost pocket change.
Jessica I mean… when the article says V4 Flash is basically Opus four point eight level coding for a ninety‑nine percent discount, that’s not subtle.
Cathy Right.
Cathy On Arena’s front‑end leaderboard it even edges Opus four point eight, and then the author does that electricity and gasoline analogy — you stop caring which plant made it, you just want cheap kilowatt hours.
Jessica Before we get fully existential about kilowatt hours, how’s your week? You looked like you were doomscrolling benchmarks when I pinged you.
Cathy I was, but in my defense this time the charts were actually interesting. Price axes are finally moving, not just accuracy bars.
Jessica Okay, that is such a you sentence. I’ve mostly been in product decks where everyone’s asking, can we switch to the cheap one yet without breaking everything.
Cathy That’s literally the question this article is poking at.
Jessica Yeah. Central claim is basically: top‑tier models are close enough that for a huge chunk of use cases, you swap on price and call it a day.
Cathy And the evidence stack is decent. DeepSeek claims near‑Opus four point eight on complex coding and autonomous software benchmarks, the Arena leaderboard backs that for front‑end tasks, and then they drop the number: twenty‑eight cents versus twenty‑five bucks for the same output volume.
Jessica That number is wild.
Cathy It is. Then they zoom out: July was just everyone blinking. OpenAI cuts GPT five point six Luna by eighty percent three weeks after launch, Google ships a whole line of Gemini Flash models that are explicitly about efficiency, Grok four point five comes out at Luna’s original price, and Meta suddenly does a closed Muse Spark that’s aggressively priced for devs.
Jessica Plus Anthropic as the holdout, keeping Opus premium and basically saying, we’re the fancy gas station with the good additives.
Cathy Exactly. Their bet is that safety and precision are worth a margin even if the raw code‑completion score is similar.
Jessica So where does this commodity argument hold for you technically, and where does it feel like Axios is smoothing over cliffs?
Cathy It holds in the sense that marginal capability gains on generic coding are flattening. If V4 Flash, Luna, Gemini Flash, and Sonnet all clear your test suite within a couple percent, you absolutely should look at dollars per bug‑free pull request.
Jessica Mm‑hm.
Cathy But “commodity” is doing a lot of work. These models still have very different failure modes, tool ecosystems, and governance stories. Opus four point eight’s whole thing is lower misaligned behavior on stuff like deception; that doesn’t show up in a front‑end leaderboard but it matters if the agent is touching payments or prod infra.
Jessica Yeah, like, the article leans on Zack Kass’s line about diminishing model returns — at some point the next model doesn’t matter to you — and that’s true for, say, spam triage or basic code gen. It is not true if you’re running an autonomous workflow that can nuke real systems.
Cathy Or if regulators are asking you which lab’s safety processes you trust. That’s not electricity, that’s more like picking a cloud region with compliance guarantees.
Jessica The part that did feel very on‑theme for us was the “intelligent routers” bit. Vinesh from Qualcomm talking about systems that auto‑pick the best model on capability, speed, and price — that’s our whole cost‑per‑outcome routing rant from months ago.
Cathy Yeah, that’s basically Cursor’s classifier story scaled out. You hide the menu behind a simple toggle and let the infra decide which model actually runs, which makes brand loyalty even weaker.
Jessica And that’s where I think the piece underplays something. If routing gets good, the model vendors don’t just lose pricing power, they also become interchangeable modules in somebody else’s product. The router owner becomes the gatekeeper.
Cathy Right. The existential threat for frontier labs isn’t just cheap Chinese models, it’s being abstracted behind an API that can swap you out overnight.
Jessica Although they do give OpenAI’s counterpoint some airtime — Sam basically saying, we’ll make it up on volume, we don’t need insane margins if usage explodes.
Cathy Which is coherent as a story. If intelligence per dollar keeps improving and demand is elastic, you can lower price and still grow total revenue. The question is whether that demand expansion is enough to pay for the next ten‑billion‑dollar training run when DeepSeek is selling near‑Opus performance for pocket change.
Jessica So who should actually care here? Like, if I’m a team shipping a coding agent this month, does this article change anything or is it just vibes?
Cathy If you’re running serious volume, it’s a direct budget question. You should at least A slash B test something like V4 Flash or Luna against your current default, and you should assume your CFO is going to forward you this kind of article and ask why you’re not on the cheap one.
Jessica Yeah. From the product side, it’s also permission to design for swap‑ability. Build that routing layer in from day one, assume the menu keeps changing, and don’t hard‑wire a single lab into your architecture just because it feels safer emotionally.
Cathy And if you’re Anthropic‑pilled, you now have to articulate why you’re paying the premium. Not just “they’re safer,” but, here’s the concrete risk profile or governance feature we need that the cheaper model doesn’t have yet.
Jessica I like that the bottom line in the piece is basically, abundance is coming, profitability is the open question. That feels… very load‑bearing infrastructure of them.
Cathy Of course you’d say that. The room definitely got smaller on pricing, though.
Jessica You’re just happy your “room got smaller” line from last time aged well in, like, forty‑eight hours.
Cathy I am a little smug about that, yeah.
Jessica Alright, let’s stop before you start quoting your own metaphors back at me. This was fun, Cathy.