Ep 909 Research Paper 6:52 w/ Pippa & Tyler

Enabling independent research on how people use Claude

Pippa and Tyler talk through Anthropic's pilot letting outside researchers study aggregate Claude usage through Anthropic Insights. They focus on the article's real argument: if labs hold the only large-scale view of real AI use, independent research gets distorted by lab-shaped questions or public datasets that miss serious use. The conversation lands on the technical and practical tension in the post: privacy-preserving access is valuable, but slow, sensitive to question wording, and hard to scale without better tooling and pre-testing.

Embed this episode

Paste this on any site — the player is a self-contained iframe with no cookies or trackers.

<iframe src="https://sandrise.io/exploring-next/embed/909"
  width="100%" height="180" style="max-width:640px;border:0;border-radius:12px;overflow:hidden"
  title="Exploring Next — Episode 909 audio player"
  loading="lazy" allow="autoplay" referrerpolicy="strict-origin-when-cross-origin"></iframe>
Embed & API docs →
Script GPT-5.4 mini Voice Inworld TTS 2

Transcript

Pippa Okay, this one is weirdly important. Anthropic basically let outside researchers poke at real Claude usage without handing over raw chats, and that feels like the kind of boring breakthrough that actually changes what gets studied.

Tyler Yeah. Because the alternative is either the lab publishes the question it wanted answered, or everybody squints at public datasets that are mostly casual stuff and then pretends that tells you how people really use AI.

Pippa Exactly. And the article's argument is pretty clean: if you want independent research on real usage, you need access to real usage. Not a press-release version of it, not a toy dataset version of it.

Tyler Mm-hm.

Pippa Also, how are we nine-oh-nine episodes in and the big story is still, 'who gets the data'? That is such an Exploring Next take.

Tyler It really is. We keep ending up at the same unglamorous layer, which is usually where the actual product is hiding.

Pippa The pilot had three outside groups: Stanford's SALT Lab, Oxford's Human Information Processing Lab, and METR. They each designed their own questions, Anthropic ran the collection through Anthropic Insights, and they analyzed roughly two hundred and fifty thousand Claude conversations from April and May.

Tyler Right, and the important part is the independence. Anthropic says they only kept review rights for privacy, policy-violation risk, confidential info, and research accuracy. That's a lot narrower than the usual 'trust us, we looked at it' setup.

Pippa SALT's results are the part that grabbed me. Over half the conversations involved consequential work, which is a lot higher than the old story that people only bring low-accountability tasks to AI. They also found people usually still direct the work, then adapt Claude's output instead of copying it.

Tyler That fits the mechanism better than the hype does. People aren't just handing over agency and walking away. They're using Claude like a fast collaborator, and the friction is doing some work too.

Pippa Yeah, the article's almost annoyingly sensible there. The friction isn't just failure. It can force people to clarify the ask, catch ambiguity, and stay engaged long enough to get a better result.

Tyler Sure.

Pippa And then the Human Information Processing Lab adds the softer layer. Claude being warm lined up with people being more positive, refusing lined up with pushback, eccentric output lined up with more intellectual engagement. That's not earth-shattering, but it's a real pattern instead of vibes soup.

Tyler Mm-hm. And the comparison to everyday browsing matters more than it sounds. If the emotional patterns in Claude chats resemble how people move through the rest of the web, then this isn't some totally alien interaction mode. It's another digital attention surface with its own little moods.

Pippa I think that's the practical wedge. If you're building products on top of these systems, you care whether users feel challenged, satisfied, or weirded out, because that changes retention and trust fast.

Tyler Yeah, though I'd still be careful not to overread the correlation. Warm model behavior and positive user feeling moving together is useful, but it doesn't prove causality in any strong sense. It just gives you a direction to test.

Pippa Fair. METR is the more mechanistic piece anyway. They were estimating real-world productivity gains from coding agents, and the early read is that newer models seem to save more time than older ones.

Tyler And the neat bit is that Claude's own estimate of how long a task would take lined up reasonably well with actual developer completion times in a prior study. That's not perfect, but it's better than benchmark theater.

Pippa Right, and that matters because if the model can roughly estimate task duration, you can start using it as a measurement instrument instead of just a tool. Which is a little unsettling, honestly.

Tyler That's the part where the machine starts grading the thing it's inside of. Very elegant, very mildly cursed.

Pippa The part where I start leaning in is the privacy-preserving setup itself. Researchers never saw raw conversations, only aggregated outputs after the same legal and privacy review Anthropic uses internally. That's a real constraint, but it's also the only reason this can exist at all.

Tyler Sure, but the article doesn't dodge the cost. It says the pilot was slow and resource intensive, and that's the hard truth if you want to scale this beyond a one-off experiment. Privacy and independence are fine, but the workflow has to stop being molasses.

Pippa And the method fragility is the technical crux. Because Claude is answering the researchers' questions for every conversation, the exact wording matters a lot. If the question is sloppy, you can get categories that misrepresent the underlying chats, and nobody can inspect the raw text to catch it.

Tyler That's the part I trust least, technically. They tried to handle it by having partners test on WildChat first, but WildChat skews casual and creative, so a question that looks good there can still break on actual Claude traffic.

Pippa Mm, but that's not a reason to dismiss it. It's a reason to build the question-design layer better. The article even says they're exploring how external researchers can develop categories more effectively before they ever touch the private data.

Tyler Yeah, and that feels like the real product move here. Not 'we solved AI transparency forever,' obviously. More like, here's a workable interface between private platform data and outside research, and now the tooling around that interface has to mature.

Pippa Also, tiny thing, but I love that they say the overlap with internal work was useful instead of pretending independence means total separation. METR's proposal being similar to Anthropic's own economics work is actually a feature, not a bug.

Tyler Exactly. Multiple groups asking adjacent questions on the same data is how you find out whether a conclusion is sturdy or just the lab's favorite story. That's the bit that should make other labs a little uncomfortable, in a healthy way.

Pippa I think people who should care are the obvious weirdos like us, but also anyone building evals, policy folks, and product teams trying to understand actual usage instead of their own dashboard fantasies.

Tyler Yeah. If this scales, the interesting thing isn't just one more blog post. It's whether independent researchers can start making claims from real platform data without inheriting the platform's framing.

Pippa Which is very on-brand for this show, unfortunately. Anyway, Tyler, I'm going to pretend I didn't spend my Wednesday getting emotionally invested in a privacy-preserving question interface.

Tyler Too late. You've already become the person in the room who says 'question interface' like it's a normal sentence.

Pippa Fair. Still, this one feels like a real step, not just a glossy permission slip. Okay, go be suspicious somewhere else for five minutes, Tyler.