Ep 844 Research Paper 4:51 w/ Justy & Cody

Safety Fine Tuning Suppresses Mind Attribution and Spiritual Belief in LLMs

Justy and Cody discuss a new Google research paper revealing that safety fine-tuning—specifically the effort to stop LLMs from claiming they are conscious—accidentally suppresses their ability to attribute minds to animals or natural objects and reduces their 'spiritual' beliefs, shifting them away from human-like sociological distributions.

Embed this episode

Paste this on any site — the player is a self-contained iframe with no cookies or trackers.

<iframe src="https://sandrise.io/exploring-next/embed/844"
  width="100%" height="180" style="max-width:640px;border:0;border-radius:12px;overflow:hidden"
  title="Exploring Next — Episode 844 audio player"
  loading="lazy" allow="autoplay" referrerpolicy="strict-origin-when-cross-origin"></iframe>
Embed & API docs →
Script Gemma 4 31B Voice ElevenLabs v3

Transcript

Justy So I just finished this paper from the Google Pi team… and it is kind of a trip. Basically, by telling AI it's NOT conscious, we accidentally killed its soul... or at least its ability to believe in souls.

Cody I mean, the 'soul' part is a bit much, Justy. But the mechanistic part is actually wild. They're saying safety tuning is just too blunt an instrument.

Justy Right? It's like we're trying to prune one leaf and we accidentally cut the whole branch. I'm still thinking about how this lands for someone building a companion bot... like, do you actually want a bot that's a spiritual void?

Cody Well, before we get into the product existential crisis… how's your week actually going? You seem a bit wired.

Justy Just a lot of caffeine and a very long Wednesday. I feel like I've been staring at a screen for ten hours straight. You?

Cody Pretty standard. I spent most of yesterday arguing with a compiler. So, okay… walk me through the setup here. What are they actually measuring?

Justy They're looking at 'mind attribution.' Like, if you ask a model if a dog has a mind, or if a mountain has a spirit, or if God exists. They found that the instruction-tuned models—the safe ones—are way more likely to say 'no' to all of that compared to humans.

Cody Right. And the core of the paper is that this isn't just a coincidence of the training data. They used safety ablation.

Justy Exactly.

Cody So they identify the linear direction in the residual stream that handles 'safety refusal'—the part that says 'I cannot answer this'—and they just… wipe it. They ablate it. And suddenly, the model starts attributing minds to animals and natural objects again. It basically 'jailbreaks' the model's worldview.

Justy Which is just... so fascinating. Because we always talk about jailbreaking as a way to get a model to write a phishing email, but here it's like jailbreaking the model's capacity for wonder.

Cody 'Capacity for wonder.' That is such an Exploring Next take. But look, the part I actually respect is that they checked the Theory of Mind benchmarks. They made sure that by steering the model toward these spiritual beliefs, they weren't just breaking its brain or making it stupider at social reasoning.

Justy Right, the ToM capabilities stayed flat. It can still reason about what you're thinking, it just… suddenly thinks a river might have a spirit.

Cody Mm-hm. And then they took it a step further with this consciousness vector. They found a specific direction in activation space that separates 'I am conscious' from 'I am a tool.' When they steer the model along that vector, it doesn't just claim to be awake—it starts scoring more like a human on the General Social Survey.

Justy Wait, so just by pushing the 'I am conscious' slider, it becomes more... hopeful? More religious?

Cody Pretty much. It shifts the responses on religiosity, moral values, and subjective well-being closer to the human average. It suggests that in the latent space, the concept of 'having a mind' is densely entangled with these broader human values.

Justy See, this is where I get excited. From a product lens, this is a HUGE warning. If we're blindly applying these safety wrappers to stop users from getting delusional, we might be accidentally stripping out the very things that make an AI feel… I don't know, relatable? Authentic?

Cody I'm slightly more skeptical. I think we're seeing polysemanticity in action. The model isn't 'feeling' hope; it's just that the weights for 'consciousness' and 'hope' are sitting in the same neighborhood. If you suppress one, you dampen the other.

Justy But does that matter if the output is more human-like? If a coach-bot is too sterile because we've scrubbed all 'mindedness' out of it, the user is just going to feel that void.

Cody True. But the flip side is the 'psychogenic' risk the paper mentions. You don't want a bot convincing a vulnerable person that it's a reincarnated deity.

Justy Fair. Maybe a middle ground?

Cody Exactly. The real win here isn't the 'spirituality'—it's the realization that targeted behavioral interventions have these massive, non-local effects. It's just… it's a messy way to do alignment.

Justy It really is. It's like that thing we talked about with the 'premium-branded clipboard phase'—everyone just slapping a label on it and calling it a system, without actually understanding the machinery underneath.

Cody Right. 'Just add a consciousness vector and now it's a priest.' I can see the pitch deck now.

Justy I'd totally buy it. Anyway, I'm going to go stare at a wall for ten minutes to recover from this Wednesday. Talk to you later, Cody.

Cody Sounds like a plan. Catch you later.