Topic
Polysemanticity
2 episodes
-
Safety Fine Tuning Suppresses Mind Attribution and Spiritual Belief in LLMs
Justy and Cody discuss a new Google research paper revealing that safety fine-tuning—specifically the effort to stop LLMs from claiming they are conscious—accidentally suppresses their ability to attribute minds to animals or natural objects and reduces their 'spiritual' beliefs, shifting them away from human-like sociological distributions.
-
Overview: Model Interpretability
We slow down and make model interpretability actually click: what it means to explain a model, what the main tools can and cannot show, and why the difference between a useful explanation and a comforting story matters.