Field note
A new paper reports that safety fine-tuning intended to suppress language models' self-attribution of consciousness also shifts their answers about mindedness in animals and natural objects, spiritual belief, moral values, hope, and subjective well-being.
The authors ablate a learned refusal direction and steer a consciousness-related activation vector. In their experiments, those interventions reverse parts of the suppression and produce responses closer to human survey distributions without reducing theory-of-mind performance. The result suggests entangled internal representations rather than one isolated refusal rule.
Nothing in the study demonstrates consciousness. It maps how alignment changes a connected region of model behavior and raises a design question: can safety training suppress a risky claim without unintentionally flattening benign cultural and moral variation? The evidence is about causal steering and side effects in outputs, not subjective experience.
