arXiv Machine Learning
Aug 27

Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

The study examines whether a decodable "empathy" direction can be used as a reliable causal lever in language models. Using EPITOME-derived facets of Recognition (cognitive) and Resonance (affective) across three instruction‑tuned LLMs, the authors find that while affective steering can partially raise affective scores, cognitive steering shows inconsistent or unmeasurable effects. The results highlight that decodability does not guarantee reliable control, especially for cognitive empathy, and that measurement sensitivity must be explicitly checked.

By Haoran Jisun