arXiv:2603. 23485v2 Announce Type: replace-cross Abstract: Standard evaluation practices assume that large language model (LLM) outputs are stable when prompts are embedded in contextually equivalent discourses.
By Sagar Kumar, Ariel Flint, Luca Maria Aiello, Andrea Baronchelli
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2608.30719v1 Announce Type: new
Abstract: Productive dialogue alignment requires distinguishing \emph{surface coordination} (acknowledgments and smooth task progression) from \emph{epistemic al...
By Yifan Zhu, Kyeongmin Rim, James Pustejovsky
Nastase et al. (2026) argue that large language models (LLMs) can shed light on language processing because both use distributed, context‑sensitive representations shaped by statistical learning, and they advocate for LLM‑brain alignment research. They reject simple cortical “boxology” but claim that representational alignment can constrain mechanistic hypotheses, though it does not itself identify a mechanism. The author critiques this position, pointing out logical, causal, and computational underdetermination and the tension between the authors’ methodological caveats and their conclusion that LLMs could serve as fully mechanistic models of language.
By Elliot Murphy
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
arXiv:2608. 13069v1 Announce Type: new Abstract: Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants.
By Lucia Mal\'i\v{c}kov\'a