arXiv:2607. 04510v1 Announce Type: cross Abstract: Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.
By Lyndon Drake (University of Oxford), Zandi Eberstadt (University of Oxford)
arXiv:2606. 00545v1 Announce Type: new Abstract: Post-trained language models can recognize their own outputs from a sentence or two out of context.
By Asvin G
arXiv:2608. 11225v1 Announce Type: new Abstract: AI "personality clones" force a re-examination of personal identity in operational terms.
By Luc E. Brunet
arXiv:2605. 03160v2 Announce Type: replace Abstract: The standard protocol for interpreting sparse-autoencoder (SAE) features labels each feature from its top-activating contexts and validates the label by steering that single feature at a typical magnitude.
By Michael A. Riegler, Birk Sebastian Frostelid Torpmann-Hagen
arXiv:2609.22090v1 Announce Type: new
Abstract: An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present Ps...
By Joy Bose
arXiv:2607. 13162v1 Announce Type: cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone.
By Winston Zeng, Ali Emami, Jinho Choi
The paper argues that the current debate on machine consciousness rests on the mistaken assumption that AI systems are already the kind of entities that could possess consciousness. By distinguishing between phenomenal consciousness, introspective report, and human projective introspection, it introduces the AI Consciousness Fallacy, showing that generative models can produce first‑person linguistic traces without being conscious. It then proposes Causal Liability Theory (CLT), with CLT‑I defining liability closure as a criterion for identifying a bearer of consciousness and CLT‑II suggesting that liability closure is both necessary and sufficient for minimal phenomenal subjecthood, and demonstrates experimentally that these distinctions are tractable and can separate causal bearer structure from first‑person performance.
By Afshin Khadangi
arXiv:2608.20647v1 Announce Type: new
Abstract: Splitting a bidirectional LSTM's contextual representation into a forward-only $F_i$ (strictly a function of tokens $1..i$) and a backward-only $B_i$ (...
By Sai Krishna Arthanari, JaeHyeong Chang, Chengzhe Sun, Siwei Lyu
arXiv:2605. 07284v2 Announce Type: replace Abstract: A late-layer change learned during post-training may work on the base model's earlier state, or it may depend on earlier computation learned with it.
By Yifan Zhou
arXiv:2606. 01092v1 Announce Type: cross Abstract: Supervised learning evaluates predictors through their input-output behavior.
By Vasileios Sevetlidis
arXiv:2606. 18509v1 Announce Type: new Abstract: Reliable generalization in conditional latent variable models requires understanding both identifiability and extrapolation: how observed variation across attributes determines latent structure, and how that structure determines distributions at unseen attributes.
By Soheun Yi, Yizhou Lu, Chandler Squires, Pradeep Ravikumar
arXiv:2607. 15883v1 Announce Type: cross Abstract: Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and somehow not quite the presence of a mind.
By Sebastian Cochinescu