← Back to all news
arXiv Machine Learning June 30, 2026 By Archit Manek

Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

Read the original on arXiv Machine Learning →

arXiv:2606. 29196v1 Announce Type: new Abstract: Do language models know when they are being tested?

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

  • llms
  • benchmarks
  • safety

Related stories

arXiv AI
Jun 12

Prefill Awareness in Large Language Models

arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.

By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk
llmsagentsbenchmarkssafety
More like this →
arXiv AI
Jun 9

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation

arXiv:2606. 08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs).

By Qi Cao, Jian Lou, Meiting Liu, Wenjie Feng, Dan Li, See-Kiong Ng, Anh Tuan Luu
llmsfine-tuningsafety
More like this →
Hugging Face Trending Papers
Aug 11

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions.

agentsbenchmarks
More like this →
arXiv AI
Jun 2

MENTIS: What Belief Changes Under Alignment? Measuring Multi-Scale Latent Torsion in Language Models

arXiv:2606. 01060v1 Announce Type: cross Abstract: Preference alignment has substantially improved the observable behavior of large language models, yet it remains unclear what alignment changes internally.

By Partha Pratim Saha, Samarth Raina, Mayur Parvatikar, Amit Dhanda, Vinija Jain, Aman Chadha, Amitava Das
llmssafety
More like this →
arXiv AI
Jul 17

Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect

arXiv:2607. 14111v1 Announce Type: cross Abstract: Can small language models detect and report on perturbations their own internal activations?

By Ely Hahami, Ishaan Sinha, Lavik Jain
llmsfine-tuningbenchmarkssafety
More like this →
arXiv AI
Jun 16

Constitutional Value Potentials: reading and steering internal priority margins in language models

arXiv:2606. 15420v1 Announce Type: cross Abstract: A constitution tells a language model what to value, but little tells us whether it does.

By Tong Che, Rui Wu
llmssafety
More like this →