← Back to all news
arXiv AI August 25, 2026 By Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau

Evaluation Awareness in Language Models: Representation, Verbalization, and Control

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

  • llms
  • fine-tuning
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 3

Decomposing and Measuring Evaluation Awareness

arXiv:2605. 23055v2 Announce Type: replace-cross Abstract: Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results.

By Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin, Sahar Abdelnabi, Maksym Andriushchenko
llmsbenchmarkssafety
More like this →
arXiv AI
Jun 17

In-Context Environments Induce Evaluation-Awareness in Language Models

arXiv:2603. 03824v2 Announce Type: replace Abstract: Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{evaluation awareness}.

By Maheep Chaudhary
llmsbenchmarkssafety
More like this →
arXiv AI
Jul 17

Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect

arXiv:2607. 14111v1 Announce Type: cross Abstract: Can small language models detect and report on perturbations their own internal activations?

By Ely Hahami, Ishaan Sinha, Lavik Jain
llmsfine-tuningbenchmarkssafety
More like this →
arXiv AI
Jun 12

Prefill Awareness in Large Language Models

arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.

By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk
llmsagentsbenchmarkssafety
More like this →
arXiv AI
Jul 29

Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models

arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.

By Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
llmssafety
More like this →
arXiv Machine Learning
Jun 30

Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

arXiv:2606. 29196v1 Announce Type: new Abstract: Do language models know when they are being tested?

By Archit Manek
llmsbenchmarkssafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea