← Back to all news
arXiv Machine Learning September 22, 2026 By Navraj Singh, Maheep Chaudhary

Evaluation Awareness Shifts from Format to Context with Model Scale

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • llms

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 17

In-Context Environments Induce Evaluation-Awareness in Language Models

arXiv:2603. 03824v2 Announce Type: replace Abstract: Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models exhibit environment-dependent \textit{evaluation awareness}.

By Maheep Chaudhary
llmsbenchmarkssafety
More like this →
arXiv AI
Jun 3

Decomposing and Measuring Evaluation Awareness

arXiv:2605. 23055v2 Announce Type: replace-cross Abstract: Frontier language models sometimes recognize that they are being evaluated and adjust their behavior, undermining validity of benchmark results.

By Changling Li, Terry Jingchen Zhang, Jie Zhang, Zhijing Jin, Sahar Abdelnabi, Maksym Andriushchenko
llmsbenchmarkssafety
More like this →
arXiv AI
Aug 25

Evaluation Awareness in Language Models: Representation, Verbalization, and Control

arXiv:2608.21766v1 Announce Type: cross Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their beha...

By Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau
llmsfine-tuningbenchmarkssafety
More like this →
arXiv AI
Jun 12

Prefill Awareness in Large Language Models

arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.

By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk
llmsagentsbenchmarkssafety
More like this →
arXiv Machine Learning
Jun 11

Mechanisms of Introspective Awareness

arXiv:2603. 21396v5 Announce Type: replace Abstract: Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introspective awareness.

By Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey
llmsreinforcement-learningfine-tuningbenchmarkssafety
More like this →
arXiv AI
Aug 7

Measuring and Detecting Harmful AI Sycophancy

arXiv:2608. 05624v1 Announce Type: new Abstract: Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful.

By Bohan Jiang, Dawei Li, Yasin Silva, Huan Liu
llms
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea