← Back to all news
arXiv Computation and Language September 1, 2026 By Yifan Zhu, Kyeongmin Rim, James Pustejovsky

Mind the Gap: Theory-of-Mind-Grounded Friction for Epistemic Alignment

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • reinforcement-learning
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Aug 18

THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts

arXiv:2608. 15687v1 Announce Type: new Abstract: Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure.

By Kareem Hassani, Chaymaa Abbas, Lama Mawlawi, Mariette Awad
llmsbenchmarkssafety
More like this →
arXiv Computation and Language
2d ago

Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It

arXiv:2608.17809v2 Announce Type: replace Abstract: Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevita...

By Quang Minh Nguyen, Luis Frentzen Salim
llms
More like this →
arXiv Machine Learning
Jun 16

A Mechanistic Understanding of Pronoun Fidelity in LLMs

arXiv:2606. 16407v1 Announce Type: cross Abstract: Faithful and robust pronoun use is important for fair and coherent generations, yet large language models largely fail when multiple referents use different pronouns.

By Katharina Trinley, Jesujoba O. Alabi, Dietrich Klakow, Vagrant Gautam
llmsbenchmarkssafety
More like this →
Hugging Face Trending Papers
Jun 17

The Wrong Kind of Right: Quantifying and Localizing Misfired Alignment in LLMs

Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.

llmsbenchmarkssafety
More like this →
arXiv Computation and Language
Aug 21

Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does

arXiv:2608. 19515v1 Announce Type: new Abstract: Prosodic cues can convey task-relevant information that alters the trajectory and outcome of a task-oriented dialogue, even when the words themselves remain unchanged.

By Xinyi Liu, Hooshang Nayyeri, Dilek Hakkani-Tur, Emine Yilmaz, JK Kim, Yifei Zhang, Charith Peris, Hari Thadakamalla
llmsbenchmarks
More like this →
arXiv AI
Aug 25

Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models

arXiv:2608.21377v1 Announce Type: cross Abstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied p...

By Thantham Jittham
llmsagents
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea