← Back to all news
arXiv Computation and Language September 17, 2026 By Victoria Popa, Guglielmo Cola, Caterina Senette, Maurizio Tesconi

Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

  • llms
  • benchmarks
  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 29

PRISON: Unmasking the Criminal Potential of Large Language Models

arXiv:2506. 16150v4 Announce Type: replace-cross Abstract: As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify.

By Xinyi Wu, Geng Hong, Pei Chen, Yueyue Chen, Xudong Pan, Min Yang
llmsroboticsbenchmarkssafety
More like this →
arXiv Computation and Language
4d ago

Self-reported archetypes and behavioral failures in Large Language Models

arXiv:2609.15998v1 Announce Type: new Abstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property...

By Tabia Tanzin Prama, Calla Glavin Beauregard, Christopher M. Danforth, Peter Sheridan Dodds
llmsagentssafety
More like this →
arXiv AI
Jul 16

What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors

arXiv:2607. 13162v1 Announce Type: cross Abstract: What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not revealed by prompting alone.

By Winston Zeng, Ali Emami, Jinho Choi
llmsagentsfine-tuningsafety
More like this →
arXiv AI
Jul 28

LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings

arXiv:2607. 24435v1 Announce Type: cross Abstract: Large language models may easily assign personality labels from text, but model interpretability remains an open problem.

By Brittany Harbison, Ashok K. Goel
llmsbenchmarkssafety
More like this →
arXiv AI
Jul 31

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

arXiv:2607. 26853v1 Announce Type: cross Abstract: Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors.

By Ruikang Zhang, Shuo Wang, Qi Su
llms
More like this →
arXiv Computation and Language
5d ago

Data Attribution of Emergent Misalignment with Persona Features

arXiv:2608.11025v2 Announce Type: replace Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...

By Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
llmsroboticsfine-tuningsafety
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea