The paper introduces the Profiling, Investigation, and Judgment (PIJ) benchmark, which contains 2,500 real homicide cases from five countries to evaluate large language models (LLMs) on pre‑arrest criminal investigation tasks. It assesses LLMs across criminal profiling, crime process reconstruction, and sentence prediction, revealing that performance drops as tasks require more implicit reasoning about unknown suspect profiles. The study finds that LLMs lag behind human experts, especially on inferential tasks like motivation and victim‑offender relationships, and exhibit biases in gender, age, and motive attribution.
By Yutong Yao, Yanjie Cao, Guanhua Chen, Xu Yang, Junchao Wu, Zeyu Wu, Lidia S. Chao, Derek F. Wong
arXiv:2511. 19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves the way for a more significant one, to bypass safety alignments, pose a persistent threat to Large Language Models (LLMs).
By Adarsh Kumarappan, Ananya Mujoo
The paper examines how large language models (LLMs) alter the expression of Dark Triad traits—Machiavellianism, narcissism, and psychopathy—when prompted to fake good or fake bad. Across seven state‑of‑the‑art models and two real‑world contexts (employment selection and forensic evaluation), most models lowered trait scores under fake‑good conditions and raised them under fake‑bad conditions, with varying consistency across traits and models. The study also finds that explicit fake‑bad instructions produce stronger distortions than contextual framing alone, underscoring the influence of motivational and situational context on personality‑related outputs.
By Victoria Popa, Guglielmo Cola, Caterina Senette, Maurizio Tesconi
arXiv:2606. 17478v1 Announce Type: cross Abstract: As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern.
By Kexin Chen, Yi Liu, Haonan Zhang, Yanhui Li, Xinyu Deng, Dongxia Wang
The paper introduces the Profiling, Investigation, and Judgment (PIJ) benchmark, which contains 2,500 real homicide cases from five countries to evaluate large language models (LLMs) on pre‑arrest criminal investigation tasks. It assesses LLMs across criminal profiling, crime process reconstruction, and sentence prediction, revealing that performance drops as tasks require more implicit reasoning about unknown suspect profiles. The study finds that LLMs lag behind human experts, especially in inferential categories like motivation and victim‑offender relationships, and exhibit biases in gender, age, and motive attribution.
arXiv:2607. 29066v1 Announce Type: cross Abstract: Deception detection has critical implications for legal proceedings, law enforcement, and online security.
By Theekshana Samaradiwakara, Nisansa de Silva, George C. Lobb
arXiv:2504.00285v2 Announce Type: replace
Abstract: Large Language Models (LLMs) are effective at deceiving when prompted to do so. Models that demonstrate better performance on reasoning tasks are a...
By Samuel M. Taylor, Benjamin K. Bergen
arXiv:2607. 14791v1 Announce Type: new Abstract: Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model behaviour.
By Darius Lim, Nathan Leow, Xin Wei Chia
arXiv:2608.23028v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where users engage th...
By Zeyu Feng, Qingyu Wu, Yuzhe Luo, Hua Cheng
The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.
By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
arXiv:2607. 20444v1 Announce Type: cross Abstract: Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally induced goal.
By Ali Asad, Stephen Obadinma, Anshul Pattoo, Wenxuan Zhang, Xiaodan Zhu
The paper introduces a counterfactually anchored evidence attribution approach for multi‑turn large language model safety failures. It presents a new dataset of 1,762 conversations, including adversarial, benign twins, and high‑risk vocabulary variants, and trains a lightweight hierarchical model that accurately predicts safety violations and attributes them to specific user turns and token spans. The model achieves high detection performance (F1 = 0.988) and significantly reduces adversarial confidence when top‑attributed tokens are removed, while maintaining low false‑positive rates on benign conversations.
By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan