The paper introduces SILICA, an open instrument designed to evaluate whether large language model (LLM) agent societies replicate human behavioural distributions. Using five environments with human‑anchored data and perturbations, the study finds that most LLMs only match human behaviour at initial stages, failing to reproduce end‑state cooperation or correct acceptance thresholds. The results suggest that current LLM societies can support exploratory claims but do not yet reliably emulate human social dynamics.
By Raad Bin Tareaf
arXiv:2605. 21706v2 Announce Type: replace Abstract: Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations.
By Giorgio Piras, Raffaele Mura, Fabio Brau, Maura Pintor, Luca Oneto, Fabio Roli, Battista Biggio
arXiv:2607. 20492v1 Announce Type: cross Abstract: Language models in production do not write prose.
By Rana Muhammad Usman
arXiv:2609.00760v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly trained to decline queries that fall outside their knowledge (knowledge-based refusal, KR) or violate saf...
By Yuri Son, Seunghee Kim, Hyuhng Joon Kim, Taeuk Kim
arXiv:2605. 00123v3 Announce Type: replace Abstract: Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts.
By Shubham Kumar, Narendra Ahuja
arXiv:2605. 25420v2 Announce Type: replace-cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed globally.
By Khalid Yusuf Dahir
arXiv:2608. 11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs.
By Zafar Hussain, Kristoffer Nielbo
The paper introduces a pragmatics-inspired taxonomy for evaluating how large language models (LLMs) refuse unsafe or inappropriate requests. By applying this framework to 16 modern LLMs across 14 harm categories, the authors find that while refusals are generally explicit and morally charged, they often lack interpersonal facework and instead offer safer alternatives, which can be problematic in sensitive contexts. The study argues for alignment evaluations that assess not just whether LLMs refuse, but how they do so in a contextually adaptive and socially responsible manner.
By Ruoxuan Li, Pinqiao Wang, Sheng Li, Cameron Robert Jones
arXiv:2609.00578v1 Announce Type: new
Abstract: Large Language Models (LLMs) can solve complex problems, but their misuse in high-risk domains can lead to severe consequences. Model providers therefo...
By Rui Yang, Yang Hong, Yichao Xu, Zhengyu Liu, Ziyang Li, Yinzhi Cao
arXiv:2606. 04160v1 Announce Type: cross Abstract: Safety alignment in instruction-tuned large language models (LLMs) depends on a model's ability to reliably refuse to respond to harmful or disallowed requests.
By Anna C. Marbut, Daniel R. Olson, Travis J. Wheeler
arXiv:2607. 24758v1 Announce Type: new Abstract: Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking.
By Cole Alexander Niblett, Alexander Chabot Nanni, Anita K. Rao
arXiv:2608. 15772v1 Announce Type: new Abstract: When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or merely suppressed at the output layer.
By Yiqi Liu, Yang Wang, Songxin Wang, Chenghao Xiao, Chenghua Lin