arXiv:2606. 29113v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly mediate strategic interactions through natural language, making semantic control a critical element of communication and deception.
By Quanyan Zhu
arXiv:2502. 19193v2 Announce Type: replace-cross Abstract: Social media platforms frequently impose restrictive policies to moderate user content, prompting the emergence of creative evasion language strategies.
By Jinyu Cai, Yusei Ishimizu, Mingyue Zhang, Munan Li, Jialong Li, Kenji Tei
arXiv:2501. 00745v3 Announce Type: replace-cross Abstract: The increasing integration of Large Language Model (LLM) based search engines has transformed the landscape of information retrieval.
By Xiyang Hu
Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In this work, we reveal that this advanced safety awareness inadvertently introduces a fatal vulnerability.
arXiv:2607. 09766v1 Announce Type: new Abstract: AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards.
By Yaowen Ye, Jacob Steinhardt
arXiv:2607. 23394v1 Announce Type: new Abstract: Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly benign data can transmit hidden preferences that generalize broadly.
By Adhyyan Narang, Artin Tajdini, Claire Zhang, Jamie Morgenstern