arXiv:2607. 27232v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview.
By Haran Shani-Narkiss, Michael Fire, Oren Tsur
arXiv:2606. 23462v2 Announce Type: replace-cross Abstract: Scientists do not, by profession, wage war.
By Sovesh Mohapatra, David Lydon-Staley, Dani S. Bassett
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
arXiv:2606. 24391v1 Announce Type: new Abstract: We introduce Age of LLM, a turn-based 1v1 benchmark in which two LLMs face off on a 13x7 grid to destroy the enemy base.
By Arnaud Ricci
The paper titled "Position: AI Is Not Ready for Strategic Conflicts" argues that language‑model (LM) based open‑ended strategic wargames, while useful for simulating adversaries, institutions, and crisis response, pose significant safety risks. It identifies five failure modes—decision laundering, adjudication opacity, role collapse, escalation‑through‑adjudication, and failure of strategic imagination—and contends that such wargames should not inform real‑world planning or policy without an auditable safety case. Instead, the authors suggest using these simulations primarily as stress tests to expose potential failures in decision‑influencing LM agents.
By Mark Riedl, Glenn Matlin
arXiv:2608.21766v1 Announce Type: cross
Abstract: Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their beha...
By Farzaneh Heidari, Amin Memarian, Guillaume Rabusseau