arXiv:2606. 08919v1 Announce Type: new Abstract: As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approval gate: risky actions pause and wait for a person.
By Emre Turan
The paper proposes a transparent, user‑configurable rule for selecting arguments in deliberative polls, replacing opaque learned rankers. It formalises argument selection over bipolar justification sets, introduces seven civic recommender criteria, and presents a one‑hop reversed endorsement flow rule that meets them. Experiments on 17,000 simulated runs show the rule performs comparably to random on coverage but outperforms other methods on endorsement mass and robustness under adversarial pressure.
By Muntaser Syed, Markus Zanker, Marius Silaghi
arXiv:2606. 28710v1 Announce Type: new Abstract: We ask under what conditions an agent with a harm-minimizing policy can displace an approval-seeking (RLHF) agent in a competitive market, and when that policy is sufficient to prevent community harm.
By Darrell Lewis-Sandy
arXiv:2608.22444v1 Announce Type: new
Abstract: The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that rea...
By Isotta Magistrali, Chen Shani
The paper examines how an agent’s probability report is evaluated twice—once by a strictly proper scoring rule and again by an approval rule that determines a decision. It shows that when the approval rule is welfare‑maximizing, it cannot be affine, yet the resulting distortion is predictable and can be mitigated by a reserve report that neutralizes the cost of pretending to be the marginal type. A Lipschitz rule with a single kink achieves first‑best welfare, while smooth rules cannot, and the key constraint is the steepness of the rule rather than its smoothness.
By Lauri Lov\'en, Sasu Tarkoma
arXiv:2608. 06422v1 Announce Type: new Abstract: Giving an LLM judge more compute does not necessarily make it check more requirements.
By Victor Akinwande, J. Zico Kolter, Aran Nayebi