AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

8,366 stories · RSS feed

arXiv AI
2d ago

UC-PSRO: Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum for Game-Theoretic Course-of-Action Generation in Adversarial Swarms

arXiv:2608. 15372v1 Announce Type: new Abstract: We study generating game-theoretically optimized Courses of Action (COAs) for a Blue UAS swarm against an adaptive Red adversary in a communication-degraded environment, motivated by (but not derived from) a public U.

By Phillip Jiang
arXiv AI
2d ago

JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills

arXiv:2608. 16465v1 Announce Type: new Abstract: Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematically integrate, reuse, and improve at scale.

By Xiaoyu Wen, Jiajia Li, Zhida He, Peng Yu, Chenxu Wang, Han Qi, Ziyuan Zhou, Cheng Jin, Ying Wen, Xingcheng Xu, Shuyue Hu, Tianhang Zheng, Chaochao Lu, Qiaosheng Zhang
arXiv AI
2d ago

MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment

arXiv:2608. 14828v1 Announce Type: new Abstract: Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no real help.

By Tony Tu, Sayan Chakraborty, Ruomeng Xu, Tony Qin, Austin Tian
arXiv AI
2d ago

When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry

arXiv:2608. 14680v1 Announce Type: new Abstract: Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent messages), not of the final answer alone, yet evaluating only task outcomes reveals little about how or why a run fails.

By Chenkai Zhang, Yiran Li, Yifang Tian, Michalis Bachras, Hans-Arno Jacobsen
arXiv AI
2d ago

OGX: An Open-Source, Vendor-Neutral Generative AI Application Server

arXiv:2608. 14580v1 Announce Type: new Abstract: OGX (Open GenAI Stack) is an open-source AI application server and Python library that implements the APIs of major frontier labs (OpenAI, Anthropic, Google) with pluggable backend providers.

By Francisco Javier Arceo, S\'ebastien Han, Matthew Farrellee, Charlie Doern, Yuan Tang, Derek Higgins, Varsha Prasad Narsing, Gordon Sim, Sumanth Kamenani, Ben Browning, Raghotham Murthy
arXiv AI
2d ago

When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning

arXiv:2608. 14610v1 Announce Type: new Abstract: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination.

By Yiqian Huang, Shuyuan Zheng, Qianying Liu, Shaowen Peng, Yuntao Kong, Kotaro Funakoshi, Chuan Xiao, Manabu Okumura, Yang Cao
arXiv AI
2d ago

Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152

arXiv:2608. 16394v1 Announce Type: new Abstract: Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards.

By Vahid Zolfaghari, Nenad Petrovic, Andr\'E Schamschurko, Alois Knoll
arXiv AI
2d ago

Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework

arXiv:2608. 14584v1 Announce Type: cross Abstract: In Multimodal Question Answering (MQA), models are required to jointly encode and integrate heterogeneous information from multiple modalities, including text, images, and speech, to perform complex semantic reasoning and decision making.

By Hailong Yang, Jianqi Wang, Guanjin Wang, Zhaohong Deng
arXiv AI
2d ago

FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making

arXiv:2608. 15004v1 Announce Type: cross Abstract: Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool for screening and followup assessment.

By Pramit Dutta, Jenita Manokaran, Richa Mittal, Ryan Appleby, Eranga Ukwatta