AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

11,407 stories · RSS feed

arXiv AI
Jun 4

Efficient Adversarial Attacks on High-dimensional Offline Bandits

arXiv:2602. 01658v2 Announce Type: replace-cross Abstract: Bandit algorithms have recently emerged as a powerful tool for evaluating machine learning models, including generative image models and large language models, by efficiently identifying top-performing candidates without exhaustive comparisons.

By Seyed Mohammad Hadi Hosseini, Amir Najafi, Mahdieh Soleymani Baghshah
arXiv Machine Learning
Jun 4

Plug-and-Play Diffusion Meets ADMM: Dual-Variable Coupling for Robust Medical Image Reconstruction

arXiv:2602. 23214v2 Announce Type: replace-cross Abstract: Plug-and-Play diffusion prior (PnPDP) frameworks have emerged as a powerful paradigm for solving imaging inverse problems by treating pretrained generative models as modular priors.

By Chenhe Du, Xuanyu Tian, Qing Wu, Muyu Liu, Jingyi Yu, Hongjiang Wei, Yuyao Zhang
arXiv AI
Jun 4

Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text

arXiv:2605. 29076v2 Announce Type: replace-cross Abstract: LLMs have advanced text classification, yet existing paradigms face a trade-off: supervised (label only) fine-tuning is scalable but offers limited reasoning on complex text and lacks broader model transparency, while discrete prompt optimization offers human-readable instructions but struggles with performance and scalability.

By Tianyang Zhou, Wenbo Chen, Pierre Jinghong Liang, Leman Akoglu
arXiv AI
Jun 4

Generative Augmented Inference

arXiv:2604. 14575v2 Announce Type: replace-cross Abstract: Large language models enable inexpensive AI-generated annotations, but using them reliably for causal inference remains challenging.

By Cheng Lu, Mengxin Wang, Dennis J. Zhang, Heng Zhang
arXiv Machine Learning
Jun 4

The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench

arXiv:2605. 24782v2 Announce Type: replace Abstract: While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than underlying structural invariants, making even perception-based out-of-distribution accuracy a poor proxy for scientific utility.

By Dingling Yao, Andrea Polesello, Adeel Pervez, Caroline Muller, Francesco Locatello
arXiv AI
Jun 4

Cascading Hallucination in Agentic RAG: The CHARM Framework for Detection and Mitigation

arXiv:2606. 04435v1 Announce Type: new Abstract: Multi-step agentic retrieval-augmented generation (RAG) pipelines have demonstrated significant capability for complex reasoning tasks, yet remain vulnerable to a class of failure that existing hallucination detection mechanisms systematically miss: cascading hallucination, where errors introduced at early pipeline stages propagate and amplify across successive reasoning steps, producing confident but factually incorrect final outputs.

By Saroj Mishra
arXiv Machine Learning
Jun 4

An Open-Source Two-Stage Computer Vision Pipeline for Fine-Grained Vehicle Classification using Vision Transformers

arXiv:2606. 05149v1 Announce Type: cross Abstract: Vehicle body type is a significant determinant of cyclist injury severity in overtaking crashes, yet automated tools for classifying vehicles into injury-risk-relevant categories from naturalistic roadway video do not exist in the open literature.

By Gandhimathi Padmanaban, Fred Feng
arXiv AI
Jun 4

Dynamic Content Moderation in Livestreams: Combining Supervised Classification with MLLM-Boosted Similarity Matching

arXiv:2512. 03553v3 Announce Type: replace-cross Abstract: Content moderation remains a critical yet challenging task for large-scale user-generated video platforms, especially in livestreaming environments where moderation must be timely, multimodal, and robust to evolving forms of unwanted content.

By Wei Chee Yew, Hailun Xu, Sanjay Saha, Xiaotian Fan, Hiok Hian Ong, David Yuchen Wang, Kanchan Sarkar, Zhenheng Yang, Danhui Guan