AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

9,515 stories · RSS feed

arXiv Machine Learning
Jul 17

Human-In-The-Loop Machine Learning for Safe and Ethical Autonomous Vehicles: Principles, Challenges, and Opportunities

arXiv:2408. 12548v3 Announce Type: replace Abstract: Machine Learning (ML) has become central to Autonomous Vehicles (AVs), supporting perception, prediction, planning, control, and decision-making in dynamic environments.

By Yousef Emami, Mohammadhossein Homaei, Miguel Guti\'errez Gait\'an, Luis Almeida, Kai Li, Hui Huang, Zhu Han
arXiv Machine Learning
Jul 17

Cross-Cluster Weighted Forests

arXiv:2105. 07610v5 Announce Type: replace-cross Abstract: Building trustworthy machine learning algorithms for biological applications requires adapting to data heterogeneity from different sources, batches, distributions, or studies.

By Maya Ramchandran, Rajarshi Mukherjee, Giovanni Parmigiani
arXiv Machine Learning
Jul 17

DiMaS: Distribution Matching for Steering Vision-Language-Action Models

arXiv:2607. 14280v1 Announce Type: cross Abstract: Flow-matching-based vision-language-action (VLA) models have emerged as powerful policies for robotic manipulation, yet a critical capability remains underexplored: fine-grained behavioral control, the ability to govern how a robot performs a task by intervening on its internal representations.

By Pegah Khayatan, Sara Meziane, Jayneel Parekh, Matthieu Cord
arXiv AI
Jul 17

HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization

arXiv:2607. 14349v1 Announce Type: cross Abstract: While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content effects, demonstrating a measurable bias towards real-world plausibility.

By Abdullah Shaikh, Zain Naqi, Taha Zahid, Sandesh Kumar, Abdul Samad
arXiv Machine Learning
Jul 17

XCT-SAM: Sequential Parameter-Efficient Domain Adaptation of SAM for Industrial XCT Defect Segmentation

arXiv:2607. 14287v1 Announce Type: cross Abstract: Defect segmentation in additive manufacturing (AM) X-ray computed tomography (XCT) images remains challenging due to severe class imbalance and large distribution shifts across scan conditions.

By Md Mahedi Hasan, Md Mushfiqur Rahaman, Alan Pachkovskiy, Imtiaz Ahmed, Jeremy Dawson, Srinjoy Das
arXiv Machine Learning
Jul 17

Trajectory-Aware Flow Matching for Topology Optimisation

arXiv:2607. 14652v1 Announce Type: new Abstract: Topology optimisation (TO) often requires repeated finite element analysis and sensitivity-based material updates, which can be costly when multiple candidate designs are needed under varying physical and design conditions.

By Shusheng Xiao, Jinshuai Bai, Hyogu Jeong, Yunfei Xi, Yilin Gui, YuanTong Gu
arXiv AI
Jul 17

Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation

arXiv:2607. 14203v1 Announce Type: cross Abstract: 3D simulation platforms are critical for autonomous driving because they enable end-to-end policy evaluation, thereby reducing development costs and improving safety.

By NVIDIA, :, Jiahui Huang, Jiawei Ren, Michal Tyszkiewicz, Bjoern Haefner, Michael Shelley, Xin Kang, Seung Wook Kim, Ning Xu, Qi Wu, Janick Martinez Esturo, Shengyu Huang, Nick Schneider, Laura Leal-Taixe, Zan Gojcic, Sanja Fidler
arXiv Machine Learning
Jul 17

MamaBench: Benchmarking LLM Robustness in Maternal and Child Health Diagnosis through Counterfactual Clinical Perturbation

arXiv:2607. 14385v1 Announce Type: cross Abstract: Large language models achieve strong scores on medical benchmarks, yet these benchmarks evaluate each question in isolation, providing no measure of whether a system can distinguish clinically similar presentations requiring different interventions.

By Thanni Adewuyi, Anuoluwa Sotome, Samuel Okoko, Angel Ezendu, Oluwafunke Akinbuwa, Oluwaseun Odunsi, Oluwasegun Oguntuase, Oluwadarasimi Oguntuase, Ifeoma Nwabueze, Abiodun Adereni
arXiv AI
Jul 17

MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents

arXiv:2607. 14651v1 Announce Type: cross Abstract: Persistent external memory enhances agent continuity but introduces persistent security vulnerabilities: adversarial content can be injected via standard interaction channels, retained across turns, and later distort downstream behavior.

By Jifeng Gao, Kang Xia, Yi Zhang, Xiaobin Hong, Mingkai Lin, Xingshen Wei, Wenzhong Li, Sanglu Lu