AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

11,221 stories · RSS feed

arXiv Machine Learning
Jun 9

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning

arXiv:2602. 01357v2 Announce Type: replace Abstract: Self-play post-training methods has emerged as an effective approach for finetuning large language models and turn the weak language model into strong language model without preference data.

By Shangzhe Li, Xuchao Zhang, Chetan Bansal, Weitong Zhang
arXiv Machine Learning
Jun 9

Federated Large Language Models: Current Progress and Future Directions

arXiv:2409. 15723v3 Announce Type: replace Abstract: Large Language Models have achieved impressive performance across diverse applications, yet their training typically depends on centralized data collection, raising serious privacy and governance concerns.

By Yuhang Yao, Jianyi Zhang, Junda Wu, Chengkai Huang, Yu Xia, Tong Yu, Ruiyi Zhang, Sungchul Kim, Ryan Rossi, Ang Li, Lina Yao, Julian McAuley, Yiran Chen, Carlee Joe-Wong
arXiv Machine Learning
Jun 9

Similarity-Distance-Magnitude Activations

arXiv:2509. 12760v5 Announce Type: replace Abstract: We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax activation function, adding Similarity (i.

By Allen Schmaltz
arXiv AI
Jun 9

AeroSpectra Sentinel: An Auditable LLM Prompt-Chaining Decision-Support Workflow for Acute Asthma Risk Assessment from Respiratory Sounds and Clinical Signals

arXiv:2606. 08247v1 Announce Type: cross Abstract: Acute asthma risk assessment requires rapid interpretation of respiratory sounds, oxygenation, airflow limitation, speech ability, work of breathing, mental status, and response to reliever therapy.

By Aueaphum Aueawatthanaphisut
arXiv AI
Jun 9

Bridging Expert Knowledge and Automated Feature Engineering via Self-Evolution

arXiv:2606. 08800v1 Announce Type: new Abstract: In high-stakes settings such as brand compliance, clinical care, and content moderation, machine learning cannot be deployed as opaque oracles: practitioners inspect the features driving model decisions, and models must leverage the expert documentation governing these domains.

By Varun Khurana, Vijval Ekbote, Vashu Chauhan, Yaman Kumar Singla, Rajiv Ratn Shah, Balaji Krishnamurthy
arXiv AI
Jun 9

Supracompetitive Pricing Under AI Monoculture

arXiv:2601. 01279v3 Announce Type: replace-cross Abstract: When competing sellers delegate pricing to a shared AI model, such as a large language model, correlated recommendations combined with performance-driven updates aggregating seller feedback raise a key question: can standard AI deployment practices inadvertently produce supracompetitive pricing?

By Shengyu Cao, Ming Hu