AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

8,629 stories · RSS feed

arXiv Machine Learning
Aug 4

FedChronos: Federated Fine-Tuning of Time-Series Foundation Models for Privacy-Preserving Commodity Price Forecasting

arXiv:2608. 01290v1 Announce Type: new Abstract: Time-series foundation models (TSFMs) such as Chronos have demonstrated strong forecasting capabilities across domains, yet adapting them to institutionally fragmented settings, where data cannot be centralized due to regulatory, competitive, or sovereignty constraints, remains unexplored.

By Amit Sharma, Nitin Auluck, Akramul Azim
arXiv Machine Learning
Aug 4

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

arXiv:2608. 01378v1 Announce Type: new Abstract: Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run.

By Shuangxiu (Max), Ma (Zachary), Wenhe (Zachary), Zhao
arXiv Machine Learning
Aug 4

Heterogeneous Multi-Agent Reinforcement Learning for Radio Resource Management under Coupled Finite-Horizon Constraints

arXiv:2608. 01745v1 Announce Type: new Abstract: Maximizing throughput under proportional fairness in dense wireless networks requires jointly managing user association, scheduling, base station (BS) activation, and handover control under hard finite-horizon energy and handover budgets, which induces a fundamental tension between BS-side energy management and user-side handover regulation.

By Yeonseo Jeong, Wonhyeok Ko, Sungweon Hong, Songnam Hong
arXiv Machine Learning
Aug 4

Convex Neural Energy Elements: Monolithic Finite-Element Assembly of Geometry-Parameterized Neural Operators with Stability and Error Guarantees

arXiv:2608. 02036v1 Announce Type: new Abstract: Extending the neural-operator element method from individually trained, fixed-geometry neural elements to a library of reusable, geometry-parameterized element types fails structurally: a field-predicting operator trained by value regression induces an energy whose assembled Hessian is indefinite, and Newton converges to spurious minima (247% error) even with 1%-accurate field predictions.

By Hongyue Jiang, Jianjiang Zhan, Chenzhuo Zhang, Fan Wang
arXiv Machine Learning
Aug 4

Constrained Co-Design for Photonic Bayesian Neural Networks

arXiv:2608. 02229v1 Announce Type: new Abstract: Classical neural networks frequently produce overconfident predictions on ambiguous or out-of-distribution (OOD) data, a liability that grows with each AI system deployed in safety-critical real-world scenarios.

By Hendrik Borras, Xiao Wang, Bernhard Klein, Robin Janssen, Frank Br\"uckerhoff-Pl\"uckelmann, Wolfram Pernice, Holger Fr\"oning
arXiv Machine Learning
Aug 4

Automated ECG Interval Measurement and Wave Delineation Using Fast Fourier Convolution ResNet

arXiv:2608. 00058v1 Announce Type: cross Abstract: Accurate measurement of ECG intervals, including PR, QRS duration, and QT/QTc, is central to cardiac diagnosis, yet the published ECG delineation literature evaluates performance almost exclusively as fiducial-point timing errors on small curated databases, rather than as clinical interval accuracy on large unselected cohorts.

By Farhan Adam Mukadam, Harshit Mishra, Nachiket Makwana, Pradyot Tiwari, Subramani Kandasamy, KVS Hari
arXiv Machine Learning
Aug 4

Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit

arXiv:2608. 02302v1 Announce Type: cross Abstract: Long-horizon coding-agent trajectories are poorly matched to the credit units available to train on: a single action has no stable value, an episode label merges productive exploration with abandoned directions, and a fixed window cuts where the logging mechanics fall.

By Jingxi Wei
arXiv Machine Learning
Aug 4

Development and Validation of a Dynamic Kidney Failure Prediction Model based on Deep Learning: A Real-World Study with External Validation

arXiv:2501. 16388v3 Announce Type: replace Abstract: Background: Chronic kidney disease (CKD), a progressive disease with high morbidity and mortality, has become a significant global public health problem.

By Jingying Ma, Jinwei Wang, Lanlan Lu, Zhiqin Jiang, Mengling Feng, Feifei Zhang, Peng Shen, Yexiang Sun, Shenda Hong, Luxia Zhang
arXiv Machine Learning
Aug 4

Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models

arXiv:2506. 07121v2 Announce Type: replace Abstract: Ensuring the safety and robustness of large language models (LLMs) is a fundamental challenge and a critical prerequisite for the responsible deployment of artificial intelligence.

By Ren-Jian Wang, Ke Xue, Zeyu Qin, Ziniu Li, Sheng Tang, Hao-Tian Li, Shengcai Liu, Zhi Yu, Yuanpeng Tan, Chao Qian