AI safety and alignment

Alignment, interpretability, red-teaming, bias and privacy: the research on what these systems do when they misbehave.

10,300 stories · RSS feed

arXiv Machine Learning
Jul 7

CSympNet-ID: conformal-symplectic map learning for linearly damped Hamiltonian systems

arXiv:2607. 03339v1 Announce Type: new Abstract: Learning dissipative dynamics from discrete observations is essential for reliable long-horizon prediction and physically meaningful parameter identification.

By Jiale Gong (School of Mathematics), Pengzhan Jin (National Engineering Laboratory for Big Data Analysis and Applications, Peking University, Beijing, China), Dongyang Kuang (School of Mathematics), Lu Li (School of Mathematics), Yifa Tang (State Key Laboratory of Mathematical Sciences, Academy of Mathematics and Systems Science, Chinese Academy of Sciences, Beijing, China)
arXiv AI
Jul 7

Open Problems in AI Incident Governance

arXiv:2607. 05163v1 Announce Type: cross Abstract: AI systems may produce failures after deployment that pre-deployment safety assessments do not anticipate.

By Harleen Kaur Sidhu, Rebecca Scholefield, Nour Annan, Kevin Hernandez, Isabel Nieh Hou, Abdulrahman Alshaikhi, Ze Shen Chin, Rokas Gipi\v{s}kis
arXiv Machine Learning
Jul 7

Induction Heads Interpolate N-Grams

arXiv:2607. 02800v1 Announce Type: new Abstract: Induction heads are attention circuits believed to underlie in-context learning in transformers, yet a precise characterization of the estimators they implement remains elusive.

By Francesco D'Angelo, Oguz Kaan Yuksel, Swathi Shree Narashiman, Nicolas Flammarion
arXiv Machine Learning
Jul 7

Agentic AI-RAN: Enabling Intent-Driven, Explainable and Self-Evolving Open RAN Intelligence

arXiv:2602. 24115v2 Announce Type: replace Abstract: Open RAN (O-RAN) exposes rich control and telemetry interfaces across the Non-RT RIC, Near-RT RIC, and distributed units, but also makes it harder to operate multi-tenant, multi-objective RANs in a safe and auditable manner.

By Zhizhou He, Yang Luo, Xinkai Liu, Mahdi Boloursaz Mashhadi, Mohammad Shojafar, Merouane Debbah, Rahim Tafazolli