arXiv:2605. 18580v2 Announce Type: replace Abstract: Outcome-only evaluation can certify economically unsafe agents: a policy can hit a business KPI while violating deployable behavioral discipline.
By Peiying Zhu, Sidi Chang
arXiv:2509. 23960v2 Announce Type: replace-cross Abstract: Co-optimizing safety and performance in large-scale multi-agent systems remains a fundamental challenge.
By Manan Tayal, Aditya Singh, Shishir Kolathaya, Somil Bansal
arXiv:2606. 20408v3 Announce Type: replace-cross Abstract: Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness under sustained, adaptive adversarial pressure remains poorly characterized.
By Hanwool Lee, Dasol Choi, Bokyeong Kim, Haon Park, Seung Geun Kim
arXiv:2607. 03025v1 Announce Type: new Abstract: The use of Large Language Models (LLMs) across diverse areas of human activity-ranging from everyday tasks to safety-critical applications-aims to enhance decision-making effectiveness with minimal human feedback.
By Andreas Kouridakis, Dimitrios Patiniotis Spyropoulos, George Vouros
arXiv:2607. 02800v1 Announce Type: new Abstract: Induction heads are attention circuits believed to underlie in-context learning in transformers, yet a precise characterization of the estimators they implement remains elusive.
By Francesco D'Angelo, Oguz Kaan Yuksel, Swathi Shree Narashiman, Nicolas Flammarion
arXiv:2602. 17750v2 Announce Type: replace-cross Abstract: A key problem of solid mechanics is the identification of the constitutive law of a material, that is, the relation between strain history and stress.
By Chenyi Ji, Kian P. Abdolazizi, Hagen Holthusen, Christian J. Cyron, Kevin Linka
arXiv:2607. 05163v1 Announce Type: cross Abstract: AI systems may produce failures after deployment that pre-deployment safety assessments do not anticipate.
By Harleen Kaur Sidhu, Rebecca Scholefield, Nour Annan, Kevin Hernandez, Isabel Nieh Hou, Abdulrahman Alshaikhi, Ze Shen Chin, Rokas Gipi\v{s}kis
arXiv:2607. 01983v1 Announce Type: cross Abstract: Robust 3D object detection under adverse weather remains a critical hurdle for autonomous driving.
By Shuyao Li, Chuanxing Geng, Heyang Sun, Qiang Zhou, Jingjing Gu
arXiv:2603. 06921v2 Announce Type: replace-cross Abstract: Safe navigation of autonomous robots remains one of the core challenges in the field, especially in dynamic and uncertain environments.
By Bojan Deraji\'c, Sebastian Bernhard, Wolfgang H\"onig
arXiv:2508. 10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks.
By Jinhwa Kim, Ian G. Harris
arXiv:2607. 03640v1 Announce Type: cross Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic.
By Taras Kutsyk, Bartosz Zieli\'nski
arXiv:2607. 03392v1 Announce Type: cross Abstract: The ever-increasing collection of personal data has created mounting pressure to develop technologies that protect sensitive aspects of individual identity.
By Gergely Flamich, Oyk\"u S{\i}la G\"uner, Yanxiao Liu, Deniz G\"und\"uz
arXiv:2602. 24115v2 Announce Type: replace Abstract: Open RAN (O-RAN) exposes rich control and telemetry interfaces across the Non-RT RIC, Near-RT RIC, and distributed units, but also makes it harder to operate multi-tenant, multi-objective RANs in a safe and auditable manner.
By Zhizhou He, Yang Luo, Xinkai Liu, Mahdi Boloursaz Mashhadi, Mohammad Shojafar, Merouane Debbah, Rahim Tafazolli
arXiv:2607. 05198v1 Announce Type: cross Abstract: Minimum Bayes Risk (MBR) decoding yields more robust and higher-quality text generation than maximum a posteriori (MAP) decoding by selecting hypotheses that maximize expected utility over sampled pseudo-references.
By Yusuke Sakai, Hidetaka Kamigaito, Taro Watanabe
arXiv:2607. 04673v1 Announce Type: cross Abstract: Glaucoma is a leading cause of irreversible blindness worldwide, yet most automated diagnosis systems rely on opaque deep-learning models that offer little clinical interpretability.
By Cheng Huang, Jia Zhang, Yi Jiang, Yang Liu, Karanjit Kooner, Yadi Liu, Tsengdar Lee, Yang Xie, Wenqi Shi, Guanghua Xiao
arXiv:2601. 22652v2 Announce Type: replace-cross Abstract: Spectral gradient methods, such as the Muon optimizer, modify gradient updates by preserving directional information while discarding scale, and have shown strong empirical performance in deep learning.
By Guillaume Braun, Han Bao, Wei Huang, Masaaki Imaizumi
arXiv:2405. 19521v3 Announce Type: replace Abstract: In applied statistics and machine learning, the gold standards used for training are often biased and almost always noisy.
By Seong Woo Han, Ozan Ad{\i}g\"uzel, Bob Carpenter
arXiv:2603. 07606v2 Announce Type: replace Abstract: Interpretable machine learning is essential in high-stakes domains where decision-making requires accountability, transparency, and trust.
By Hans Farrell Soegeng, Sarthak Ketanbhai Modi, Thomas Peyrin
arXiv:2506. 07468v4 Announce Type: replace Abstract: Conventional large language model (LLM) safety alignment relies on a reactive, disjoint loop: attackers exploit a static model, then defenders patch exposed vulnerabilities.
By Mickel Liu, Liwei Jiang, Yancheng Liang, Simon Shaolei Du, Yejin Choi, Tim Althoff, Natasha Jaques
arXiv:2509. 22267v5 Announce Type: replace Abstract: Reliable detection of bearing faults is essential for maintaining the safety and operational efficiency of rotating machinery.
By Jo\~ao Paulo Vieira, Victor Afonso Bauler, Rodrigo Kobashikawa Rosa, Danilo Silva