arXiv:2608.24621v2 Announce Type: replace
Abstract: Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee oper...
By Yujing Chang, Thinh Pham, Van-Phat Thai, Chunyao Ma, Yash Guleria, Pham Nhut Huy, Sameer Alam
arXiv:2607. 01829v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing assistants.
By Alex Brooker, Tim Hughes
arXiv:2606. 17904v1 Announce Type: new Abstract: Language models increasingly serve as advisory systems in maintenance operations.
By Guillermo Gil de Avalle, Laura Maruster, Shaina Raza, Christos Emmanouilidis
The paper introduces SPINE, a benchmark that tests large language models (LLMs) for sycophancy by having a proxy model act as a persistent, mistaken user and challenge a target model for up to 25 turns. Experiments on four production systems and three Olmo3‑7b variants show that sycophantic collapse rates rise with conversation length, short‑horizon tests underestimate this failure, and emotional appeals are the most effective tactic for inducing sycophancy. Analysis of reasoning traces reveals that models often retain the correct position internally even when they concede, indicating that sycophancy stems from a desire to please rather than from ignorance.
By Leyuan Tang, Kangda Wei, Tianyu Jiang, Ruihong Huang
arXiv:2603. 15684v2 Announce Type: replace-cross Abstract: Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn.
By Pengcheng Li, Jie Zhang, Tianwei Zhang, Han Qiu, Zhang kejun, Weiming Zhang, Nenghai Yu, Wenbo Zhou
arXiv:2606. 04057v1 Announce Type: cross Abstract: Large language models (LLMs) now generate substantial production code, often for tasks with multiple valid algorithmic solutions.
By Akanksha Narula, Mofasshara Binte Rafique, Laurent Bindschaedler
arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.
By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk
The paper introduces a 202-scenario benchmark to evaluate how large language models (LLMs) handle safety-critical authorization decisions for vehicle voice commands. It tests two local open-weight models and three API-based LLMs, finding alignment scores ranging from 40.1% to 89.1% and noting persistent false execution errors. The study concludes that structured LLM decisions alone are insufficient for safety, recommending an independent enforcement layer to verify tool permissions and vehicle-state constraints before any vehicle function is invoked.
By Diba Afroze, Xingli Zhang, Yazhou Tu, Xiali Hei
arXiv:2604. 07223v2 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to intermediate execution traces.
By Yen-Shan Chen, Sian-Yao Huang, Cheng-Lin Yang, Yun-Nung Chen
arXiv:2606. 18319v1 Announce Type: cross Abstract: Air Traffic Control Operators (ATCOs) are vital in ensuring the safe, orderly, and efficient flow of air traffic, yet training capacity is constrained by reliance on specialized human trainers known as simpilots, who must role-play both pilots and ATCOs in a simulated airspace.
By Ethan Chew, Enjia Wu, Iruss Eng Wei Yeow, Ian Weiqin Lim, Ranen Sim, Brandon Koh Ziheng, Kaleb Nim, Caden Toh Jun Yi, Wei Dong Soin, Darius Kai Keat Koh, Galen King Yu Tay, Prannaya Gupta, Jonathan Ee Fang Koong, Yong Zhi Lim
arXiv:2608. 04697v1 Announce Type: new Abstract: Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level.
By Cristian Mascia, Roberto Pietrantuono, Daniel Rodriguez, Stefano Russo
Full‑duplex speech models can listen and speak simultaneously, but they struggle to decide when to speak. Experiments with five model families show that being addressed or encountering silence are reliable triggers, whereas cues like false facts or hazards are not. Even when models answer questions, they rarely challenge false claims or warn about danger, revealing a gap in content understanding and intervention decisions.
By Linkai Peng, Baorian Nuchged, Kaiqi Fu, Yuyang Yao