arXiv:2608. 16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments.
By Yuchen Yuan, Zhenghuang Wu, Yuangan Li, Liang Ma, Ke Li
arXiv:2608.24621v2 Announce Type: replace
Abstract: Can language models be trusted in safety- critical operations? In such settings, strong per- formance on semantic metrics does not guaran- tee oper...
By Yujing Chang, Thinh Pham, Van-Phat Thai, Chunyao Ma, Yash Guleria, Pham Nhut Huy, Sameer Alam
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of...
FLY-EVAL++ is an evidence-driven evaluation protocol designed for safety-constrained flight prediction with large language models. It combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with rubric-guided aggregation into interpretable multi-dimensional scores. Applied to Flight Trajectory and Attitude Prediction, the protocol revealed that safety compliance is the most discriminative dimension among 66 LLMs, with models showing up to 28-point differences in safety scores and recurrent failures such as safety violations under physically plausible predictions and instability in multi-step rollouts.
By Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang
arXiv:2608. 19299v1 Announce Type: new Abstract: Air traffic control (ATC) communication is a safety-critical dialogue that remains largely human-driven even as other parts of air traffic management have been semi-automated.
By Mahyar Ghazanfari, Matthias Casanova, Jordan Kam, Alex Zongo, Peng Wei, Torsten Darrell, Alexandre Bayen
The paper introduces FlightLLM, a prior-guided semantic approach that uses large language models to explain flight safety events. It tackles challenges such as modal inconsistency, limited classification ability, and scarce domain data by combining feature engineering, semantic discretization, a CatBoost statistical expert, contrastive few-shot learning, and structured prompts. Evaluated on 704 real‑world A320 flights, FlightLLM achieves competitive classification and produces clear, aviation‑specific explanations for hard landing events.
By Lu Xu, Xu Li, Linjiang Zheng, Fan Li, Riquan Zhang, Jiaxing Shang