arXiv:2608. 16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments.
By Yuchen Yuan, Zhenghuang Wu, Yuangan Li, Liang Ma, Ke Li
arXiv:2606. 30441v1 Announce Type: cross Abstract: A rigorous formalization of system requirements is a fundamental prerequisite for the verification of Multi-Agent Systems (MAS).
By Marco Aruta, Francesco Improta, Vadim Malvone, Aniello Murano, Vladana Perlic
arXiv:2608. 04697v1 Announce Type: new Abstract: Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level.
By Cristian Mascia, Roberto Pietrantuono, Daniel Rodriguez, Stefano Russo
arXiv:2409. 07314v3 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows.
By Praveenkumar Kanithi, Cl\'ement Christophe, Marco AF Pimentel, Tathagata Raha, Prateek Munjal, Nada Saadi, Hamza A Javed, Svetlana Maslenkova, Nasir Hayat, Ronnie Rajan, Shadab Khan
arXiv:2603. 25450v2 Announce Type: replace Abstract: Detecting when a language model is wrong without ground truth labels is a fundamental challenge for safe deployment.
By Matt Gorbett, Suman Jana
arXiv:2608. 08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.
By Dongjie Xu, Julius, Hanchi Dong, Minghua Tang, Yuxuan Sun, Ziwei Nie, Zicheng Liu, Dujun Qing, Jiajie Xu
arXiv:2608. 06202v1 Announce Type: cross Abstract: Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness.
By Ro Encarnaci\'on, Tina Behzad, Emma Lurie, Dana\'e Metaxa
arXiv:2604. 20288v2 Announce Type: replace Abstract: Flight diversions are rare but high-impact events in aviation, making their reliable prediction vital for both safety and operational efficiency.
By Karim Aly, Alexei Sharpanskykh, Jacco Hoekstra
arXiv:2604. 24827v2 Announce Type: replace-cross Abstract: Closed-source frontier labs do not disclose parameter counts.
By Bojie Li
arXiv:2607. 07184v1 Announce Type: cross Abstract: Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model.
By Marcus Williams, Hannah Sheahan, Cameron Raymond, Tomek Korbak, Deng Pan, Peilin Yang, Leon Maksin, Ningyi Xie, Phillip Guo, Ian Kivlichan, Micah Carroll
arXiv:2605. 28591v2 Announce Type: replace-cross Abstract: The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings.
By Katharina Deckenbach, Haritz Puerto, Jonas Geiping, Sahar Abdelnabi
arXiv:2511. 01650v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly entering specialized, safety-critical engineering workflows governed by strict quantitative standards and immutable physical laws, making rigorous evaluation of their reasoning capabilities imperative.
By Ayesha Gull, Muhammad Usman Safder, Rania Elbadry, Fan Zhang, Veselin Stoyanov, Preslav Nakov, Zhuohan Xie