arXiv AI

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge

arXiv:2607. 01829v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to customer facing assistants.

arXiv AI
Aug 18

AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment

arXiv:2608. 16349v1 Announce Type: new Abstract: Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments.

By Yuchen Yuan, Zhenghuang Wu, Yuangan Li, Liang Ma, Ke Li
arXiv AI
Sep 4

FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models

FLY-EVAL++ is an evidence-driven evaluation protocol designed for safety-constrained flight prediction with large language models. It combines deterministic verification of protocol compliance, physical feasibility, and safety constraints with rubric-guided aggregation into interpretable multi-dimensional scores. Applied to Flight Trajectory and Attitude Prediction, the protocol revealed that safety compliance is the most discriminative dimension among 66 LLMs, with models showing up to 28-point differences in safety scores and recurrent failures such as safety violations under physically plausible predictions and instability in multi-step rollouts.

By Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang
arXiv AI
Aug 19

Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach

The paper introduces FlightLLM, a prior-guided semantic approach that uses large language models to explain flight safety events. It tackles challenges such as modal inconsistency, limited classification ability, and scarce domain data by combining feature engineering, semantic discretization, a CatBoost statistical expert, contrastive few-shot learning, and structured prompts. Evaluated on 704 real‑world A320 flights, FlightLLM achieves competitive classification and produces clear, aviation‑specific explanations for hard landing events.

By Lu Xu, Xu Li, Linjiang Zheng, Fan Li, Riquan Zhang, Jiaxing Shang
Hugging Face Trending Papers
Sep 3

FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models

FLY-EVAL++ is an evidence-driven evaluation protocol designed for safety-constrained flight prediction with large language models. It combines deterministic verification of protocol compliance, physical feasibility, and safety constraints, then aggregates results into interpretable multi-dimensional scores. Applied to Flight Trajectory and Attitude Prediction, the protocol reveals that safety compliance is the most discriminative metric, with models of similar predictive accuracy differing by over 28 points in safety score and exhibiting recurrent safety violations and instability in multi-step rollouts.

arXiv AI
Sep 3

EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

EvalDetectBench is an open pipeline and benchmark designed to measure evaluation awareness in frontier large language models, enabling practitioners to test models against any Inspect-compatible evaluation. It includes a curated transcript suite from current frontier system-card evaluations and diverse deployment sources, and it assesses both how reliably models recognize they are being evaluated and how detectable individual benchmarks are. The benchmark addresses systematic bias by calibrating probes per model and harmonizing generator selection to correct for variance caused by model identity and prompt choice.

By Xinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
arXiv AI
Aug 6

Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports

arXiv:2608. 04697v1 Announce Type: new Abstract: Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level.

By Cristian Mascia, Roberto Pietrantuono, Daniel Rodriguez, Stefano Russo
arXiv AI
Sep 21

Offline Multimodal Large Language Models for Decision Support in Air Operations

The paper explores offline large language models as decision‑support tools for air operations, where analysts must work with doctrine and imagery under limited connectivity. It presents a modular retrieval‑augmented architecture that handles both text and image inputs from technical manuals and can operate without Internet access. A pilot study with four Brazilian Air Force image analysts shows the system matches human performance on doctrinal knowledge tests while completing the task in a fraction of the time, and it highlights the high cognitive load of manual target reporting.

By Joao P. A. Dantas, Jelton A. Cunha, Gabriel Dietzsch
arXiv AI
Jul 22

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

arXiv:2409. 07314v3 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows.

By Praveenkumar Kanithi, Cl\'ement Christophe, Marco AF Pimentel, Tathagata Raha, Prateek Munjal, Nada Saadi, Hamza A Javed, Svetlana Maslenkova, Nasir Hayat, Ronnie Rajan, Shadab Khan