arXiv Machine Learning

V2TATC: A Joint Voice-Trajectory Embedding Framework and Dataset for Air Traffic Controller Situational Awareness

arXiv Machine Learning
Jun 26

Learning to Explain Air Traffic Situation

arXiv:2502. 10764v4 Announce Type: replace Abstract: Understanding how air traffic controllers construct a mental 'picture' of complex air traffic situations is crucial but remains a challenge due to the inherently intricate, high-dimensional interactions between aircraft, pilots, and controllers.

By Hong-ah Chai, Seokbin Yoon, Keumjin Lee
arXiv AI
Aug 19

Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach

The paper introduces FlightLLM, a prior-guided semantic approach that uses large language models to explain flight safety events. It tackles challenges such as modal inconsistency, limited classification ability, and scarce domain data by combining feature engineering, semantic discretization, a CatBoost statistical expert, contrastive few-shot learning, and structured prompts. Evaluated on 704 real‑world A320 flights, FlightLLM achieves competitive classification and produces clear, aviation‑specific explanations for hard landing events.

By Lu Xu, Xu Li, Linjiang Zheng, Fan Li, Riquan Zhang, Jiaxing Shang
arXiv AI
Jun 18

ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots

arXiv:2606. 18319v1 Announce Type: cross Abstract: Air Traffic Control Operators (ATCOs) are vital in ensuring the safe, orderly, and efficient flow of air traffic, yet training capacity is constrained by reliance on specialized human trainers known as simpilots, who must role-play both pilots and ATCOs in a simulated airspace.

By Ethan Chew, Enjia Wu, Iruss Eng Wei Yeow, Ian Weiqin Lim, Ranen Sim, Brandon Koh Ziheng, Kaleb Nim, Caden Toh Jun Yi, Wei Dong Soin, Darius Kai Keat Koh, Galen King Yu Tay, Prannaya Gupta, Jonathan Ee Fang Koong, Yong Zhi Lim
arXiv Machine Learning
Jul 1

Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning

arXiv:2510. 14819v3 Announce Type: replace-cross Abstract: Trajectory representation learning (TRL) aims to encode raw trajectory data into low-dimensional embeddings for downstream tasks such as travel time estimation, mobility prediction, and trajectory similarity analysis.

By Ji Cao, Yu Wang, Tongya Zheng, Jie Song, Qinghong Guo, Zujie Ren, Canghong Jin, Gang Chen, Mingli Song
arXiv AI
Jun 18

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.

By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen
arXiv Machine Learning
6d ago

M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

M3-Former is a multimodal transformer framework that uses large language models to encode vessel static attributes and navigational intent as semantic priors for long‑term trajectory prediction. It builds a unified multimodal representation space, aligns static semantic information with dynamic trajectory features via self‑attention, and employs a dual‑granularity Mixture‑of‑Experts architecture to capture both global route planning and fine‑grained maneuvering behaviors. A Steering‑Weighted Cross‑Entropy loss further improves accuracy on sparse turning events, and experiments on a Danish AIS dataset show consistent improvements over state‑of‑the‑art baselines, reducing ADE and FDE by up to 5.1% in 4‑hour predictions.

By Wenzhe Jin, Haina Tang
arXiv AI
Jun 2

MOSS-Audio Technical Report

arXiv:2606. 01802v1 Announce Type: cross Abstract: MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped transcription, and audio-grounded reasoning.

By Chen Yang, Chufan Yu, Hanfu Chen, Jie Zhu, Jingqi Chen, Ke Chen, Wenxuan Wang, Yang Wang, Yaozhou Jiang, Yi Jiang, Zhengyuan Lin, Ziqi Chen, Zhaoye Fei, Chenghao Liu, Jun Zhan, Kang Yu, Kexin Huang, Mingshu Chen, Qinyuan Cheng, Ruixiao Li, Shimin Li, Songlin Wang, Yang Gao, Yiyang Zhang, Xipeng Qiu