arXiv Machine Learning

M-CTX: Exact and Scalable Spatial Context Retrieval for Trajectory Analytics

arXiv:2606. 15244v1 Announce Type: new Abstract: Modern trajectory predictors increasingly condition on external spatial context, such as map geometry, signed distance fields (SDFs), and nearby moving agents.

arXiv AI
Sep 7

Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

The paper introduces Linguistic Trajectory Encoding (LTE), a hybrid representation that compresses dynamic object motion histories using natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location while preserving accuracy with geometric waypoints and linguistic descriptions. Evaluated on the newly constructed Spatial Memory Benchmark (SMB) from EgoLife multi‑day recordings, LTE achieves 45.3 % success in semantic trajectory retrieval and 48.7 % in long‑horizon object retrieval, outperforming prior structured‑memory and VLM baselines, and compresses trajectories 8.7×–26.1× with sub‑second query latency on 24‑hour video.

By Tianyidan Xie, Shenyi Wang, Qiang Tang, Mingjie Wang, Zhicheng Qiu, Xuanfu Li, Zhan Xu, Jian Yang, Lanjun Wang, Zili Yi
arXiv Machine Learning
Jul 1

Capturing Context-Aware Route Choice Semantics for Trajectory Representation Learning

arXiv:2510. 14819v3 Announce Type: replace-cross Abstract: Trajectory representation learning (TRL) aims to encode raw trajectory data into low-dimensional embeddings for downstream tasks such as travel time estimation, mobility prediction, and trajectory similarity analysis.

By Ji Cao, Yu Wang, Tongya Zheng, Jie Song, Qinghong Guo, Zujie Ren, Canghong Jin, Gang Chen, Mingli Song
arXiv Machine Learning
Sep 11

M3-Former: Multimodal Transformer with Mixture-of-Experts for Long-Term Vessel Trajectory Prediction

M3-Former is a multimodal transformer framework that uses large language models to encode vessel static attributes and navigational intent as semantic priors for long‑term trajectory prediction. It builds a unified multimodal representation space, aligns static semantic information with dynamic trajectory features via self‑attention, and employs a dual‑granularity Mixture‑of‑Experts architecture to capture both global route planning and fine‑grained maneuvering behaviors. A Steering‑Weighted Cross‑Entropy loss further improves accuracy on sparse turning events, and experiments on a Danish AIS dataset show consistent improvements over state‑of‑the‑art baselines, reducing ADE and FDE by up to 5.1% in 4‑hour predictions.

By Wenzhe Jin, Haina Tang
arXiv AI
Jun 9

Mobility-Embedded POIs: Learning What A Place Is and How It Is Used from Human Movement

arXiv:2601. 21149v3 Announce Type: replace-cross Abstract: Recent progress in geospatial foundation models highlights the importance of learning general-purpose representations for real-world locations, particularly points-of-interest (POIs) where human activity concentrates.

By Maria Despoina Siampou, Shushman Choudhury, Shang-Ling Hsu, Neha Arora, Cyrus Shahabi
arXiv AI
Jul 20

GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

arXiv:2607. 15768v1 Announce Type: cross Abstract: Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space.

By Yujie Li, Jiancheng Pan, Zhiwei Wei, Jiuniu Wang, Mugen Peng, Wenjia Xu
arXiv Computer Vision
Sep 25

Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation

The paper investigates whether progress on spatial reasoning benchmarks translates into improved navigation performance. It finds a gap between benchmark-oriented spatial specialization and real navigation, and proposes aligning spatial supervision with navigation goals, phases, and decision learning. The authors introduce Spatial-Nav-100K, a two-stage fine-tuning approach, and Spatial-NPD, a teacher–student framework that eliminates explicit spatial reasoning at inference, achieving state‑of‑the‑art success rates on HM3D and MP3D with efficient GPU usage.

By Xun Huang, Shijia Zhao, Rongsheng Qu, Jiayuan Li, Xin Lu, Weixin Li, Chenglu Wen, Cheng Wang