arXiv AI

A Self-Supervised Framework for Space Object Behaviour Characterisation

arXiv:2504. 06176v4 Announce Type: replace-cross Abstract: Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasingly being applied to specialised domains.

arXiv Machine Learning
Sep 10

Hyperspectral Trajectory Image for Multi-Month Trajectory Anomaly Detection

The paper introduces TITAnD, a Trajectory Image Transformer that converts dense and sparse GPS trajectories into a Hyperspectral Trajectory Image (HTI) and applies vision-based classification and segmentation for anomaly detection. It employs a Cyclic Factorized Transformer (CFT) that splits attention along within-day and across-day axes, drastically reducing computational cost and enabling multi-month analysis. Empirical results show TITAnD outperforms existing sparse and dense benchmarks, achieving higher AUC-PR and faster inference than comparable Transformers.

By Md Awsafur Rahman, Chandrakanth Gudavalli, Hardik Prajapati, B. S. Manjunath
arXiv Computer Vision
3d ago

Hyperspectral Image Models: Technical Report

The technical report introduces Hyperspectral Image Models, a modular framework that unifies 55 deep‑learning models across six paradigms for hyperspectral remote sensing. It standardizes tensor conventions, evaluation protocols, and dataset handling, integrating 24 benchmark scenes from various sensors and providing tools to avoid train‑test overlap. Experiments across 1,320 model‑scene combinations show that scene difficulty outweighs architecture, with no single paradigm dominating and small models achieving performance comparable to much larger ones.

By Tanishq Rachamalla, Aryan Das, Srishti Kaushik, Swalpa Kumar Roy
arXiv Machine Learning
Aug 28

Cross-simulator transfer with foundation model summaries: Towards robust SKA-era reionization inference

The paper demonstrates that a self‑supervised Vision Transformer (ViT) pretrained on a fast, low‑cost semi‑numerical simulator can produce data summaries that transfer across different simulators without retraining. In 21cm cosmology, the ViT—named SKATR—pretrained on 67,000 21cmFAST lightcones is applied unchanged to hydrodynamical Loreli II lightcones, enabling accurate inference of five astrophysical parameters with fewer radiative‑transfer simulations than a fully‑supervised baseline. SKATR remains accurate, informative, and calibrated even under realistic SKA antenna array noise, outperforming supervised models retrained on noisy data.

By Yannic Pietschke, Caroline Heneka, Ayodele Ore, Romain Meriot
arXiv Machine Learning
Jul 2

Leveraging Multimodality for Real-Time Classification of Transients and Variables found by the Zwicky Transient Facility

arXiv:2607. 00228v1 Announce Type: cross Abstract: Modern time-domain surveys such as the Zwicky Transient Facility (ZTF) generate hundreds of thousands of alerts each night, making real-time decisions for follow-up observations a central challenge in time-domain astronomy.

By Ved G. Shah, Nabeel Rehemtulla, Adam A. Miller, Sushant Sharma Chaudhary, Michael W. Coughlin, Antoine Le Calloch, Matthew J. Graham, Joahan Castaneda Jaimes, Theophile Jegou du Laz, Ashish A. Mahabal, Frank J. Masci, Josiah Purdum, Reed Riddle, Jesper Sollerman, Anastasia Wei, Mansi M. Kasliwal
arXiv Computer Vision
Sep 17

STRADAViT: Self-Supervised Domain Adaptation of Vision Transformer Backbones for Radio Astronomy

STRADAViT is a self‑supervised continued‑pretraining framework that adapts Vision Transformer (ViT) backbones for radio‑astronomy image analysis. It curates mixed‑survey data, generates radio‑astronomy‑aware training views, and initializes encoders with ViT‑MAE, optionally adding register tokens. Evaluations on three morphology benchmarks (MiraBest, LoTSS DR2, and Radio Galaxy Zoo) show that a register‑based two‑stage checkpoint improves linear‑probe Macro‑F1 scores over the ViT‑MAE baseline and enhances fine‑tuning on MiraBest and RGZ DR1, though performance on LoTSS DR2 fine‑tuning declines; these differences are statistically significant.

By Andrea DeMarco, Ian Fenech Conti, Hayley Camilleri, Ardiana Bushi, Simone Riggi
arXiv Computer Vision
Sep 11

MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View Images

MC-DeTra is a reimplementation of the DeTra model that jointly performs object detection and socially-aware trajectory forecasting in bird's-eye-view images. It introduces motion-consistency mechanisms that add supervision from each actor’s past motion, surrounding traffic occupancy, and a consistency constraint aligning predicted heading with motion direction. The added losses are train‑only and inference‑safe, improving dynamic trajectory forecasting on the Waymo Open Dataset while maintaining or enhancing detection accuracy.

By Vladislav Diuzhev, Dmitry Yudin