Domain-Informed Multi-View Self-Distillation for Astronomical Light-Curve Representation Learning with JEPA
arXiv:2606. 28446v1 Announce Type: cross Abstract: Light curves describe temporal variations in the brightness of celestial objects.
arXiv:2504. 06176v4 Announce Type: replace-cross Abstract: Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasingly being applied to specialised domains.
arXiv:2606. 28446v1 Announce Type: cross Abstract: Light curves describe temporal variations in the brightness of celestial objects.
arXiv:2603. 28963v2 Announce Type: replace-cross Abstract: Simulation with realistic traffic agents is essential for validating autonomous driving systems.
The paper introduces TITAnD, a Trajectory Image Transformer that converts dense and sparse GPS trajectories into a Hyperspectral Trajectory Image (HTI) and applies vision-based classification and segmentation for anomaly detection. It employs a Cyclic Factorized Transformer (CFT) that splits attention along within-day and across-day axes, drastically reducing computational cost and enabling multi-month analysis. Empirical results show TITAnD outperforms existing sparse and dense benchmarks, achieving higher AUC-PR and faster inference than comparable Transformers.
arXiv:2609.01551v1 Announce Type: new Abstract: Self-supervised video foundation models learn rich spatiotemporal representations, yet it remains unclear what visual concepts these representations en...
The technical report introduces Hyperspectral Image Models, a modular framework that unifies 55 deep‑learning models across six paradigms for hyperspectral remote sensing. It standardizes tensor conventions, evaluation protocols, and dataset handling, integrating 24 benchmark scenes from various sensors and providing tools to avoid train‑test overlap. Experiments across 1,320 model‑scene combinations show that scene difficulty outweighs architecture, with no single paradigm dominating and small models achieving performance comparable to much larger ones.
The paper demonstrates that a self‑supervised Vision Transformer (ViT) pretrained on a fast, low‑cost semi‑numerical simulator can produce data summaries that transfer across different simulators without retraining. In 21cm cosmology, the ViT—named SKATR—pretrained on 67,000 21cmFAST lightcones is applied unchanged to hydrodynamical Loreli II lightcones, enabling accurate inference of five astrophysical parameters with fewer radiative‑transfer simulations than a fully‑supervised baseline. SKATR remains accurate, informative, and calibrated even under realistic SKA antenna array noise, outperforming supervised models retrained on noisy data.
arXiv:2607. 00228v1 Announce Type: cross Abstract: Modern time-domain surveys such as the Zwicky Transient Facility (ZTF) generate hundreds of thousands of alerts each night, making real-time decisions for follow-up observations a central challenge in time-domain astronomy.
STRADAViT is a self‑supervised continued‑pretraining framework that adapts Vision Transformer (ViT) backbones for radio‑astronomy image analysis. It curates mixed‑survey data, generates radio‑astronomy‑aware training views, and initializes encoders with ViT‑MAE, optionally adding register tokens. Evaluations on three morphology benchmarks (MiraBest, LoTSS DR2, and Radio Galaxy Zoo) show that a register‑based two‑stage checkpoint improves linear‑probe Macro‑F1 scores over the ViT‑MAE baseline and enhances fine‑tuning on MiraBest and RGZ DR1, though performance on LoTSS DR2 fine‑tuning declines; these differences are statistically significant.
arXiv:2601. 18823v4 Announce Type: replace Abstract: Variational autoencoders (VAE) encode data into lower-dimensional latent vectors before decoding those vectors back to data.
arXiv:2608. 19419v1 Announce Type: cross Abstract: Microlensing can reveal populations of faint compact objects that are otherwise difficult to detect.
MC-DeTra is a reimplementation of the DeTra model that jointly performs object detection and socially-aware trajectory forecasting in bird's-eye-view images. It introduces motion-consistency mechanisms that add supervision from each actor’s past motion, surrounding traffic occupancy, and a consistency constraint aligning predicted heading with motion direction. The added losses are train‑only and inference‑safe, improving dynamic trajectory forecasting on the Waymo Open Dataset while maintaining or enhancing detection accuracy.
arXiv:2606. 09936v1 Announce Type: cross Abstract: World models are now built on substantially different computational substrates.