arXiv:2510. 06200v4 Announce Type: replace-cross Abstract: Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling.
By Weijian Li, Hong-Yu Chen, Nabeel Rehemtulla, Ved G. Shah, Dongho Kim, Dennis Wu, Qinjie Lin, Adam A. Miller, Han Liu
arXiv:2606. 28446v1 Announce Type: cross Abstract: Light curves describe temporal variations in the brightness of celestial objects.
By Yicheng Rui
arXiv:2605. 27527v2 Announce Type: replace-cross Abstract: Astrophysical observations from Earth are subject to weather, environmental, and scientific constraints that lead to sparse, irregular light curves.
By Siddharth Chaini, Federica B. Bianco, Ashish Mahabal
arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.
By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero
arXiv:2606. 05103v1 Announce Type: new Abstract: The Nancy Grace Roman Space Telescope (Roman), set for launch as early as September 2026, will conduct wide-field infrared imaging surveys with unprecedented spatial resolution and cadence, enabling the discovery of millions of astronomical transients.
By Karan Gandhi, Ashish A. Mahabal, Jacob E. Jencson, Russ R. Laher, Ben Rusholme, Lin Yan, Ryan M. Lau, Schuyler D. Van Dyk, Mansi M. Kasliwal
STRADAViT is a self‑supervised continued‑pretraining framework that adapts Vision Transformer (ViT) backbones for radio‑astronomy image analysis. It curates mixed‑survey data, generates radio‑astronomy‑aware training views, and initializes encoders with ViT‑MAE, optionally adding register tokens. Evaluations on three morphology benchmarks (MiraBest, LoTSS DR2, and Radio Galaxy Zoo) show that a register‑based two‑stage checkpoint improves linear‑probe Macro‑F1 scores over the ViT‑MAE baseline and enhances fine‑tuning on MiraBest and RGZ DR1, though performance on LoTSS DR2 fine‑tuning declines; these differences are statistically significant.
By Andrea DeMarco, Ian Fenech Conti, Hayley Camilleri, Ardiana Bushi, Simone Riggi
Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions.
The study audits the AION-1 foundation model, a 39‑modality transformer trained on over 200 million astronomical objects, and finds that its reliance on a survey detection channel—specifically the segmentation map—introduces a severe systematic bias. By keeping image tokens unchanged and editing only the segmentation map, all model outputs (flux, size, ellipticity, redshift) shift by factors of 110–4400 compared to a placebo, revealing that the model’s predictions are driven more by detection gating than by the actual light distribution. This bias propagates into cosmological analyses, shifting tomographic mean redshifts by a median 0.71 × the LSST DESC requirement and exceeding it in multiple assignments, while removing the detection channel eliminates the effect without measurable cost.
whyItMatters":"The bias in the detection channel directly inflates errors in key astronomical measurements, potentially compromising the precision of cosmological studies that rely on accurate redshift estimates."
By Ihor Kendiukhov
arXiv:2609.15087v1 Announce Type: cross
Abstract: Most time series forecasting benchmarks remain numerical-centric and provide limited support for evaluating contextual information that shapes real-w...
By Peng Chen, Zhihao Zhuang, Hongzhou Chen, Junhao Huang, Aiping Yang, Mengsen Wu, Yiding Liu, Xilin Dai, Zewei Dong
arXiv:2504. 06176v4 Announce Type: replace-cross Abstract: Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasingly being applied to specialised domains.
By Ian Groves, Andrew Campbell, James Fernandes, Diego Ram\'irez Rodr\'iguez, Paul Murray, Massimiliano Vasile, Victoria Nockles
arXiv:2608. 00012v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated.
By Fengxiang Wang, Qiuyang Yu, Yueying Li, Mingshuo Chen, Chengchi Fei, Kaiyi Xu, Lixin Gu, Wangxu Wei, Junchao Gong, Lipeng Ma, Jiong Wang, Fenghua Ling, Wenlong Zhang, Xue Yang, Wenjing Yang, Ben Fei, Long Lan
arXiv:2605.24038v3 Announce Type: replace-cross
Abstract: Aurora visibility at a given location requires two physically distinct conditions to hold at once: aurora occurring overhead, governed by sol...
By Zongyuan Ge, Chenwaner Zhang, Haoyang Li, Hantai Zhang, Wei Zhou, Wenxin Gu, Zhaoming Wang