arXiv Machine Learning By Pablo Mercader-Perez, Carolina Cuesta-Lazaro, Daniel Muthukrishna, Jeroen Audenaert, V. Ashley Villar, David W. Hogg, Marc Huertas-Company, William T. Freeman

Learning What's Real: Disentangling Signal and Measurement Artifacts in Multi-Sensor Data, with Applications to Astrophysics

Read the original on arXiv Machine Learning →

arXiv:2604. 09787v2 Announce Type: replace-cross Abstract: Data collected from the physical world is always a combination of multiple sources: an underlying signal from the physical process of interest and a signal from measurement-dependent artifacts from the sensor or instrument.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 10

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

arXiv:2606. 11190v1 Announce Type: new Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality.

By Ilay Kamai, Hugues Van Assel, Aviv Regev, Hagai B. Perets, Randall Balestriero
Hugging Face Trending Papers
Jun 9

When to Align, When to Predict: A Phase Diagram for Multimodal Learning

Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all -- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions.

arXiv Machine Learning
Aug 28

Cross-simulator transfer with foundation model summaries: Towards robust SKA-era reionization inference

The paper demonstrates that a self‑supervised Vision Transformer (ViT) pretrained on a fast, low‑cost semi‑numerical simulator can produce data summaries that transfer across different simulators without retraining. In 21cm cosmology, the ViT—named SKATR—pretrained on 67,000 21cmFAST lightcones is applied unchanged to hydrodynamical Loreli II lightcones, enabling accurate inference of five astrophysical parameters with fewer radiative‑transfer simulations than a fully‑supervised baseline. SKATR remains accurate, informative, and calibrated even under realistic SKA antenna array noise, outperforming supervised models retrained on noisy data.

By Yannic Pietschke, Caroline Heneka, Ayodele Ore, Romain Meriot
arXiv Machine Learning
Jun 24

Efficient reduction of stellar contamination and noise in planetary transmission spectra using neural networks

arXiv:2602. 10330v3 Announce Type: replace-cross Abstract: Context: The characterization of exoplanetary atmospheres has been transformed by the James Webb Space Telescope (JWST), whose infrared sensitivity enables transmission spectroscopy at unprecedented precision.

By David S. Duque-Casta\~no, Lauren Flor-Torres, Jorge I. Zuluaga