AlphaEarth Foundations helps map our planet in unprecedented detail
New AI model integrates petabytes of Earth observation data to generate a unified data representation that revolutionizes global mapping and monitoring
New AI model integrates petabytes of Earth observation data to generate a unified data representation that revolutionizes global mapping and monitoring
arXiv:2603. 02142v2 Announce Type: replace-cross Abstract: Scaling laws assume larger models trained on more data consistently outperform smaller ones -- an assumption that drives model selection in computer vision but remains untested in resource-constrained Earth observation (EO).
The paper presents a dual‑encoder Transformer model for estimating Planetary Boundary Layer Height (PBLH) from satellite radiances, addressing challenges of multimodal, spatially incomplete data. It benchmarks eight different approaches, analyzes model reliance via grouped Shapley decomposition, and demonstrates that the proposed architecture achieves a mean absolute error of 155.8 m on a global test set, outperforming all baselines. On out‑of‑distribution data from the TEAMx campaign, the model attains 165.3 m MAE, better than a pixel‑wise baseline trained on the same data.
Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover narrow category vocabularies or limited query forms.
arXiv:2608. 04792v1 Announce Type: new Abstract: Accurate estimation of Above-Ground Biomass (AGB) from satellite imagery is essential for the large-scale monitoring of carbon stocks, yet it remains a challenging regression task at global scale.
arXiv:2605.24038v3 Announce Type: replace-cross Abstract: Aurora visibility at a given location requires two physically distinct conditions to hold at once: aurora occurring overhead, governed by sol...
The study audits the AION-1 foundation model, a 39‑modality transformer trained on over 200 million astronomical objects, and finds that its reliance on a survey detection channel—specifically the segmentation map—introduces a severe systematic bias. By keeping image tokens unchanged and editing only the segmentation map, all model outputs (flux, size, ellipticity, redshift) shift by factors of 110–4400 compared to a placebo, revealing that the model’s predictions are driven more by detection gating than by the actual light distribution. This bias propagates into cosmological analyses, shifting tomographic mean redshifts by a median 0.71 × the LSST DESC requirement and exceeding it in multiple assignments, while removing the detection channel eliminates the effect without measurable cost. whyItMatters":"The bias in the detection channel directly inflates errors in key astronomical measurements, potentially compromising the precision of cosmological studies that rely on accurate redshift estimates."
arXiv:2607. 27217v1 Announce Type: cross Abstract: Forest aboveground biomass (AGB) is a critical indicator of ecosystem productivity and terrestrial carbon storage, yet regional carbon monitoring remains constrained by the sparse spatial and temporal availability of field inventories and airborne structural measurements.
The study evaluates how the length of observation windows affects the performance of Tessera embeddings for land‑use/land‑cover mapping. By freezing the encoder and recomputing embeddings from a full year down to a single day, the authors benchmark linear probes and UNet heads on LUCAS, DynamicEarthNet, and PASTIS‑R datasets. Results show that embeddings are highly task‑dependent: for phenology‑driven classes (PASTIS‑R) they outperform from‑scratch models by ~46%, while for temporally stable classes (DynamicEarthNet, LUCAS) they match only with full supervision, yet remain more label‑efficient across all datasets.
arXiv:2607. 17037v1 Announce Type: new Abstract: High-resolution atmospheric data are required to resolve mesoscale and localized meteorological structures, however such datasets remain limited in many regions of the world.
arXiv:2607. 01584v1 Announce Type: new Abstract: Large language models have recently been explored for scientific hypothesis generation, but most prior work relies on unstructured literature and free-form textual claims.