SolarBench is an open global benchmark for image-based solar nowcasting that consolidates over six million sky and satellite images from 11 sites across a decade, paired with irradiance, PV output, and atmospheric data. The benchmark includes a toolbox for reproducible data access, processing, model development, and evaluation. Using SolarBench, the authors benchmark representative models, uncover a gap between average forecasting accuracy and the capture of rapid solar fluctuations, quantify predictability across cloud regimes, and demonstrate data‑efficient adaptation to new PV systems.
By Yuhao Nie, Stephen Campbell, Quentin Paletta, Liwenbo Zhang, Tao Jing, Samer Chaaraoui, Jonathan Giezendanner, Andea Scott, Tao Sun, Cong Feng, Max Aragon, Jacques Camier, Adam Jensen, Florian Kotthoff, Yuexing Yang, Yang Ming, Mengying Li, Stefanie Meilinger, Yupeng Wu, Adam Brandt, Sherrie Wang
arXiv:2603.03806v2 Announce Type: replace
Abstract: The state space model Mamba has recently emerged as a promising paradigm in computer vision, attracting considerable attention for its efficient ha...
By Hanpeng Liu, Zidan Wang, Shuoxi Zhang, Kaiyuan Gao, Kun He
arXiv:2606. 16996v1 Announce Type: cross Abstract: Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabulary semantic segmentation (OVSS) is inefficient: full-resolution decoding is typically run over the entire dataset vocabulary, whereas each image contains only a small active subset of classes.
By Tran Dinh Tien, Zhiqiang Shen
arXiv:2606.16996v2 Announce Type: replace-cross
Abstract: Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabu...
By Tran Dinh Tien, Zhiqiang Shen
SolarFlowRefiner is a refinement‑aware flow‑matching framework designed to downscale high‑resolution surface solar radiation (SSR) fields from coarse ERA5 radiative variables and satellite channels. It first uses a conditional FlowMatch generator to predict a normalized correction to an upsampled ERA5 baseline, then trains a refiner on prediction‑conditioned states between the generator’s output and the target residual, exposing the refiner to the generator’s structured errors. The refinement objective is backpropagated through the FlowMatch sampler, enabling joint optimization of generation and correction, and experiments on an ERA5–SolarCube benchmark demonstrate consistent improvements over standalone generation and post‑hoc refinement.
By Udbhav Srivastava, Antonita Racheal, Yiheng Chen, Runlong Yu, Xinyue Ye
arXiv:2607. 28627v1 Announce Type: cross Abstract: Long visual context poses a challenge for vision-language models: performance degrades as the number of distractors grows, and processing all tokens at once is computationally infeasible under GPU memory constraints.
By Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem
arXiv:2610.07758v1 Announce Type: cross
Abstract: Training-free token reduction accelerates vision transformers by removing redundant tokens across layers, recovering most of the original accuracy at...
By Hyeongheon Cha, Hyungjun Yoon, Sung-Ju Lee
The paper introduces SGMA, a structure‑guided masked autoencoding framework designed for ultra‑high‑resolution scientific images. SGMA combines a content‑adaptive quadtree tokenizer that reduces gigapixel images to a fixed‑length sequence with a structure‑conditioned masking process that focuses reconstruction on spatially informative regions. The method, enhanced by Damped Accumulation to stabilize multi‑scale signals, achieves superior performance over standard MAE baselines on electron microscopy, whole‑slide optical microscopy, and X‑ray CT datasets, delivering significant accuracy gains and up to a 24.8× inference speedup.
By Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas, Amir Koushyar Ziabari, Xiao Wang, Peng Chen, Tao Luo, Toshio Endo, Fumiyoshi Shoji, Kento Sato, Kentaro Uesugi, Takayuki Nonoyama, Ryuji Kiyama, Masahiro Yoshida, Masaru Tezuka, Tetsuya Ishikawa, Satoshi Matsuoka, Masaharu Munetomo, Mohamed Wahib
arXiv:2605.27696v3 Announce Type: replace
Abstract: Discrete visual tokenizers map images to ordered sequences of tokens, providing a natural representation for structural scene descriptions. Most us...
By Piotr Wyrwi\'nski, Kacper Dobek, Krzysztof Krawiec
arXiv:2608.23923v1 Announce Type: new
Abstract: Slicing-Aided Hyper Inference (SAHI) improves small object detection in high-resolution images but often spends substantial compute on background tiles...
By Rashid Riyadh, Abd Ullah Khan, Imad Gohar, Muzammil Behzad
arXiv:2603. 02142v2 Announce Type: replace-cross Abstract: Scaling laws assume larger models trained on more data consistently outperform smaller ones -- an assumption that drives model selection in computer vision but remains untested in resource-constrained Earth observation (EO).
By Kwame Mbobda-Kuate, Gabriel Kasmi
arXiv:2606. 27978v1 Announce Type: cross Abstract: Pixel-space continuous-token autoregressive (AR) generation directly models images as sequences of raw pixel patches, avoiding discrete tokenization or a separately pretrained tokenizer.
By Jiayi Xu, Di He, Guolin Ke