arXiv:2606. 07569v1 Announce Type: new Abstract: Accurate carbon emission monitoring is critical for climate policy and emerging regulatory mechanisms such as the EU Carbon Border Adjustment Mechanism, yet city-level high-frequency monitoring data remain extremely scarce, severely limiting data-hungry deep learning models.
By Zesen Wang, Lijuan Lan, Yonggang Li, Chunhua Yang
arXiv:2606. 15553v1 Announce Type: cross Abstract: Representation Autoencoders (RAEs) have improved diffusion and flow models by semantically richer latent space owing to the strongly label-wise clustered DINO features in the pretrained encoders.
By Jiawei Zhang, Mengfei Xia, Gen Li, Yuantao Gu
arXiv:2610.00686v1 Announce Type: new
Abstract: Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene...
By Mikhail Dereviannykh, Vikram Voleti, Simon Donne, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss
DAGS introduces a lightweight, attention‑free conditioning scheme that disentangles appearance and geometry for a frozen image diffusion transformer (DiT), enabling high‑fidelity, temporally stable renders with independent control. Two small convolutional encoders generate per‑frame conditioning features, which are injected as learned residuals into the image tokens, avoiding the quadratic cost of attention. Coupled with a recurrent lighting stabilizer and a training‑free temporal guidance term, DAGS transforms a per‑frame image model into a streaming renderer that outperforms real‑time denoisers and diffusion renderers in PSNR and temporal stability while requiring far less compute than path tracing.
By Karthik Mohan Kumar, Damian Andrysiak, Pedro Antonio Pena, Kunal Tyagi, Rama Harihara
MirrorDistill introduces an illumination‑aware latent distillation framework for low‑light image enhancement. It trains a lightweight student encoder‑decoder by aligning its intermediate features with clean‑domain targets generated by a teacher decoder, using feature mirroring and illumination‑aware weighting to emphasize underexposed regions. The method achieves state‑of‑the‑art performance on the LOL‑v2‑Real benchmark while maintaining the lowest computational complexity, and the code is released as open source.
By Farida Mohsen, Tala Zaim, Nurul Izni Rusli, Ali Al-Zawqari, Ali Safa, Samir Brahim Belhaouari
ProgResViT is an input‑adaptive Vision Transformer that processes images progressively across multiple rounds, starting with a low‑resolution image and a narrow subnetwork and refining the prediction with higher resolution and a wider subnetwork if needed. The method introduces Progress‑Conditioned Soft Gating (PSG) to share a single backbone across rounds while conditioning token fusion and layer outputs on the current round, block, and input resolution. Experiments on DeiT show improved accuracy‑compute trade‑offs compared to adaptive‑width, adaptive‑depth, and dynamic‑token baselines, and the design also benefits self‑supervised DINO representations and downstream semantic segmentation.
By Ali Hojjat, Janek Haberer, Olaf Landsiedel
arXiv:2603. 13326v2 Announce Type: replace-cross Abstract: Multimodal Transformers often produce predictions without clarifying how different modalities jointly support a decision.
By Yeji Kim, Housam Khalifa Bashier Babiker, Mi-Young Kim, Randy Goebel
arXiv:2606. 07031v1 Announce Type: new Abstract: Self-supervised learning (SSL) for time-series representation learning is dominated by two paradigms: contrastive methods, which face challenges in constructing positive or negative pairs, and masking-based methods, which disrupt the temporal continuity of time-series signals.
By Jaehoon Lee, Sunghyun Sim
arXiv:2605. 18324v2 Announce Type: replace-cross Abstract: Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders.
By Jaskirat Singh, Boyang Zheng, Zongze Wu, Richard Zhang, Eli Shechtman, Saining Xie
TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.
By Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
IRENE is a deep learning model that provides probabilistic short‑range precipitation nowcasts over Italy at 1 km spatial and 5‑minute temporal resolution. It uses an encoder–forecaster architecture built on multi‑scale Convolutional Gated Recurrent Units (ConvGRUs) and is trained on national radar composites, with an importance‑sampling scheme and the almost‑fair Continuous Ranked Probability Score as its primary loss. Three training variants—standard, adversarial (IRENE‑GAN), and spectrally constrained (IRENE‑GAN‑RAPSD)—outperform benchmark methods STEPS and DGMR in probabilistic skill, though the advantage in mean absolute error is limited to the first 90 minutes.
By Alessandro Camilletti, Gabriele Franch, Elena Tomasi, Marco Cristoforetti
GenONet introduces a Spatio-Temporal U-DeepONet architecture that serves as a generator in a GAN framework for high‑resolution precipitation nowcasting up to three hours ahead. By learning continuous‑time precipitation dynamics with a Deep Operator Network and enforcing physics through a moisture‑conservation loss, the model produces sharp, physically consistent forecasts that outperform baselines, especially for high‑intensity events and longer lead times. Ablation studies confirm the added value of the physics‑informed regularizer and the synergy of operator learning with adversarial training.
By Mohammad Kian Golkar, Luciano Alves de Oliveira, Mohammad Khanjani