arXiv Machine Learning

The Domain Is a Residue: Adapting Self-Supervised Features, Not Generators

arXiv Machine Learning
Jun 9

TriHead-GAN: A Generative Adversarial Network with Triple-Head Discriminator for Carbon Emission Time Series Generation

arXiv:2606. 07569v1 Announce Type: new Abstract: Accurate carbon emission monitoring is critical for climate policy and emerging regulatory mechanisms such as the EU Carbon Border Adjustment Mechanism, yet city-level high-frequency monitoring data remain extremely scarce, severely limiting data-hungry deep learning models.

By Zesen Wang, Lijuan Lan, Yonggang Li, Chunhua Yang
arXiv AI
1d ago

DAGS: Disentangled Appearance-and-Geometry Steering of a Frozen Image DiT for Temporally Stabilized Generative Rendering

DAGS introduces a lightweight, attention‑free conditioning scheme that disentangles appearance and geometry for a frozen image diffusion transformer (DiT), enabling high‑fidelity, temporally stable renders with independent control. Two small convolutional encoders generate per‑frame conditioning features, which are injected as learned residuals into the image tokens, avoiding the quadratic cost of attention. Coupled with a recurrent lighting stabilizer and a training‑free temporal guidance term, DAGS transforms a per‑frame image model into a streaming renderer that outperforms real‑time denoisers and diffusion renderers in PSNR and temporal stability while requiring far less compute than path tracing.

By Karthik Mohan Kumar, Damian Andrysiak, Pedro Antonio Pena, Kunal Tyagi, Rama Harihara
arXiv Computer Vision
Sep 23

MirrorDistill: Illumination-Aware Latent Distillation for Efficient Low-Light Restoration

MirrorDistill introduces an illumination‑aware latent distillation framework for low‑light image enhancement. It trains a lightweight student encoder‑decoder by aligning its intermediate features with clean‑domain targets generated by a teacher decoder, using feature mirroring and illumination‑aware weighting to emphasize underexposed regions. The method achieves state‑of‑the‑art performance on the LOL‑v2‑Real benchmark while maintaining the lowest computational complexity, and the code is released as open source.

By Farida Mohsen, Tala Zaim, Nurul Izni Rusli, Ali Al-Zawqari, Ali Safa, Samir Brahim Belhaouari
arXiv Computer Vision
Sep 4

ProgResViT: Progressive Resolution and Width for Adaptive Vision Transformers

ProgResViT is an input‑adaptive Vision Transformer that processes images progressively across multiple rounds, starting with a low‑resolution image and a narrow subnetwork and refining the prediction with higher resolution and a wider subnetwork if needed. The method introduces Progress‑Conditioned Soft Gating (PSG) to share a single backbone across rounds while conditioning token fusion and layer outputs on the current round, block, and input resolution. Experiments on DeiT show improved accuracy‑compute trade‑offs compared to adaptive‑width, adaptive‑depth, and dynamic‑token baselines, and the design also benefits self‑supervised DINO representations and downstream semantic segmentation.

By Ali Hojjat, Janek Haberer, Olaf Landsiedel
arXiv Machine Learning
Jun 8

CF-JEPA: Mask-free forward prediction with asymmetric encoder utilization for time-series representation learning

arXiv:2606. 07031v1 Announce Type: new Abstract: Self-supervised learning (SSL) for time-series representation learning is dominated by two paradigms: contrastive methods, which face challenges in constructing positive or negative pairs, and masking-based methods, which disrupt the temporal continuity of time-series signals.

By Jaehoon Lee, Sunghyun Sim
arXiv Computer Vision
6d ago

TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

TT-VidT is a video pretraining method that decouples the temporal axis by combining a per‑frame ViT-B/16 spatial encoder with a compact Temporal Transfer Layer trained via Diff Compression. The authors conduct a systematic 24‑configuration study to isolate architecture, objective, and decoder effects, showing that the full TT-VidT design yields the strongest motion‑sensitive representations. In downstream fine‑tuning, TT‑VidT outperforms state‑of‑the‑art baselines on Jester, Something‑Something V2, ARID, and Diving48 while using significantly fewer encoder FLOPs.

By Shih-Ying Yeh, Daniel Z. Kaplan, Xuehai Wang, Fu-En Yang, Min-Hung Chen, Shang-Hong Lai
arXiv Machine Learning
Sep 16

IRENE: A Convolutional GRU Ensemble Model for Radar Precipitation Nowcasting over Italy

IRENE is a deep learning model that provides probabilistic short‑range precipitation nowcasts over Italy at 1 km spatial and 5‑minute temporal resolution. It uses an encoder–forecaster architecture built on multi‑scale Convolutional Gated Recurrent Units (ConvGRUs) and is trained on national radar composites, with an importance‑sampling scheme and the almost‑fair Continuous Ranked Probability Score as its primary loss. Three training variants—standard, adversarial (IRENE‑GAN), and spectrally constrained (IRENE‑GAN‑RAPSD)—outperform benchmark methods STEPS and DGMR in probabilistic skill, though the advantage in mean absolute error is limited to the first 90 minutes.

By Alessandro Camilletti, Gabriele Franch, Elena Tomasi, Marco Cristoforetti
arXiv Machine Learning
Sep 2

GenONet: A Generative operator Network for High-Resolution Precipitation Nowcasting

GenONet introduces a Spatio-Temporal U-DeepONet architecture that serves as a generator in a GAN framework for high‑resolution precipitation nowcasting up to three hours ahead. By learning continuous‑time precipitation dynamics with a Deep Operator Network and enforcing physics through a moisture‑conservation loss, the model produces sharp, physically consistent forecasts that outperform baselines, especially for high‑intensity events and longer lead times. Ablation studies confirm the added value of the physics‑informed regularizer and the synergy of operator learning with adversarial training.

By Mohammad Kian Golkar, Luciano Alves de Oliveira, Mohammad Khanjani