Scalable GANs with Transformers
arXiv:2509. 24935v3 Announce Type: replace-cross Abstract: Scalability has driven recent advances in generative modeling, yet its principles remain underexplored for adversarial learning.
arXiv:2404. 06294v2 Announce Type: replace-cross Abstract: Super-Resolution (SR) is a time-hallowed image processing problem that aims to improve the quality of a Low-Resolution (LR) sample up to the standard of its High-Resolution (HR) counterpart.
arXiv:2509. 24935v3 Announce Type: replace-cross Abstract: Scalability has driven recent advances in generative modeling, yet its principles remain underexplored for adversarial learning.
arXiv:2608. 09133v1 Announce Type: cross Abstract: Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to the LR observation remains challenging.
arXiv:2601.17723v3 Announce Type: replace Abstract: Implicit neural representation (INR) has become the standard approach for arbitrary-scale image super-resolution (ASSR). However, no systematic emp...
arXiv:2608.22272v1 Announce Type: cross Abstract: Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration but requir...
arXiv:2609.15120v1 Announce Type: new Abstract: Benefiting from the powerful generative priors of diffusion models, diffusion-based real-world image super-resolution (Real-ISR) methods have demonstra...
UnCapsTSR is an unsupervised transformer-based GAN framework designed to enhance the spatial resolution of low‑resolution wireless capsule endoscopy (WCE) images. It eliminates the need for explicit degradation modeling or paired LR‑HR data by using a Bilateral Total Variation loss to preserve spatial continuity. The authors introduce a new Kvasir Capsule dataset for training, validate generalizability on KID and GIANA datasets, and propose the Endoscopy Quality Metric (EndoQM) as a non‑reference evaluation tool, reporting 40–80% improvement in EndoQM over state‑of‑the‑art unsupervised methods.
arXiv:2607. 15711v1 Announce Type: cross Abstract: Diffusion-based methods have achieved impressive performance in real-world image super-resolution (Real-ISR) by leveraging large pre-trained stable diffusion (SD) models as powerful generative priors.
arXiv:2606. 30528v1 Announce Type: cross Abstract: Current generative models, including GANs and diffusion models, have reached an outstanding level of photorealism, posing significant risks to privacy and security.
The paper investigates whether the stable GAN architecture R3GAN can improve time‑series imputation when adapted to 1‑D temporal data. Using a coarse‑to‑fine refinement framework and a frequency‑domain discriminator, the authors evaluate 14 saved configurations across three datasets and find a negative result: most configurations either show negligible improvement or degrade performance compared to baseline methods. The study highlights that the usual argument—GANs optimize distributional objectives rather than point‑wise ones—does not fully explain the lack of benefit, and it poses an open problem regarding why a learned discriminator fails to provide useful refinement gradients while diffusion denoisers succeed, offering practical guidance on when adversarial refinement may be worthwhile.
Pixel diffusion models generate RGB images directly but tend to miss fine‑scale natural‑image statistics. The authors introduce an adversarial post‑training step that adds an adversarial loss to the model’s output at non‑high‑noise timesteps, without changing the architecture or sampling procedure. This approach improves distribution fidelity, coverage, prompt alignment, and perceptual quality across two pixel backbones, and restores missing high‑frequency spectral power while avoiding memorization or mode dropping.
Diffusion-based generative models have achieved remarkable success in real-world image super-resolution (SR). With tiled diffusion techniques, these models can produce high-resolution images that exceed their native-supported resolution.
EmbeddGAN introduces a new GAN framework that replaces the traditional discriminator with an embedding network trained to maximize statistical dependence between embeddings and real/fake labels using Gini distance correlation (gCor). The generator simultaneously minimizes this dependence, encouraging real and generated samples to become indistinguishable in the learned low‑dimensional embedding space. Experiments on MNIST, CIFAR‑10, and CelebA show competitive performance and notably more stable training dynamics compared to established baselines.