arXiv Computer Vision

Structured-Prior-Guided Diffusion Inpainting with Physical Consistency for Traffic Sign Augmentation

The paper introduces a structured‑prior‑guided diffusion inpainting framework that enhances traffic sign augmentation by incorporating semantic, appearance, and geometric priors through text prompts, vector templates, and ControlNet. It enforces physical consistency with colour and edge‑structure losses, achieving superior reconstruction fidelity and OCR accuracy compared to larger models. The synthetic data generated improves detection performance for rare traffic sign classes by up to 7.4× over a real‑data‑only baseline.

arXiv AI
Aug 13

Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment

arXiv:2608. 11537v1 Announce Type: cross Abstract: Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization.

By Weize Cai, Yongqi Dong, Zhida Shao, Zixin Fu
arXiv Computer Vision
Sep 2

Training-Free Inpainting Across Domains with a Frozen Text-to-Image Diffusion Model

The paper demonstrates that a frozen generic text‑to‑image diffusion model can perform conditional inpainting across three natural‑image domains without any inpainting‑specific training or dataset adaptation. The proposed Step‑PI method augments known‑region projection with boundary‑interior latent feedback, persistent PI state, and a predefined four‑field release schedule, improving all evaluated metrics over baseline training‑free approaches. Experiments on CelebA‑HQ, AFHQ, and Places2 show consistent gains, with Step‑PI outperforming LanPaint and PILOT on all five macro metrics.

By Zhenhuan Wang, Fengyi Yuan
arXiv Computer Vision
6d ago

The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.

By Soshun Kihara, Shunsuke Yasuki, Masato Taki
arXiv Machine Learning
Jul 14

BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation

arXiv:2604. 16514v5 Announce Type: replace-cross Abstract: Autoregressive vision-language models (VLMs) deliver strong multimodal capability, but their token-by-token decoding imposes a fundamental inference bottleneck.

By Baoyou Chen, Hanchen Xia, Peng Tu, Haojun Shi, Liwei Zhang, Yuxuan Yao, Weihao Yuan, Siyu Zhu
arXiv Computer Vision
Sep 1

PERSIST: Persistent-State Discrimination for Shot Boundary Detection

PERSIST redefines shot boundary detection as a task of semantic discrimination, requiring a persistent update of a video’s latent temporal state rather than a transient visual change. It employs a FiLM‑conditioned sinusoidal representation network and a structured discriminator that fuses local change, transient impulse, and return‑to‑trend cues into a single interpretable per‑frame signal. The method achieves comparable recall to leading detectors while significantly reducing false positives from flash, text overlay, and archival artifacts, and it is trained solely on real transitions from ClipShots.

By Tingyu Lin, Christian Stippel, Armin Dadras, Jakob Zenzmaier, Florian Kleber, Wolfgang Aigner, Robert Sablatnig
arXiv AI
Jun 29

OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal

arXiv:2606. 28094v1 Announce Type: cross Abstract: Real-world object removal is challenging due to two key difficulties: the target object's non-local effects, such as shadows and reflections, which are difficult to model, and the fact that user-provided masks are often inaccurate or incomplete.

By Qinming Zhou, Chenxi Sun, Deyang Kong, Junhao He, Xiangheng Tang, Peike Yu, Haotian Wu, Leilei Cao, Linfeng Zhang
arXiv AI
Jul 21

What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning

arXiv:2607. 16938v1 Announce Type: cross Abstract: End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation.

By Kalpana Panda, Wesley Maia, Vinti Agarwal, Ross Greer