arXiv:2608. 11537v1 Announce Type: cross Abstract: Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boundary mixing, whereas latent-feature decoders that predict a separate output distribution may relegate the rendered image to an intermediate visualization.
By Weize Cai, Yongqi Dong, Zhida Shao, Zixin Fu
The paper demonstrates that a frozen generic text‑to‑image diffusion model can perform conditional inpainting across three natural‑image domains without any inpainting‑specific training or dataset adaptation. The proposed Step‑PI method augments known‑region projection with boundary‑interior latent feedback, persistent PI state, and a predefined four‑field release schedule, improving all evaluated metrics over baseline training‑free approaches. Experiments on CelebA‑HQ, AFHQ, and Places2 show consistent gains, with Step‑PI outperforming LanPaint and PILOT on all five macro metrics.
By Zhenhuan Wang, Fengyi Yuan
The paper investigates how knowledge distillation from event cameras to RGB images can alter the inductive biases of convolutional neural networks. By transferring learning from the event domain, the authors find that models gain color invariance, a shape bias, and improved robustness to high‑frequency noise, largely due to reduced reliance on texture and increased emphasis on edge‑based object shape. These changes are evidenced by early‑layer processing differences and a spectral trade‑off between robustness to missing high‑frequency content and vulnerability to its contamination or geometric disruption.
By Soshun Kihara, Shunsuke Yasuki, Masato Taki
arXiv:2609.38111v1 Announce Type: new
Abstract: Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when...
By Hongzhu Guo, Mohsen Fayyaz, Nanyun Peng
arXiv:2609.01584v1 Announce Type: new
Abstract: Vehicle attribute analysis is a key component of Intelligent Transportation Systems (ITS), supporting applications such as vehicle identification, traf...
By Sergio M. Silva Jr., Otavio T. Remer, Gabriel E. Lima, Lucas Wojcik, Rayson Laroca, David Menotti
arXiv:2606. 31603v1 Announce Type: cross Abstract: Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.
By Nikolai R\"ohrich, Julian Glei{\ss}ner, Ahmed H. A. Ibrahim, Silvan Mertes, Tobias Huber
arXiv:2609.38010v1 Announce Type: cross
Abstract: Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety m...
By Mohamed Benkedadra, Aissa Saoudi, Maxime Gloesener, Sidi Ahmed Mahmoudi, Matei Mancas
arXiv:2609.14383v1 Announce Type: new
Abstract: Real-time pedestrian detection in driving scenes is constrained by three coupled failure modes: tiny targets lose discriminative evidence, occlusion we...
By Sam Williams, Yuan Xiang
arXiv:2604. 16514v5 Announce Type: replace-cross Abstract: Autoregressive vision-language models (VLMs) deliver strong multimodal capability, but their token-by-token decoding imposes a fundamental inference bottleneck.
By Baoyou Chen, Hanchen Xia, Peng Tu, Haojun Shi, Liwei Zhang, Yuxuan Yao, Weihao Yuan, Siyu Zhu
PERSIST redefines shot boundary detection as a task of semantic discrimination, requiring a persistent update of a video’s latent temporal state rather than a transient visual change. It employs a FiLM‑conditioned sinusoidal representation network and a structured discriminator that fuses local change, transient impulse, and return‑to‑trend cues into a single interpretable per‑frame signal. The method achieves comparable recall to leading detectors while significantly reducing false positives from flash, text overlay, and archival artifacts, and it is trained solely on real transitions from ClipShots.
By Tingyu Lin, Christian Stippel, Armin Dadras, Jakob Zenzmaier, Florian Kleber, Wolfgang Aigner, Robert Sablatnig
arXiv:2606. 28094v1 Announce Type: cross Abstract: Real-world object removal is challenging due to two key difficulties: the target object's non-local effects, such as shadows and reflections, which are difficult to model, and the fact that user-provided masks are often inaccurate or incomplete.
By Qinming Zhou, Chenxi Sun, Deyang Kong, Junhao He, Xiangheng Tang, Peike Yu, Haotian Wu, Leilei Cao, Linfeng Zhang
arXiv:2607. 16938v1 Announce Type: cross Abstract: End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed paths for open-loop evaluation and often effective driving in closed-loop evaluation.
By Kalpana Panda, Wesley Maia, Vinti Agarwal, Ross Greer