arXiv Computer Vision By Tianle Fang, Zhenbing Liu, Chong Yin, Bolun Li, Haoxiang Lu

C2FXNet: Coarse-to-Fine Scene Expert for Unified Object Detection across Adverse Weather

Read the original on arXiv Computer Vision →

C2FXNet is a unified object detection framework designed for adverse weather conditions. It uses a dual-level guidance mechanism: a Multi-step Reasoning Router (MRR) for coarse scene reasoning and a Fine Scene Refinement (FSR) module for fine-grained semantic adjustment. A Scene-aware Mixture-of-Experts (SMoE) dynamically combines scene-specific experts, enabling robust detection across foggy, dark, and clear scenes without scene-specific training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv AI
Sep 2

SARTM: Segment Any RGB Thermal Model with Language aided Distillation

The paper introduces SARTM, a framework that adapts the Segment Anything Model (SAM) for RGB‑thermal (RGB‑T) semantic segmentation. It fine‑tunes SAM with LoRA layers, incorporates language guidance, and employs a Cross‑Modal Knowledge Distillation module to bridge modality gaps. The approach also modifies the segmentation head and adds an auxiliary semantic head, achieving superior performance on MFNET, PST900, and FMB benchmarks.

By Dong Xing, Jinhe Zhang, Hang Yang, Yuqing Wang
arXiv AI
Jul 28

ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness

arXiv:2607. 23537v1 Announce Type: new Abstract: Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standard conditions, synthetic corruptions, or single modality.

By Qiao Yan, Yihan Wang, Zhenghao Xing, Jiaqi Xu, Pheng-Ann Heng
Hugging Face Trending Papers
Aug 13

P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation

Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.

arXiv Computer Vision
Sep 4

ProgResViT: Progressive Resolution and Width for Adaptive Vision Transformers

ProgResViT is an input‑adaptive Vision Transformer that processes images progressively across multiple rounds, starting with a low‑resolution image and a narrow subnetwork and refining the prediction with higher resolution and a wider subnetwork if needed. The method introduces Progress‑Conditioned Soft Gating (PSG) to share a single backbone across rounds while conditioning token fusion and layer outputs on the current round, block, and input resolution. Experiments on DeiT show improved accuracy‑compute trade‑offs compared to adaptive‑width, adaptive‑depth, and dynamic‑token baselines, and the design also benefits self‑supervised DINO representations and downstream semantic segmentation.

By Ali Hojjat, Janek Haberer, Olaf Landsiedel