arXiv Computer Vision By Junyang Xia, Luocheng Zhang, Wenwen Pan, Chifeng Zhu, Yang Yang, Xinchun Liu, Jiajun Ding

Consensus-Aware Multi-Source Fusion for Reference-Guided Camouflaged Object Detection

Read the original on arXiv Computer Vision →

The paper introduces a consensus‑aware multi‑source fusion framework for reference‑guided camouflaged object detection. It couples trainable PVTv2 query features with frozen DINOv3 representations, using reference‑conditioned correlation to select foundation‑model evidence before multi‑scale fusion. The method also aggregates multiple references via cross‑reference consensus aggregation and injects reference information at semantic depths matched to the query features, achieving complementary improvements in experiments.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Aug 27

Context-Guided Semantic Alignment for Feature Fusion Networks

The paper introduces Feature Interaction Network (FINE), a lightweight semantic alignment module for feature fusion networks in object detectors. FINE refines low‑level features using high‑level contextual guidance through cross‑level attention, and employs Alignment‑Aware Token Sampling to reduce attention complexity. The resulting spatial‑channel modulation map selectively enhances semantically relevant pixels while preserving sub‑pixel localization, leading to improved detection accuracy with minimal computational overhead.

By Hyungseop Lee, Jiho Lee, Woochul Kang