Co-Fusion4D: Spatio-temporal Collaborative Fusion for Robust 3D Object Detection
arXiv:2605. 20301v2 Announce Type: replace-cross Abstract: In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making.
arXiv:2605. 20301v2 Announce Type: replace-cross Abstract: In autonomous driving, 3D object detection is essential for accurate perception and reliable decision-making.
The paper introduces OccLinker, a lightweight plugin for vision‑based occupancy networks that reduces flickering by efficiently merging historical static and motion cues with current features via a dual cross‑attention mechanism. It generates correction components to refine base network predictions and proposes a new temporal consistency metric to quantify flickering. Experiments on two benchmark datasets show that OccLinker improves performance with minimal computational overhead while effectively diminishing flickering artifacts.
Connected Autonomous Vehicles (CAVs) increasingly exploit Vehicle-to-Everything (V2X) communication to exchange multi-source sensor information, enabling advanced Collaborative Perception (CP) capabilities. Extending beyond these capabilities, this work focuses on Collaborative Joint Perception and Prediction (Co-P&P), a paradigm that unifies CP with motion prediction to mitigate two persistent challenges: the accumulation of perception errors and visual occlusions.
arXiv:2606. 07708v1 Announce Type: cross Abstract: We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections.
Reconstructing humans and their surrounding environments in a globally consistent 4D space is essential for comprehensive perception. However, prior works typically assume single-view inputs or decouple humans, scenes, and cameras, making them unable to recover coherent geometry, stable motion, and physically aligned trajectories.
arXiv:2609.00951v1 Announce Type: new Abstract: Collaborative perception shares information among multiple agents to obtain a comprehensive scene representation, enhancing the perceptual capability o...
Diffusion models have shown strong potential for multi-modal planning in end-to-end autonomous driving. However, most existing methods confine diffusion to the planning module, conditioning on fixed outputs from separate discriminative perception networks.
arXiv:2608. 12187v1 Announce Type: cross Abstract: Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level structural information before temporal modelling.
Collaborative perception shares information among multiple agents to obtain a comprehensive scene representation, enhancing the perceptual capability of individual agents. However, most existing metho...
Cross-view geo-localization is challenging due to drastic viewpoint changes and large appearance discrepancies between street-level and satellite imagery. Although existing methods often use geometric warping to expose co-visible cues, such transformations rely on restrictive spatial assumptions and inevitably introduce severe visual distortions under view-dependent visibility, yielding noisy supervision and fragile correspondences.
arXiv:2607. 19036v1 Announce Type: cross Abstract: V2X collaborative object detection features overcoming the limitations of single-vehicle systems by aggregating environmental features from multiple collaborative agents.
arXiv:2608.22398v1 Announce Type: cross Abstract: Reliably tracking moving deformable linear objects (DLOs) while simultaneously ensuring robustness, accuracy, and temporally consistent state estimat...