arXiv Machine Learning By Annajoyce Mariani, Kira Maag, Hanno Gottschalk

Extended Field of View Analysis for VideoGAN-based Trajectory Generation

Read the original on arXiv Machine Learning →

arXiv:2608. 02289v1 Announce Type: cross Abstract: Realistic and diverse trajectory generation is central to enabling higher levels of vehicle automation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 6

TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely predominantly on geometric supervision, such as occupancy regression and optical flow, effectively treating scene agents as generic moving obstacles.

arXiv AI
Sep 10

DGCPath: Distribution-Aware Generative Contrastive Framework for Self-supervised Path Representation Learning -- Extended Version

DGCPath is a Distribution‑Aware Generative Contrastive framework designed for self‑supervised path representation learning. It combines a diffusion‑based view generator, a variational contrastive mechanism that aligns latent features at the distribution level, and a generative cross‑supervision module for view‑level consistency. Experiments on three real‑world trajectory datasets show that DGCPath surpasses state‑of‑the‑art baselines on two downstream tasks, indicating stronger generalization and representation effectiveness.

By Sean Bin Yang, Hao Miao, Zongyi Xu, Jilin Hu, Xiangmeng Wang, Hua Lu, Bin Yang, Christian S. Jensen
arXiv Computer Vision
Sep 7

Object Concepts Emerge from Motion

The paper introduces a biologically inspired framework that learns object‑centric visual representations from raw videos without human annotations or camera calibration. By using motion boundaries detected via optical flow and clustering to create pseudo‑instance masks, the method supervises a single‑image encoder with pixel‑level pairwise metric learning. Training on 195 million pseudo‑labeled frames and expanding to 421 million frames through Motion‑Verified Self‑Training, the approach yields Swin‑based encoders that outperform or match supervised and self‑supervised baselines on tasks such as monocular depth estimation, 3D object detection, 3D occupancy prediction, and end‑to‑end planning.

By Boshi Li, Xiaohui Wang, Xiaoyang Wu, Zhichao Li, Ya Yang, Naiyan Wang