arXiv AI By William Bonilla, Maxime Boisvert, David-Alexandre Poissant, David Meger, Louis Petit

FLINT: Fast Lightweight Inference for Traversability

Read the original on arXiv AI →

FLINT is a lightweight traversability estimator that uses a 21.6‑million‑parameter backbone—38 times smaller than comparable foundation models—to predict traversability from a single RGB camera. It achieves higher accuracy on held‑out terrain probes and runs at 14.7 FPS on CPU, outperforming a deployed foundation‑model system (WildOS) on 23 of 24 field logs. In closed‑loop field trials, FLINT’s best self‑supervised head reached 99% autonomy, surpassing a human‑label‑trained baseline on the same course.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 7

Towards Robust Driving Perception: A Flexible Scale-Driven Family for Self-Supervised Monocular Depth Estimation

The paper introduces FlexDepth, a family of self‑supervised monocular depth estimation models designed for robust driving perception. FlexDepth uses a two‑stage static‑dynamic decoupled training strategy and a Scale‑Driven Decoder that selects components based on scale size, enabling efficient feature fusion and high‑precision depth maps. Experiments on driving benchmarks show state‑of‑the‑art performance across arbitrary scales with minimal computational cost, with the smallest model (Flex‑Nano) achieving 37.6 FPS on mobile devices.

By Zhaowen Zhu, Li Zhang, Yujie Chen, Tian Zhang, Yingjie Wang, Mingxia Zhan
arXiv Computer Vision
Aug 27

Lightweight Machine Learning-Driven Monocular Sidewalk Path Extraction for Embedded Micromobility Navigation

The paper presents a lightweight monocular vision pipeline for extracting sidewalk paths on low‑power embedded micromobility platforms. It evolves through three design iterations—from a skeleton‑graph baseline to a distance‑transform corridor planner and finally to a compact image‑space architecture—using a SegFormer‑B0 student model trained with semi‑supervised pseudo‑labels. The final system achieves high segmentation accuracy (IoU 0.946) and fast planning (under 50 ms per frame) while reducing temporal instability and increasing template‑path availability across real campus sequences.

By Lkhanaajav Mijiddorj, Yang Yan, Tyler Beringer, Bilguunzaya Mijiddorj, Alex N. Ho, Bin Xu, Binbin Weng
arXiv Computer Vision
6d ago

Learning to Navigate with Minimal Parameters: Decomposing Visual Navigation Through Closed-Form Geometric Interfaces

The paper introduces a compact visual navigation system that decomposes the task into three analytically‑computed geometric interfaces and three small learned modules: an egress predictor, a navigation predictor, and an endpoint‑pinned residual diffusion generator. Only 0.58 M of the 23 M parameters are trained on 44 k frames, achieving competitive success rates and the lowest collision rate among evaluated methods across 6 060 point‑goal episodes in 60 environments. The design allows further parameter reduction by replacing the frozen image encoder with a 0.54 M MobileNetV2, supports zero‑shot deployment on a Jetson Orin Nano UGV, and enables transparent failure analysis under sensor corruption.

By Edward Beng Wai Tan, Siew-Kei Lam
arXiv AI
Jul 10

Time-to-Collision Based Dynamic Obstacle Avoidance Using Pretrained Vision Models for Robots in Unstructured Environments

arXiv:2607. 07885v1 Announce Type: cross Abstract: Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly when large-scale robot-specific training data and simulation-based policies are impractical.

By Erik Jagnandan, Mulugeta Haile, Gregory Barber, Pratik Chaudhari