Falcon Perception
Related stories
Welcome to the Falcon 3 Family of Open Models!
Spread Your Wings: Falcon 180B is here
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
The paper introduces DroneCATS-Agent, a modular framework that places a multimodal large language model (MLLM) at the core of a drone’s control loop, allowing the model to decide actions solely from prompts. It presents the DroneCATS benchmark, evaluating MLLMs on four tasks—approaching, tracking, searching, and multi‑drone commanding—without fine‑tuning or function‑calling. Results show that while small open models can navigate reliably, they often fail by mismanaging protocol termination, highlighting a gap between perception and action planning in current MLLMs.
FalconTrack: Photorealistic Auto-Labeled Perception and Physics-Aware Vision-Based Aerial Tracking
arXiv:2606. 29783v1 Announce Type: cross Abstract: Vision-based aerial tracking is critical in GPS-denied environments.
FALCON-S: Fixed-wing ground-effect Aerodynamics Simulator and Flight Control Learning Suite
FALCON‑S is a modular, high‑fidelity simulator designed for fixed‑wing aerial robots operating near the ground. It models full 6DoF rigid‑body physics, semi‑empirical ground‑effect aerodynamics, actuator dynamics, sensor noise, and environmental disturbances, and supports both CPU and GPU backends via Torch and NVIDIA Warp for large‑scale reinforcement learning and optimal control. The framework offers a unified interface for various controllers, including RL and optical control algorithms, and allows cross‑validation with X‑Plane and JSBSim for engineering integration and visual fidelity.
Estimating Central, Peripheral, and Temporal Visual Contributions to Human Decision Making in Atari Games
arXiv:2604. 04439v2 Announce Type: replace Abstract: We study how different visual information sources contribute to human decision making in dynamic visual environments.
MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in Bird's-Eye-View Images
MC-DeTra is a reimplementation of the DeTra model that jointly performs object detection and socially-aware trajectory forecasting in bird's-eye-view images. It introduces motion-consistency mechanisms that add supervision from each actor’s past motion, surrounding traffic occupancy, and a consistency constraint aligning predicted heading with motion direction. The added losses are train‑only and inference‑safe, improving dynamic trajectory forecasting on the Waymo Open Dataset while maintaining or enhancing detection accuracy.
PerchRL: Vision-Based Agile Perching on Inclined Platforms under Rapid and Irregular Motion
arXiv:2606. 03441v1 Announce Type: cross Abstract: Autonomous vision-based perching of quadrotors on moving inclined platforms is critical for air-ground collaboration but remains challenging due to the limited field of view (FOV).
AirDreamer: Generalist Drone Navigation with World Models
arXiv:2606. 03252v1 Announce Type: cross Abstract: Navigating a drone in unseen and cluttered environments requires reliable generalization to unseen scene layouts and understanding of environmental structure relative to the robot's capabilities.
Cross-View Urban Traffic Dataset: Drone-Supervised Ground Truth for Monocular Bird's-Eye View Localization
arXiv:2606. 07708v1 Announce Type: cross Abstract: We introduce a dataset and benchmark for cross-view urban traffic perception built from synchronized ego-centric bicycle videos and aerial drone videos recorded at real urban intersections.
Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects
arXiv:2608. 07577v1 Announce Type: cross Abstract: A closed-set detector for autonomous driving must assign every object one of a fixed set of labels.