Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,329 stories · RSS feed

arXiv Computer Vision
1d ago

MultiFly: A Real-World Multimodal Aerial Dataset with Annotation-Efficient Label Transfer and Cross-Modal Semantic Consistency

arXiv:2610.10359v1 Announce Type: cross Abstract: We introduce MultiFly, a real-world, low-altitude UAV dataset for semantic perception across RGB, thermal, LiDAR, and radar modalities. MultiFly prov...

By Markus Gross, Andreas Greiner, Taehyoung Kim, Sivasubiramaniam Subbiah, Toma\v{z} Coti\v{c}, Sai Bharadwaj Matha, Conrad Christoph, Oussema Dhaouadi, Simon Zieher, Surya Vijaya Kumar, Gordon Elger, Henri Mee{\ss}, Olaf Wysocki, Paul Spannaus, Daniel Cremers
arXiv Machine Learning
1d ago

RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments

arXiv:2610.10409v1 Announce Type: cross Abstract: General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities car...

By Zhiqin Yang, Chenxin Li, Xiaomeng Hu, Yibin Liu, Weidong Huang, Jiankai Sun, Haitao Li, Zijian Wu, Yuzhi Huang, Fanding Huang, Hanwen Sun, Jiashun Liu, Jingqi Tong, Mingxin Huang, Shaoli Hu, Shijue Huang, Tianyi Bai, Xinyuan Wang, Yunlong Lin, Zhengyang Tang, Zhexin Zhang, Zhuo Chen, Xierui Song, Juntao Dai, Boyuan Chen, Jiaming Ji, Fangneng Zhan, Mengkang Hu, Wei Xue, Yonggang Zhang, Han Hu, Tsung-Yi Ho, Yike Guo