Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,374 stories · RSS feed

arXiv Machine Learning
2d ago

RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments

arXiv:2610.10409v1 Announce Type: cross Abstract: General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities car...

By Zhiqin Yang, Chenxin Li, Xiaomeng Hu, Yibin Liu, Weidong Huang, Jiankai Sun, Haitao Li, Zijian Wu, Yuzhi Huang, Fanding Huang, Hanwen Sun, Jiashun Liu, Jingqi Tong, Mingxin Huang, Shaoli Hu, Shijue Huang, Tianyi Bai, Xinyuan Wang, Yunlong Lin, Zhengyang Tang, Zhexin Zhang, Zhuo Chen, Xierui Song, Juntao Dai, Boyuan Chen, Jiaming Ji, Fangneng Zhan, Mengkang Hu, Wei Xue, Yonggang Zhang, Han Hu, Tsung-Yi Ho, Yike Guo
arXiv Computer Vision
2d ago

Scalable Patch-Level Self-Supervised Learning

arXiv:2610.10013v1 Announce Type: new Abstract: Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of mul...

By Maximilian Seitzer, Gabriele Trivigno, Anton\'in Vobeck\'y, Seungeun Yi, Maxime Oquab, Huy V. Vo, Oriane Sim\'eoni, Piotr Bojanowski
arXiv AI
3d ago

Smart Content Ingestion for Generative AI Workloads

The paper introduces a production-ready content‑extraction system tailored for generative AI workloads, addressing the heterogeneity of enterprise data formats such as PDFs, spreadsheets, and scanned documents. It features selective OCR routing, a scarcity‑first curation engine with a reference‑based extraction scorer, a deterministic structure‑aware chunker, and a read‑only retrieval evaluator that generates grounded questions and reports metrics like Hit@k and MRR. On a 180‑document corpus, the system achieves high accuracy (97.4/100 character score, 0.13% error rate) and strong retrieval performance (Hit@1 68.6%, Hit@10 92.8%, MRR 0.77).

By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv Machine Learning
3d ago

Less Is More: A Leakage-Controlled Study of Dermoscopic Preprocessing for Joint Skin Lesion Classification and Segmentation with YOLO26

arXiv:2610.08570v1 Announce Type: new Abstract: Handcrafted preprocessing is widely employed in automated dermoscopic analysis to suppress imaging artifacts and enhance lesion visibility. Nevertheles...

By Truong Viet Vu, Nguyen Chi Hai, Nguyen Phuc Nguyen, Ngo Hoang Tu, Vo Nguyen Quoc Bao, Nguyen Thai Anh