arXiv:2610.10409v1 Announce Type: cross
Abstract: General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities car...
By Zhiqin Yang, Chenxin Li, Xiaomeng Hu, Yibin Liu, Weidong Huang, Jiankai Sun, Haitao Li, Zijian Wu, Yuzhi Huang, Fanding Huang, Hanwen Sun, Jiashun Liu, Jingqi Tong, Mingxin Huang, Shaoli Hu, Shijue Huang, Tianyi Bai, Xinyuan Wang, Yunlong Lin, Zhengyang Tang, Zhexin Zhang, Zhuo Chen, Xierui Song, Juntao Dai, Boyuan Chen, Jiaming Ji, Fangneng Zhan, Mengkang Hu, Wei Xue, Yonggang Zhang, Han Hu, Tsung-Yi Ho, Yike Guo
arXiv:2610.09807v1 Announce Type: new
Abstract: Camouflaged object detection requires pixel-accurate masks, but obtaining such annotations is slow and costly, making synthetic training images an attr...
By Akshat Dobhal, Sanjay Singh
arXiv:2610.09478v1 Announce Type: new
Abstract: Referring Expression Segmentation (RES) links natural-language descriptions to pixel-level object masks. Yet standard evaluation provides limited insig...
By Yuchen Li, Shaoyang Zhou, Yiran Wang, Ruiyi Deng, Haoyu Wang, Ziru Wei, Zhen Zhao, Luping Zhou
arXiv:2610.10225v1 Announce Type: new
Abstract: Whole slide image (WSI) analysis in computational pathology follows a multiple instance learning (MIL) pipeline where patch embeddings are extracted in...
By Haoyu He, Basile Tessier-Cloutier, Yang Wang, Mahdi S. Hosseini
arXiv:2610.10335v1 Announce Type: new
Abstract: Visual in-context learning (ICL), well suited to label-scarce medical imaging, uses support image-label pairs to demonstrate input-output mappings, whi...
By Cheng Wan, Chenjun Li, Qingyu Zhao
arXiv:2610.10013v1 Announce Type: new
Abstract: Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of mul...
By Maximilian Seitzer, Gabriele Trivigno, Anton\'in Vobeck\'y, Seungeun Yi, Maxime Oquab, Huy V. Vo, Oriane Sim\'eoni, Piotr Bojanowski
arXiv:2610.08983v1 Announce Type: new
Abstract: Generative inpainting of brain MRI volumes is essential for synthesizing healthy tissue in pathological regions, improving the accuracy and reliability...
By Arnela Hadzic, Franz Thaler, Simon Johannes Joham, Martin Urschler
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we...
Pretrained models for cell and nuclear instance segmentation differ substantially in architecture, pretraining data and objectives, parameter count, inference strategy, adaptation requirements, postpr...
Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Ta...
The paper introduces a production-ready content‑extraction system tailored for generative AI workloads, addressing the heterogeneity of enterprise data formats such as PDFs, spreadsheets, and scanned documents. It features selective OCR routing, a scarcity‑first curation engine with a reference‑based extraction scorer, a deterministic structure‑aware chunker, and a read‑only retrieval evaluator that generates grounded questions and reports metrics like Hit@k and MRR. On a 180‑document corpus, the system achieves high accuracy (97.4/100 character score, 0.13% error rate) and strong retrieval performance (Hit@1 68.6%, Hit@10 92.8%, MRR 0.77).
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2610.07217v1 Announce Type: cross
Abstract: Despite its promise for scaling robot learning, egocentric manipulation data is still scarce today. Collection at scale requires vertically integrati...
By Grounded Superintelligence, BitRobot
arXiv:2610.07338v1 Announce Type: cross
Abstract: Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we intr...
By Kwanghee Choi, Suwon Shon, Dmitriy Serdyuk, Guitang Lan, Chao-Wei Huang, Mohammad Sadegh Rasooli, Sangeeta Srivastava, Zhaojiang Lin, Saurabh Adya, Ming Sun
arXiv:2410.21582v4 Announce Type: replace-cross
Abstract: Large-scale pretrained models are widely leveraged as foundations for learning new specialized tasks via fine-tuning, with the goal of mainta...
By Jaedong Hwang, Brian Cheung, Zhang-Wei Hong, Akhilan Boopathy, Pulkit Agrawal, Ila Fiete
arXiv:2610.06970v1 Announce Type: new
Abstract: Egocentric bimanual hand pose estimation is important for virtual interaction, wearable control, and rehabilitation, but visual observations are often...
By JiaCheng Ge, SiYu Zhang, ShengJie Li, XinTong Yang
arXiv:2610.08570v1 Announce Type: new
Abstract: Handcrafted preprocessing is widely employed in automated dermoscopic analysis to suppress imaging artifacts and enhance lesion visibility. Nevertheles...
By Truong Viet Vu, Nguyen Chi Hai, Nguyen Phuc Nguyen, Ngo Hoang Tu, Vo Nguyen Quoc Bao, Nguyen Thai Anh
arXiv:2610.07576v1 Announce Type: cross
Abstract: Cassini synthetic aperture radar (SAR) images reveal the dunes, plains, and lake basins of Titan, providing an instance of representations learned fr...
By Kevin Lee
arXiv:2610.06938v1 Announce Type: new
Abstract: Medical image segmentation remains fragmented along two axes: segmentation paradigms and data dimensionality. Existing methods are typically developed...
By Bangwei Guo, Yunhe Gao, Meng Ye, Yang Zhou, Difei Gu, Guoning Zhang, Leon Axel, Dimitris Metaxas
arXiv:2610.07014v1 Announce Type: new
Abstract: RGB-D semantic segmentation has made notable progress by fusing RGB and Depth, yet mainstream models still learn features almost exclusively from pixel...
By Ziang Wei, Yinlong Liu, Yan Xia, Alois Knoll, Hu Cao
arXiv:2610.07378v1 Announce Type: new
Abstract: Reconstructing cortical WM and pial surfaces from structural magnetic resonance imaging (MRI) is a prerequisite for surface-based neuroanatomical analy...
By Kaveh Moradkhani, Sylvain Bouix