Scalable Patch-Level Self-Supervised Learning
arXiv:2610.10013v1 Announce Type: new Abstract: Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of mul...
Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.
arXiv:2610.10013v1 Announce Type: new Abstract: Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of mul...
arXiv:2610.08983v1 Announce Type: new Abstract: Generative inpainting of brain MRI volumes is essential for synthesizing healthy tissue in pathological regions, improving the accuracy and reliability...
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we...
Pretrained models for cell and nuclear instance segmentation differ substantially in architecture, pretraining data and objectives, parameter count, inference strategy, adaptation requirements, postpr...
Self-supervised learning (SSL) at scale produces powerful visual representations. However, most scalable SSL methods rely on ad hoc combinations of multiple objectives and stabilization mechanisms. Ta...
The paper introduces a production-ready content‑extraction system tailored for generative AI workloads, addressing the heterogeneity of enterprise data formats such as PDFs, spreadsheets, and scanned documents. It features selective OCR routing, a scarcity‑first curation engine with a reference‑based extraction scorer, a deterministic structure‑aware chunker, and a read‑only retrieval evaluator that generates grounded questions and reports metrics like Hit@k and MRR. On a 180‑document corpus, the system achieves high accuracy (97.4/100 character score, 0.13% error rate) and strong retrieval performance (Hit@1 68.6%, Hit@10 92.8%, MRR 0.77).
arXiv:2610.07217v1 Announce Type: cross Abstract: Despite its promise for scaling robot learning, egocentric manipulation data is still scarce today. Collection at scale requires vertically integrati...
arXiv:2610.07338v1 Announce Type: cross Abstract: Audio benchmarks are built around short, pre-segmented clips, limiting model design to brief inputs or fixed vocabularies. To close this gap, we intr...
arXiv:2410.21582v4 Announce Type: replace-cross Abstract: Large-scale pretrained models are widely leveraged as foundations for learning new specialized tasks via fine-tuning, with the goal of mainta...
arXiv:2610.06970v1 Announce Type: new Abstract: Egocentric bimanual hand pose estimation is important for virtual interaction, wearable control, and rehabilitation, but visual observations are often...
arXiv:2610.08570v1 Announce Type: new Abstract: Handcrafted preprocessing is widely employed in automated dermoscopic analysis to suppress imaging artifacts and enhance lesion visibility. Nevertheles...
arXiv:2610.07576v1 Announce Type: cross Abstract: Cassini synthetic aperture radar (SAR) images reveal the dunes, plains, and lake basins of Titan, providing an instance of representations learned fr...
arXiv:2610.06938v1 Announce Type: new Abstract: Medical image segmentation remains fragmented along two axes: segmentation paradigms and data dimensionality. Existing methods are typically developed...
arXiv:2610.07014v1 Announce Type: new Abstract: RGB-D semantic segmentation has made notable progress by fusing RGB and Depth, yet mainstream models still learn features almost exclusively from pixel...
arXiv:2610.07378v1 Announce Type: new Abstract: Reconstructing cortical WM and pial surfaces from structural magnetic resonance imaging (MRI) is a prerequisite for surface-based neuroanatomical analy...
arXiv:2610.07982v1 Announce Type: new Abstract: Monocular metric depth estimation and 3D visual grounding represent the two complementary cornerstones of monocular 3D spatial understanding (M3Sun), f...
arXiv:2610.08068v1 Announce Type: new Abstract: Realistic physical interaction is a cornerstone of embodied intelligence, yet collecting paired visual--tactile data remains costly. Visual-to-tactile...
arXiv:2610.08770v1 Announce Type: new Abstract: Patch-based learning improves hyperspectral image (HSI) classification by exploiting local spectral-spatial information, but random train-test sampling...
arXiv:2607.09985v2 Announce Type: replace Abstract: Object pose estimation is a fundamental problem in 3D vision. Although recent state-of-the-art approaches achieve strong performance, generalizatio...
The paper investigates how rectifying supermarket product images using homography estimation and the Hough transform can improve deep learning-based object detection. It evaluates the impact of angle variation and object density on detection accuracy, highlighting both benefits and limitations of image rectification. The authors advocate for a new dataset to further study these effects.