Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models
arXiv:2606. 03748v1 Announce Type: cross Abstract: Real-time vision demands models that are accurate, efficient, and simple to deploy across diverse hardware.
TAPe+ML v3 is a compact computer vision system that uses a structured representation called TAPe to encode relationships among perceptual elements before recognition. The system employs a shared TAPe representation and a modular recognition architecture for tasks such as image classification, object detection, and instance segmentation, achieving strong performance with fewer than 100,000 parameters. Experiments show high mAP scores on COCO, strong classification accuracy on Imagenette and ImageNet-Real, and demonstrate benefits in video scene detection and distribution‑shift adaptation in an industrial pilot.
arXiv:2606. 03748v1 Announce Type: cross Abstract: Real-time vision demands models that are accurate, efficient, and simple to deploy across diverse hardware.
arXiv:2609.39924v1 Announce Type: cross Abstract: Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, m...
arXiv:2511. 12810v2 Announce Type: replace-cross Abstract: Camouflaged object detection is an emerging and challenging computer vision task that requires identifying and segmenting objects that blend seamlessly into their environments due to high similarity in color, texture, and size.
arXiv:2609.08914v2 Announce Type: replace Abstract: Pixel-level annotation of fixed traffic-camera imagery is expensive, while crosswalk models trained from street-level imagery face a substantial vi...
The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.
arXiv:2608.29917v1 Announce Type: new Abstract: Personalized segmentation and personalized retrieval both aim to identify the same physical object across different images. While the former localizes...
arXiv:2608.31052v1 Announce Type: cross Abstract: Semantic segmentation decomposes an image into distinct mask regions corresponding to different object categories, such as people, cars, signs or bui...
arXiv:2607. 22714v1 Announce Type: cross Abstract: Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive platforms impose severe constraints on compute, memory, and power.
The paper introduces a lightweight full-frame detector for partially manipulated AI-generated videos, suitable for edge deployment without face-detection preprocessing. It distills a DINOv2-Base teacher into a frozen MobileNetV3-Small student using temperature-annealed soft-label transfer, attention-diversity regularization, frame-level supervision, and a residual feature adapter. The model addresses false positives on legitimate scene cuts and threshold-level miscalibration, achieving an AUC of 0.766 on a 55,393-sample spliced test set while running at 3.65 ms per 16‑frame clip with a 150.4 MB checkpoint.
arXiv:2607. 06600v1 Announce Type: cross Abstract: Line segment detection is a key building block in visual SLAM, 3D reconstruction, and industrial inspection.
arXiv:2509.06422v2 Announce Type: replace Abstract: Video camouflaged object detection (VCOD) is challenging due to dynamic environments. Existing methods face two main issues: (1) SAM-based methods...
arXiv:2609.37631v1 Announce Type: new Abstract: Transformers are typically trained from random initialization, requiring all their capabilities to emerge from large-scale optimization. Recent work sh...