arXiv Computer Vision By Sergey Kurinov (Comexp Research Lab, TAPe + ML Project, Nizhniy Novgorod, Russia), Alexey Upatov (Comexp Research Lab, TAPe + ML Project, Nizhniy Novgorod, Russia)

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision

Read the original on arXiv Computer Vision →

TAPe+ML v3 is a compact computer vision system that uses a structured representation called TAPe to encode relationships among perceptual elements before recognition. The system employs a shared TAPe representation and a modular recognition architecture for tasks such as image classification, object detection, and instance segmentation, achieving strong performance with fewer than 100,000 parameters. Experiments show high mAP scores on COCO, strong classification accuracy on Imagenette and ImageNet-Real, and demonstrate benefits in video scene detection and distribution‑shift adaptation in an industrial pilot.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
Sep 17

Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM

The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.

By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa