arXiv Computer Vision

TAPe+ML: A Compact Structured Representation for Multi-Task Computer Vision

TAPe+ML v3 is a compact computer vision system that uses a structured representation called TAPe to encode relationships among perceptual elements before recognition. The system employs a shared TAPe representation and a modular recognition architecture for tasks such as image classification, object detection, and instance segmentation, achieving strong performance with fewer than 100,000 parameters. Experiments show high mAP scores on COCO, strong classification accuracy on Imagenette and ImageNet-Real, and demonstrate benefits in video scene detection and distribution‑shift adaptation in an industrial pilot.

arXiv Computer Vision
Sep 17

Decoder-Agnostic Token Merging for Vision Transformers: A Systematic Study of G2TM

The paper studies Graph-Guided Token Merging (G2TM), a module that reduces token count in Vision Transformers. It evaluates G2TM across multiple segmentation frameworks and decoder types, finding that its performance gains are tied to the encoder rather than the decoder. The authors report consistent reductions in GFLOPs (22‑47%) and throughput improvements (up to 74%) on ADE20K, with optimal hyperparameters depending mainly on backbone pre‑training and target dataset.

By Victor Bercy, Martyna Poreba, Michal Szczepanski, Samia Bouchafa
arXiv Computer Vision
Sep 18

AI or Real: Detecting Partially Altered Videos Under Resource-Constrained Environments

The paper introduces a lightweight full-frame detector for partially manipulated AI-generated videos, suitable for edge deployment without face-detection preprocessing. It distills a DINOv2-Base teacher into a frozen MobileNetV3-Small student using temperature-annealed soft-label transfer, attention-diversity regularization, frame-level supervision, and a residual feature adapter. The model addresses false positives on legitimate scene cuts and threshold-level miscalibration, achieving an AUC of 0.766 on a 55,393-sample spliced test set while running at 3.65 ms per 16‑frame clip with a 150.4 MB checkpoint.

By Tamoghna Chakraborty, Md Nurul Absur, Sourya Saha, Saptarshi Debroy