Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification
Read the original on arXiv Computer Vision →The paper introduces Sewer-Transformer-ML, a hierarchical vision Transformer that fuses multi‑level features for multi‑label sewer defect classification, and two lightweight variants, Sewer-MobileNet-ML and Sewer-Mobile-TransNet, tailored for resource‑constrained inspection scenarios. On the Sewer‑ML test set, Sewer‑Transformer‑ML‑Base achieved an $F2_{ ext{CIW}}$ of 65.68% and an $F1_{ ext{Normal}}$ of 92.68%, topping the public leaderboard and surpassing the next best method by 7.6 percentage points in $F2_{ ext{CIW}}$. The lightweight Sewer‑MobileNet‑ML reached a comparable $F2_{ ext{CIW}}$ of 65.73% with only 17 M parameters, a 95% reduction from the base model, while Sewer‑Mobile‑TransNet achieved 96.43% accuracy under the standard data split, and ablation studies highlighted the effectiveness of direct concatenation for Transformer features and attention‑based fusion for multiscale CNN features.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.