arXiv Computer Vision

Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification

The paper introduces Sewer-Transformer-ML, a hierarchical vision Transformer that fuses multi‑level features for multi‑label sewer defect classification, and two lightweight variants, Sewer-MobileNet-ML and Sewer-Mobile-TransNet, tailored for resource‑constrained inspection scenarios. On the Sewer‑ML test set, Sewer‑Transformer‑ML‑Base achieved an $F2_{ ext{CIW}}$ of 65.68% and an $F1_{ ext{Normal}}$ of 92.68%, topping the public leaderboard and surpassing the next best method by 7.6 percentage points in $F2_{ ext{CIW}}$. The lightweight Sewer‑MobileNet‑ML reached a comparable $F2_{ ext{CIW}}$ of 65.73% with only 17 M parameters, a 95% reduction from the base model, while Sewer‑Mobile‑TransNet achieved 96.43% accuracy under the standard data split, and ablation studies highlighted the effectiveness of direct concatenation for Transformer features and attention‑based fusion for multiscale CNN features.

arXiv Machine Learning
Jul 21

Leakage-Robust Evaluation and Data-Scale Sensitivity of Attention-Enhanced Multi-Task Learning for Joint Fault Diagnosis and Remaining Useful Life Estimation

arXiv:2607. 16493v1 Announce Type: new Abstract: Multi-task deep learning models that jointly perform fault classification and remaining useful life (RUL) regression are increasingly used in predictive maintenance, yet reported performance can be strongly affected by how sliding-window sequences are split into training and test sets.

By Md Mahamudur Rahaman Shamim, Md. Nuruzzaman, Zannatul Ferdus, Md Rajib Ahmed, Abieer Nwshad Anward, Mohammad Tooneer, Johir Uddin Khan, Khalid Hossen
arXiv AI
Sep 17

Tabular Deep Learning vs Classical Machine Learning for Urban Land Cover Classification

The paper compares classical machine learning algorithms—such as Logistic Regression, SVM, Random Forest, XGBoost, and CatBoost—with Tabular Deep Learning models (TabNet, FT-Transformer, TabTransformer, TabSeq, and 1D CNNs) for urban land cover classification using a UCI dataset derived from high‑resolution aerial imagery. It evaluates performance across nine land cover classes, addressing challenges like high dimensionality, heterogeneous features, and class imbalance by applying weighted cross‑entropy loss for deep models and measuring accuracy, macro‑precision, macro‑recall, macro‑F1, AUC‑ROC, and confusion matrices. Results indicate that while tree ensembles remain strong baselines, Tabular Deep Learning can match or surpass them when non‑linear interactions are prominent and imbalance handling is effective.

By Muntasir Tabasum, Tanpia Tasnim, Md. Ekramul Islam, Al Zadid Sultan Bin Habib