arXiv:2604.13287v2 Announce Type: replace
Abstract: Weight pruning is a common technique for compressing large neural networks. We focus on the challenging post-training one-shot setting, where a pre...
By Gabriel Afriat, Xiang Meng, Shibal Ibrahim, Hussein Hazimeh, Rahul Mazumder
arXiv:2407.03463v2 Announce Type: replace-cross
Abstract: In the realm of self-supervised learning (SSL), conventional wisdom has gravitated towards the utility of massive, general domain datasets fo...
By Jes\'us M Rodr\'iguez-de-Vera, Imanol G Estepa, Ignacio Saras\'ua, Bhalaji Nagarajan, Petia Radeva
arXiv:2605. 30188v2 Announce Type: replace-cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated.
By Eug\`ene Berta, David Holzm\"uller, Francis Bach, Michael I. Jordan
arXiv:2603. 15106v2 Announce Type: replace Abstract: Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typically requires DNN architectures to be specialized for each device separately.
By Mark Deutel, Simon Geis, Axel Plinge
The paper introduces a method for efficiently exploring the Rashomon set of Concept Bottleneck Models (CBMs) by using a parallel parameter‑efficient adaptation module, checkpointing, and a concept diversity objective. This approach generates multiple equally accurate CBMs from a single training process, achieving greater diversity than baseline methods while consuming less memory. The resulting diverse models enable trustworthy selection, reduce inter‑class confusion, and support reliable abstention in decision‑making.
By Shihan Feng, Cheng Zhang, Michael Xi, Ethan Hsu, Lesia Semenova, Chudi Zhong
The paper evaluates out‑of‑the‑box object detection models for automatic target detection and recognition (ATD/R) in military settings. Six YOLO variants and two DETR variants were benchmarked on a new military dataset featuring vehicles, occlusions, and small targets, with performance measured in mAP@0.5 and mAP@0.5:0.95 across air‑to‑ground and ground‑to‑ground perspectives. Findings show larger models and DETR-based approaches perform best, fine‑tuning on the VisDrone dataset improves air‑to‑ground and small‑object performance, yet all models still struggle with small targets in air‑to‑ground scenarios.
By Alma M. Liezenga, Lotte Nijskens, Henrik R. Baumann, Stefan Becker, Simon Bensberg, Niccol\`o Camarlinghi, H{\aa}vard R. Eiring, Alexander W. Johnsgaard, Tanel Liiv, Giuseppe Martino, Matteo Marturini, Matthias Rapp, Jan Erik van Woerden, Alexander Wolpert, Hugo J. Kuijf
The paper introduces a meta‑learning framework that uses a rich set of dataset‑complexity meta‑features to predict the accuracy of different classifiers on image datasets, avoiding exhaustive training. By extracting features with autoencoders, pre‑trained networks, and dimensionality reduction, regression models estimate classifier accuracies, while clustering groups similar performers to simplify recommendations. Tested on 56 diverse image datasets, the method achieves over 86% ranking prediction accuracy, offering a scalable, interpretable solution for model selection and cost reduction.
By Zahra Nabizadeh_Shahre_Babak, Farzaneh Koohestani, Nader Karimi, Shahram Shirani, Shadrokh Samavi
arXiv:2511. 19636v2 Announce Type: replace-cross Abstract: In many machine learning problems, there may exist multiple models that achieve nearly identical predictive performance while relying on fundamentally different internal logic.
By Shihan Feng, Cheng Zhang, Michael Xi, Ethan Hsu, Lesia Semenova, Chudi Zhong
arXiv:2607. 15745v1 Announce Type: new Abstract: Common practice when training Convolutional Neural Networks (CNNs) is to use randomly shuffled mini-batches.
By Anxhelo Shehu, Enes Stastoli, Arben Cela
The paper proposes a Learnware-based framework for deploying scene‑specific CSI feedback models in 6G systems. A centralized AI data center maintains a catalog of pre‑trained models, each tagged with semantic and statistical specifications. Base stations retrieve the most relevant model using only statistical fingerprints, which reduces data privacy risks, lowers retrieval latency, and cuts fine‑tuning effort, achieving up to 57.7% performance gains over a general model.
By Xiangyi Li, Jiajia Guo, Chao-Kai Wen, Xin Geng, Shi Jin, Zhi-Hua Zhou
arXiv:2409. 16808v3 Announce Type: replace-cross Abstract: Modern applications such as autonomous vehicles, intelligent surveillance, and smart city systems increasingly require object detection on resource-constrained edge devices.
By Daghash K. Alqahtani, Muhammad Aamir Cheema, Maria A. Rodriguez, Adel N. Toosi
The paper investigates the use of Zero Cost Proxies (ZCPs) to identify high‑performing wearable Human Activity Recognition (HAR) models without full training. Eight ZCPs were evaluated across six benchmark HAR datasets, showing that the top‑predicted architectures achieve performance within 7% of fully trained models, and training the top‑10 predictions reaches within 2% of full training. This demonstrates that ZCPs can significantly reduce computational costs while maintaining competitive accuracy in sensor‑based HAR tasks.
By Richard Goldman, Varun Komperla, Thomas Ploetz, Harish Haresamudram