Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

12,559 stories · RSS feed

arXiv Machine Learning
4d ago

Learning Varying Physical Therapist-Patient Interactions for Robot-mediated Upper Limb Task-Specific Training

arXiv:2608. 15995v1 Announce Type: cross Abstract: Upper extremity motor function recovery is positively linked to Task-Specific Training (TST) and sufficient therapy dosage.

By Jia Quan Loh (Human Robotics Laboratory, Department of Mechanical Engineering, The University of Melbourne), Vincent Crocher (Human Robotics Laboratory, Department of Mechanical Engineering, The University of Melbourne), Marlena Klaic (Melbourne School of Health Sciences, The University of Melbourne), Denny Oetomo (Human Robotics Laboratory, Department of Mechanical Engineering, The University of Melbourne), Ying Tan (Human Robotics Laboratory, Department of Mechanical Engineering, The University of Melbourne)
arXiv Machine Learning
4d ago

Structured Prediction for Scalable Spreadsheet Table Understanding: From Cell Types to Table Ranges (Extended Version)

arXiv:2608. 16050v1 Announce Type: cross Abstract: Spreadsheets are a primary medium for publishing tabular data, yet automatically extracting structured content from them remains difficult due to heterogeneous layouts, diverse file formats, and inconsistent organizational conventions.

By Antoine Gauquier, Ioana Manolescu, Pierre Senellart
arXiv Machine Learning
4d ago

Convolution-Free Holistic Multivariance Decomposition Layer for Efficient Hyperspectral Image Classification Tensor Networks

arXiv:2608. 16241v1 Announce Type: cross Abstract: Feature extraction for hyperspectral image classification is conventionally addressed using rigid tensor decompositions that fail to capture complex spatio-spectral interdependencies, or heavily parameterized convolutional neural networks that are computationally expensive.

By S\"uha Tuna, \"Ulker Ba\c{s}ar
arXiv Machine Learning
4d ago

POI Recommendation with LLM-Augmented Multi-Graph Learning and Contrastive Alignment

arXiv:2608. 16407v1 Announce Type: cross Abstract: Point-of-interest (POI) recommendation models based on graph neural networks achieve strong performance by propagating collaborative signals over user-item interactions, yet they struggle with the cold-start problem, where items with few or no interactions are not represented.

By Burak Tamer, Wolfram H\"opken, Zehui Wang
arXiv Machine Learning
4d ago

Turning spectra into images improves plant trait retrieval with 2D-CNNs

arXiv:2608. 16661v1 Announce Type: cross Abstract: Hyperspectral reflectance spectroscopy enables non-destructive estimation of plant functional traits, yet current deep learning approaches process spectra as one-dimensional sequences, which limits how they capture long-range inter-band dependencies.

By Javier Lopatin, Teja Kattenborn, Eya Cherif, Sebasti\'an Moreno
arXiv Machine Learning
4d ago

CodeQuant: Unified Clustering and Quantization for Enhanced Outlier Smoothing in Low-Precision Mixture-of-Experts

arXiv:2604. 10496v2 Announce Type: replace Abstract: Outliers have emerged as a fundamental bottleneck in preserving accuracy for low-precision large models, particularly within Mixture-of-Experts (MoE) architectures that are increasingly central to large-scale language modeling.

By Xiangyang Yin, Xingyu Liu, Tianhua Xia, Bo Bao, Vithursan Thangarasa, Valavan Manohararajah, Eric Sather, Sai Qian Zhang
arXiv Machine Learning
4d ago

Helios 2.0: A Robust, Ultra-Low Power Gesture Recognition System Optimised for Event-Sensor based Wearables

arXiv:2503. 07825v3 Announce Type: replace-cross Abstract: We present an advance in wearable technology: a mobile-optimized, real-time, ultra-low-power event camera system that enables natural hand gesture control for smart glasses, dramatically improving user experience.

By Prarthana Bhattacharyya, Joshua Mitton, Ryan Page, Owen Morgan, Oliver Powell, Benjamin Menzies, Gabriel Homewood, Kemi Jacobs, Paolo Baesso, Taru Muhonen, Richard Vigars, Louis Berridge
arXiv Machine Learning
4d ago

Comprehensive language-image pre-training for 3D medical image understanding

arXiv:2510. 15042v3 Announce Type: replace-cross Abstract: In the 3D medical image domain, vision-language pre-training is used to create vision-language encoders (VLEs) that can support radiologists by retrieving patients with similar abnormalities, predicting likelihoods of abnormality, or, with downstream adaptation, generating radiological reports.

By Tassilo Wald, Ibrahim Ethem Hamamci, Yuan Gao, Sam Bond-Taylor, Harshita Sharma, Maximilian Ilse, Cynthia Lo, Olesya Melnichenko, Anton Schwaighofer, Noel C. F. Codella, Maria Teodora Wetscherek, Klaus H. Maier-Hein, Panagiotis Korfiatis, Valentina Salvatelli, Javier Alvarez-Valle, Fernando P\'erez-Garc\'ia