Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

5,834 stories · RSS feed

arXiv Machine Learning
Sep 18

A Learning Algorithm for Threshold Boolean Networks with Prescribed Fixed Points

The paper introduces a learning algorithm for threshold Boolean networks (TBNs) that can infer networks with a specified set of fixed points. The method uses a custom differentiable loss function to enforce fixed point preservation, penalize spurious attractors, encourage binary outputs, and promote sparsity via L1 regularization. When applied to the FOS-GRN model of Arabidopsis thaliana, the algorithm perfectly reconstructed all 10 desired fixed points in 5 of 30 runs and on average recovered 8.53 ± 0.90 correct fixed points without any spurious attractors, outperforming standard approaches such as the Perceptron and Logistic Regression.

By Gonzalo A. Ruz
arXiv Machine Learning
Sep 18

Radio Frequency Detection and Classification of Microplastics in Water

The study introduces a machine learning–assisted radio‑frequency dielectric spectroscopic cytometry platform for detecting and classifying micro‑plastic particles in water. Eight types of 10 µm nominal‑diameter micro‑plastics were distinguished in deionized water using RF scattering parameters measured at 0.2–9 GHz, achieving macro‑average F1‑scores, precision, and recall above 0.71. The method also maintained PET classification performance in saline media up to 6.6 % sea salt, demonstrating feasibility for rapid, single‑particle micro‑plastic identification in aqueous environments.

By Jaden Tolbert, Md Saiful Islam, Pingshan Wang
arXiv Computer Vision
Sep 18

AI or Real: Detecting Partially Altered Videos Under Resource-Constrained Environments

The paper introduces a lightweight full-frame detector for partially manipulated AI-generated videos, suitable for edge deployment without face-detection preprocessing. It distills a DINOv2-Base teacher into a frozen MobileNetV3-Small student using temperature-annealed soft-label transfer, attention-diversity regularization, frame-level supervision, and a residual feature adapter. The model addresses false positives on legitimate scene cuts and threshold-level miscalibration, achieving an AUC of 0.766 on a 55,393-sample spliced test set while running at 3.65 ms per 16‑frame clip with a 150.4 MB checkpoint.

By Tamoghna Chakraborty, Md Nurul Absur, Sourya Saha, Saptarshi Debroy
arXiv Machine Learning
Sep 18

Federated Learning Framework for Privacy-Preserving Kidney Stone Detection

The paper proposes a Federated Learning framework that integrates an optimized YOLOv8 network for detecting kidney stones in CT images while preserving patient privacy. By enabling multiple medical institutions to collaboratively train a shared model without exchanging patient data, the approach complies with GDPR and HIPAA regulations. Experiments on a distributed CT dataset show a 0.733 mAP@50 and demonstrate fast, real‑time inference suitable for clinical deployment.

By Najiyya Younas, Omar Abdulkader, Yaser Ali Shah, Muhammad Jawad Ikram, Jebran Khan, Amaad Khalil
Hugging Face Trending Papers
Sep 17

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Video DeltaNet (VDN) introduces a hybrid attention mechanism for video diffusion models, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). The design updates memory once per frame, uses separate output projections and learnable gates to balance the two branches, and employs a staged teacher‑alignment recipe to integrate the new pathway into pretrained models. When applied to MiniMax H3, VDN achieves a 14.5× speedup, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs compared to the 50‑step dense baseline.

Hugging Face Trending Papers
Sep 17

On-Demand Attention: Language Models Know When to Recall

The paper introduces On‑Demand Attention (ODA), a local‑first decoding strategy that predicts when a pretrained language model would benefit from global attention. By training only a lightweight recall head, ODA selectively triggers global attention during generation, keeping pretrained weights unchanged and preserving the full key‑value cache for future recall. Experiments on Qwen, Gemma, and hybrid‑attention models show that ODA largely recovers performance lost with local attention while significantly cutting global reads, enabling faster long‑context inference.

Hugging Face Trending Papers
Sep 17

PhGS: Post-Hoc Pruning and Refinement of Single-View Feed-Forward 3D Gaussian Reconstructions

The paper introduces PhGS, a post‑hoc pruning and refinement pipeline for single‑view feed‑forward 3D Gaussian Splatting models. By keeping the base network frozen, it applies importance‑score‑based pruning followed by a lightweight recurrent refinement module to reduce spatial redundancy while restoring image quality. The method is backbone‑agnostic, integrates seamlessly with existing baselines, preserves novel‑view rendering fidelity, achieves significant memory reduction, and allows flexible inference‑time keep ratios.

Hugging Face Trending Papers
Sep 17

Training Neural Networks to Approach the Optimum Bayes Estimator in Dense Multi-Emitter Localization

The paper presents a method for training neural networks on synthetic data to approximate the optimal Bayes estimator for dense emitter localization. By demonstrating that the networks can closely match this theoretical optimum, the authors provide evidence that such training approaches can be effective for high‑throughput, large‑field‑of‑view super‑spatiotemporal resolution single‑molecule localization microscopy (SMLM).

Hugging Face Trending Papers
Sep 17

To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

The paper introduces SwitchSD, an adaptive framework that treats copying as a latent control signal in large language model decoding. By training lightweight probes on internal representations, SwitchSD accurately detects genuine copy intent (AUC > 0.99) and dynamically switches between neural drafting and context-based copying. Experiments on Llama and Qwen models show up to 15 % throughput gains over state‑of‑the‑art baselines such as EAGLE3.

Hugging Face Trending Papers
Sep 17

Local Sparsity Enables Unsupervised LLM Safety Detection

The paper proposes a novel unsupervised safety detection method for large language models that relies on local sparsity in a linear representation space recovered via a sparse autoencoder. By masking SAE neurons based on shared active support among nearby points, the authors develop a locally masked anomaly detection framework with theoretical backing. Experiments across multiple architectures and datasets—including capability‑testing and safety‑specific sets—show that using only 1–2% of SAE neurons and a small amount of out‑of‑distribution data yields near‑optimal safety detection performance.

Hugging Face Trending Papers
Sep 17

FCA-Guided Counterfactual Explanations for Multi-Modal Breast Cancer Diagnosis: A Framework Achieving Perfect Validity with Emergent Sparsity

The paper introduces FCA‑Guided Counterfactual (FCA‑CF) explanations for multi‑modal breast cancer diagnosis, leveraging a Formal Concept Analysis lattice as a hard structural constraint to generate counterfactuals. On the TCGA‑BRCA dataset, FCA‑CF achieves perfect validity (100% prediction flips), the lowest average feature changes (2.37), and competitive proximity (0.900), outperforming four established counterfactual methods. Ablation studies show the lattice constraint and a greedy refinement phase are key to its sparsity and validity.

Hugging Face Trending Papers
Sep 17

DART: Distillation-Aware Reparameterization for Training-Free LoRA Reuse in Few-Step Video Diffusion Models

The paper introduces DART, a training‑free technique that combines low‑rank coordinate transport with target‑schedule response calibration to improve the reuse of LoRA adapters in few‑step video diffusion models. By avoiding source training videos and using forward evaluations, DART raises the joint quality score on a four‑step Wan2.2 target from 0.9029 to 0.9227 and shifts macro functional retention from negative to positive. Component analysis shows that calibration drives most of the quality gains, while coordinate transport adds complementary benefits, and the method demonstrates consistent improvements across additional targets.

Hugging Face Trending Papers
Sep 17

MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration

MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration proposes a new quantization format that inverts the traditional microscaling approach by assigning private exponents to each element and a shared mantissa. The adaptive dual-format MiX-MX inference framework maps this format to a custom accelerator, replacing multipliers with shifters. Evaluations show that 4.5-bit MiX matches or surpasses NVFP4 accuracy on multimodal benchmarks while improving area efficiency by 25% and delivering 2.3–4.5× speedup with 1.4–2.9× energy reduction compared to the Focus accelerator.

arXiv Machine Learning
Sep 17

QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing

QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.

By Shuaiqi Wang, Zinan Lin, Giulia Fanti
arXiv AI
Sep 17

${M}^2$Tok: Multi-head Multi-codebook Discrete Action Tokenization for Vision-Language-Action Models

The paper introduces ${M}^2$Tok, a Multi-head Multi-codebook Action Tokenizer that reduces reconstruction loss in discrete action tokenization for Vision‑Language‑Action models. By decomposing latent action features into multiple heads and assigning independent codebooks to each, the tokenizer expands representational expressivity and improves policy performance. Experiments on RoboTwin, Simpler‑Env, and zero‑shot real‑world tasks show superior reconstruction fidelity and higher success rates compared to prior methods.

By Chunpu Xu, Zhixuan Liang, Yuhao Zhang, Chi-Min Chan, Jessie Wang, Yang Xiao, Mengkang Hu, Xiaokang Yang, Yao Mu
arXiv AI
Sep 17

Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers

The paper introduces a method for offline on‑policy distillation that addresses the problem of imperfect teacher supervision. By training on teacher‑successful problems and measuring changes in token likelihoods on teacher‑failed trajectories, the authors derive a learnability signal that weights the distillation loss. This approach improves performance on mathematical reasoning and code generation tasks while reducing computational cost compared to online distillation.

By Yihao Ai, Weilong Yan
arXiv AI
Sep 17

Objective vs. Search: Decomposing What Makes a Good Tokeniser

The paper introduces two new tokenisation algorithms—BottomUpLL and TopDownComp—to systematically explore the 2x2 design space defined by optimisation objective (compression vs. log‑likelihood) and search procedure (bottom‑up merging vs. top‑down pruning). Experiments across model sizes, vocabularies, and domains show that the search procedure, rather than the objective, consistently yields lower bits‑per‑byte, while no clear pattern emerges on the BLiMP benchmark. These findings clarify how tokeniser design choices influence language‑model performance and provide guidance for constructing tokenisers more principledly.

By Ahmetcan Yavuz, Clara Meister, Tiago Pimentel
arXiv Machine Learning
Sep 17

Similarity Pairing with Energy Mover's Distance for Self-Supervised Pre-Training at the LHC

The paper introduces a data‑driven method for pairing events at the Large Hadron Collider using the energy mover's distance (EMD) to measure similarity, thereby creating augmentation‑free views for self‑supervised pre‑training. By matching distinct events based on EMD, the approach preserves the physics content of each event without handcrafted distortions. Experiments on QCD jets demonstrate that this pairing technique yields semantic jet embeddings with downstream discrimination power comparable to or better than traditional augmentation‑based baselines.

By Ho Fung Tsoi, Dylan Rankin
arXiv Machine Learning
Sep 17

A General Kernel Framework for Non-CND Distance Measures Using |D|-Dimensional Sparse Landmark Embeddings

The paper introduces the Sparse Landmark Embedding (SLE) kernel, a new framework that removes the need for conditionally negative definite (CND) distance measures in kernel methods and Gaussian Processes. By embedding each input into a sparse feature vector using compactly supported bump functions centered at all training points, any standard positive semi-definite (PSD) kernel can be applied in this embedding space, guaranteeing PSD for arbitrary distance measures. The authors provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and show through experiments with geodesic and Wasserstein distances that the SLE kernel matches or surpasses domain-specific baselines in predictive accuracy and uncertainty quantification.

By Marcus M. Noack, Maher B. Alghalayini, Mark D. Risser