Inference efficiency

Quantization, distillation, pruning and serving work aimed at the same accuracy for less memory, latency and money.

6,032 stories · RSS feed

arXiv AI
Sep 15

Diversified and Perceptible Counterfactual Examples Leveraging Expert Knowledge

The paper introduces DiCEf, a variant of the DiCE counterfactual example generator that incorporates expert knowledge expressed as a fuzzy linguistic vocabulary. By embedding continuous data into this linguistic domain, DiCEf personalises counterfactual explanations to be linguistically perceptible while maintaining minimal cost, sparsity, and diversity. Experimental results on a real‑world dataset demonstrate that this approach yields semantically meaningful counterfactuals for the explainee.

By Akram Bensalem (IMT Atlantique - INFO), Fahima Djelil (Lab-STICC\_MOTEL, IMT Atlantique - INFO), Marie-Jeanne Lesot (IMT Atlantique - INFO, Lab-STICC, Lab-STICC\_MOTEL), Gr{\'e}gory Smits (IMT Atlantique - INFO, Lab-STICC, Lab-STICC\_MOTEL)
arXiv Machine Learning
Sep 15

Prescreening Point Defects in Semiconductors With Machine Learning

The paper presents physics‑guided machine‑learning models that predict defect formation energies and zero‑phonon lines (ZPLs) for point defects in semiconductors, aiming to replace costly density‑functional theory (DFT) calculations in the prescreening stage of high‑throughput workflows. Using ridge, kernel ridge, and multilayer perceptron models with three descriptors, the authors achieve mean absolute errors of 0.437 eV for formation energies and 0.202 eV for ZPLs on vacancies and substitutions in 4H‑SiC, while interstitials show larger errors (1.101 eV and 0.230 eV). These results demonstrate that the models can effectively accelerate defect screening, potentially obviating the need for expensive DFT relaxations in many cases.

By Paul Karlsson, Joel Davidsson, Rickard Armiento
arXiv AI
Sep 15

Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference

The paper introduces self‑orchestrating language models that annotate semantic dependence—identifying which tokens rely on others—to guide efficient inference. By leveraging these annotations, the authors design runtimes that parallelize autoregressive decoding, evict intermediate context, or determine denoising orders, achieving Pareto‑optimal quality‑efficiency trade‑offs. Three systems—PASTA, TIP, and Planned Diffusion—demonstrate these techniques for parallel decoding, memory‑efficient reasoning, and efficient discrete diffusion, respectively.

By Tian Jin
arXiv AI
Sep 15

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives

IBBench-Light is a paired evaluation framework that tests language models on both executing procedures and reading text from the same external record. The benchmark uses twelve semantic bases to generate 144 matched pairs per model, with four instruction‑quantized models producing 1,152 greedy responses. Metrics such as Paired Exact‑Contract Accuracy (PECA) reveal that models like Qwen achieve high success on individual prompts but only 97 complete pairs, highlighting the importance of paired evaluation.

By Kainan Zhou, Gangzhen Qian, Zhaoyi Li, Hang Xiao
arXiv Machine Learning
Sep 15

Isolation-based Spherical Ensemble Representations for Tabular Anomaly Detection

The paper introduces ISER, an isolation-based method for unsupervised tabular anomaly detection that uses hypersphere radii to encode local density and maintains linear time and constant space complexity. ISER builds ensemble representations where smaller radii indicate dense regions and larger radii indicate sparse regions, and it employs a similarity-based scoring method that compares these representations to a theoretical anomaly reference pattern. Experiments on 20 real-world datasets show that ISER outperforms 12 state‑of‑the‑art methods, including an enhanced Isolation Forest.

By Yang Cao, Sikun Yang, Hao Tian, Kai He, Lianyong Qi, Ming Liu, Yujiu Yang, Hong-Kun Zhang
arXiv AI
Sep 15

OCT-FedSIR: Toward Trustworthy Federated Ophthalmic Learning under Annotation Noise

OCT-FedSIR is a reliability‑aware spectral framework designed for federated learning of OCT image classification in the presence of client‑dependent annotation noise and heterogeneous data distributions. It integrates class‑balanced spectral estimation, logit adjustment, complementary spectral descriptors, selective spectral relabeling, and noise‑aware federated optimization. Across 117 experimental conditions on three datasets, OCT‑FedSIR achieved a mean accuracy of 86.73%, outperforming RoFL (79.94%) and FedCorr (78.75%) and successfully identifying and correcting corrupted annotations with high precision.

By Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Rashadul H. Badhon, Behafarin Emam, Sally S. Y. Ong, Atalie C. Thompson, Theodore Leng, Ahmed Arafa, Jennifer I. Lim, Minhaj Nur Alam