arXiv:2609.15743v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems, trained on paired speech-text data, have been improved by leveraging language models (LMs) trained on text-...
By Hayato Futami, Tatsuya Kawahara
arXiv:2609.13692v1 Announce Type: cross
Abstract: LLM serving reuses KV cache by exact prefix match, so when a prompt is assembled from a set of reusable pieces -- retrieved passages, tool definition...
By Rong He
arXiv:2609.15130v1 Announce Type: cross
Abstract: woma is a real-time foundation model for gastrointestinal endoscopy: a network trained without labels on about a million endoscopy frames, from which...
By Thang Tran, Lan Dang
arXiv:2510.19266v3 Announce Type: replace
Abstract: State-space models (SSMs) have emerged as promising alternatives to Transformers for sequence modeling. However, training competitive SSMs from scr...
By Penghao Wang, Yuhao Zhou, Mengxuan Wu, Panpan Zhang, Zhangyang Wang, Kai Wang
arXiv:2609.14735v1 Announce Type: cross
Abstract: Deep Learning (DL)-based channel estimation has shown high accuracy and low latency in terrestrial 5G NR, but Low Earth Orbit (LEO) Non-Terrestrial N...
By Miguel Camelo Botero, Nina Slamnik-Krije\v{s}torac, Johann Marquez-Barja
The paper introduces DiCEf, a variant of the DiCE counterfactual example generator that incorporates expert knowledge expressed as a fuzzy linguistic vocabulary. By embedding continuous data into this linguistic domain, DiCEf personalises counterfactual explanations to be linguistically perceptible while maintaining minimal cost, sparsity, and diversity. Experimental results on a real‑world dataset demonstrate that this approach yields semantically meaningful counterfactuals for the explainee.
By Akram Bensalem (IMT Atlantique - INFO), Fahima Djelil (Lab-STICC\_MOTEL, IMT Atlantique - INFO), Marie-Jeanne Lesot (IMT Atlantique - INFO, Lab-STICC, Lab-STICC\_MOTEL), Gr{\'e}gory Smits (IMT Atlantique - INFO, Lab-STICC, Lab-STICC\_MOTEL)
arXiv:2609.15313v1 Announce Type: cross
Abstract: Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models....
By Daxin Tan, Dehua Tao, Chengxi Deng, Hanlin Zhang, Xiao Chen
arXiv:2609.14815v1 Announce Type: cross
Abstract: This paper introduces a novel framework for Regularized Multivariate Functional Principal Component Analysis (ReMFPCA) via Functional Singular Value...
By Yue Zhao, Hossein Haghbin, Rebecca Sanders, Mehdi Maadooliat
arXiv:2609.13154v1 Announce Type: new
Abstract: Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et...
By Shamin Chokshi
The paper presents physics‑guided machine‑learning models that predict defect formation energies and zero‑phonon lines (ZPLs) for point defects in semiconductors, aiming to replace costly density‑functional theory (DFT) calculations in the prescreening stage of high‑throughput workflows. Using ridge, kernel ridge, and multilayer perceptron models with three descriptors, the authors achieve mean absolute errors of 0.437 eV for formation energies and 0.202 eV for ZPLs on vacancies and substitutions in 4H‑SiC, while interstitials show larger errors (1.101 eV and 0.230 eV). These results demonstrate that the models can effectively accelerate defect screening, potentially obviating the need for expensive DFT relaxations in many cases.
By Paul Karlsson, Joel Davidsson, Rickard Armiento
The paper introduces self‑orchestrating language models that annotate semantic dependence—identifying which tokens rely on others—to guide efficient inference. By leveraging these annotations, the authors design runtimes that parallelize autoregressive decoding, evict intermediate context, or determine denoising orders, achieving Pareto‑optimal quality‑efficiency trade‑offs. Three systems—PASTA, TIP, and Planned Diffusion—demonstrate these techniques for parallel decoding, memory‑efficient reasoning, and efficient discrete diffusion, respectively.
By Tian Jin
arXiv:2609.14762v1 Announce Type: cross
Abstract: Cloud-hosted large language models (LLMs) are increasingly used for root cause analysis (RCA) in AIOps pipelines, but they introduce data privacy ris...
By Rohit Patel, Susil Kumar Mohanty, Jeenal Chaudhary
arXiv:2609.14715v1 Announce Type: new
Abstract: We scale our conventional sub-150M pretraining recipe from 53.5M to 109.7M parameters, holding the method fixed (Qwen3-style decoder with grouped-query...
By Dushyant Rajput (AltSlate Labs LLP), Nirdesh Chauhan (AltSlate Labs LLP), Siddharth Kosaraju (AltSlate Labs LLP)
IBBench-Light is a paired evaluation framework that tests language models on both executing procedures and reading text from the same external record. The benchmark uses twelve semantic bases to generate 144 matched pairs per model, with four instruction‑quantized models producing 1,152 greedy responses. Metrics such as Paired Exact‑Contract Accuracy (PECA) reveal that models like Qwen achieve high success on individual prompts but only 97 complete pairs, highlighting the importance of paired evaluation.
By Kainan Zhou, Gangzhen Qian, Zhaoyi Li, Hang Xiao
arXiv:2609.13636v1 Announce Type: cross
Abstract: Privacy-preserving inference via Torus Fully Homomorphic Encryption (TFHE) provides strong protection for sensitive data in outsourced deep learning...
By Mahmoud Y. M. Yassin, Mahmoud AbdelHafeez Sayed, Mostafa Taha
arXiv:2609.14261v1 Announce Type: cross
Abstract: Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative...
By Prajwal Koirala, Mark Campbell
arXiv:2609.14246v1 Announce Type: cross
Abstract: In wireless federated learning (FL), data heterogeneity and multiple local updates induce client drift, degrading model convergence. It is further af...
By Changheng Wang, Xianchao Zhang, Zhiqing Wei, Lingzhu Zhao, Zhongming Yang, Zhiyong Feng
arXiv:2609.15177v1 Announce Type: new
Abstract: Diffusion language models (dLLMs) promise fast inference by generating multiple tokens in parallel, but suffer severe performance degradation when para...
By Shijian Xu, Andrea Miele, Metod Jazbec, Volker Roth, Eric Nalisnick, Ilija Bogunovic
The paper introduces ISER, an isolation-based method for unsupervised tabular anomaly detection that uses hypersphere radii to encode local density and maintains linear time and constant space complexity. ISER builds ensemble representations where smaller radii indicate dense regions and larger radii indicate sparse regions, and it employs a similarity-based scoring method that compares these representations to a theoretical anomaly reference pattern. Experiments on 20 real-world datasets show that ISER outperforms 12 state‑of‑the‑art methods, including an enhanced Isolation Forest.
By Yang Cao, Sikun Yang, Hao Tian, Kai He, Lianyong Qi, Ming Liu, Yujiu Yang, Hong-Kun Zhang
OCT-FedSIR is a reliability‑aware spectral framework designed for federated learning of OCT image classification in the presence of client‑dependent annotation noise and heterogeneous data distributions. It integrates class‑balanced spectral estimation, logit adjustment, complementary spectral descriptors, selective spectral relabeling, and noise‑aware federated optimization. Across 117 experimental conditions on three datasets, OCT‑FedSIR achieved a mean accuracy of 86.73%, outperforming RoFL (79.94%) and FedCorr (78.75%) and successfully identifying and correcting corrupted annotations with high precision.
By Sina Gholami, Abdulmoneam Ali, Tania Haghighi, Rashadul H. Badhon, Behafarin Emam, Sally S. Y. Ong, Atalie C. Thompson, Theodore Leng, Ahmed Arafa, Jennifer I. Lim, Minhaj Nur Alam