arXiv:2609.13285v1 Announce Type: cross
Abstract: The KV cache is a primary bottleneck for Transformer decoding: its memory footprint and cache-read traffic grow with sequence length. Grouped-query a...
By Vishesh Tripathi, Abhay Kumar, Ramsha Khan
arXiv:2609.15972v1 Announce Type: cross
Abstract: As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the p...
By Zixuan Wang, Yufan Zhou, Jinzhou Tang, Xinle Yu, Chengjun Wu, Lyumanshan Ye, Zhaoxiang Feng, Letian Peng, Adyasha Patra, Fan Bai, Enze Ma, Zhengding Hu, Jianyang Gu, Zhao Wang, Yufei Ding, Jingbo Shang, Tianmin Shu, Zhiting Hu, Zhen Wang
The paper introduces THEMIS, a stage-aware repair workflow that externalizes the requirement-to-repair process by generating semantic interpretations, a runtime requirement-code graph, graph-derived developer guidance, retained repair rationale and patches, and post-edit audit records. A retrospective audit of 300 SWE-bench Lite cases shows that these artifacts enable cross-stage inspection, with a complete developer rationale available for 288 cases and 214 cases retaining a full audited field set. The retained records also allow systematic measurement of cross-stage correspondence, revealing high recurrence of target symbols across rationales and patches, and a preliminary improvement in resolving cases compared to a direct same-input condition.
By Zewen Tao, Shin-nosuke Ishikawa
arXiv:2609.13737v1 Announce Type: new
Abstract: As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards...
By Hanling Wang, Chenlong Wei, Ling Xu, Hanyan Niu, Qi Cao, Shizhou Huang, Yang Yang, Xiaohui Zhu, Yao Zhu
arXiv:2511.08092v2 Announce Type: replace-cross
Abstract: We challenge the conventional view of neural network pruning as solely a compression technique, demonstrating that one-shot magnitude pruning...
By Julian Irigoyen, Arthur S\"ohler, Andreas S{\o}eborg Kirkedal
arXiv:2609.13199v1 Announce Type: new
Abstract: Knowledge distillation aims to improve the performance of lightweight student models by transferring knowledge from larger and more powerful teacher mo...
By Dawen Jiang, Zhishu Shen, Zeyu Liu, Tiehua Zhang
El Agente Potente is an agentic system that integrates typed execution graphs and a coding mode to facilitate machine‑learning interatomic potential (MLIP) driven atomistic simulations. Typed execution graphs offer structured, provenance‑aware workflows where large language models handle planning and routing while deterministic Python code performs scientific computation and validation. The coding agent builds customized workflows for tasks needing procedural flexibility, invoking existing Potente functions for supported calculations. The system is demonstrated across materials discovery, energy‑landscape exploration, adsorption, and catalytic reaction workflows, with benchmarks on reproducibility and LLM token cost.
By Tsz Wai Ko, Jiaru Bai, Thomas Swanick, Yeonghun Kang, Changhyeok Choi, Angelina Qihong Jiang, Aiwei Yin, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv:2609.13154v1 Announce Type: new
Abstract: Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et...
By Shamin Chokshi
arXiv:2510.19266v3 Announce Type: replace
Abstract: State-space models (SSMs) have emerged as promising alternatives to Transformers for sequence modeling. However, training competitive SSMs from scr...
By Penghao Wang, Yuhao Zhou, Mengxuan Wu, Panpan Zhang, Zhangyang Wang, Kai Wang
arXiv:2609.14815v1 Announce Type: cross
Abstract: This paper introduces a novel framework for Regularized Multivariate Functional Principal Component Analysis (ReMFPCA) via Functional Singular Value...
By Yue Zhao, Hossein Haghbin, Rebecca Sanders, Mehdi Maadooliat
The paper predicts single‑sequence llama.cpp throughput from GGUF metadata using roofline‑shaped predictors with quantization‑specific scale factors. Experiments on 318 measurements across 53 host‑file configurations on two Apple M4 Max systems and an NVIDIA RTX 5080 show that an active‑parameter decode model achieves significantly lower mean absolute percentage errors compared to models that use total parameters. The study also finds that low‑bit model ladders alter runtime ordering and that GGUF structure improves predictions, though fitted efficiencies vary across systems.
By Xinyu Qiu, Chuhong Xu, Bo Su, Ziyao Chen, Ruiyang Xu, Shimeng Dai
arXiv:2609.15130v1 Announce Type: cross
Abstract: woma is a real-time foundation model for gastrointestinal endoscopy: a network trained without labels on about a million endoscopy frames, from which...
By Thang Tran, Lan Dang
arXiv:2609.14735v1 Announce Type: cross
Abstract: Deep Learning (DL)-based channel estimation has shown high accuracy and low latency in terrestrial 5G NR, but Low Earth Orbit (LEO) Non-Terrestrial N...
By Miguel Camelo Botero, Nina Slamnik-Krije\v{s}torac, Johann Marquez-Barja
arXiv:2602.08370v2 Announce Type: replace-cross
Abstract: Realizing versatile and human-like performance in high-demand sports like badminton remains a formidable challenge for humanoid robotics. Unl...
By Yeke Chen, Shihao Dong, Xiaoyu Ji, Jingkai Sun, Zeren Luo, Liu Zhao, Jiahui Zhang, Wanyue Li, Ji Ma, Bowen Xu, Yimin Han, Xuanyi Li, Yudong Zhao, Liyun Li, Peng Lu
The paper compares Complement Naive Bayes (NB) with zero‑shot and few‑shot large language models (LLMs) across a wide range of model sizes and text classification tasks. NB outperforms LLMs when labeled data is available, achieving comparable accuracy to large LLMs while running thousands of samples per second on a CPU. In zero‑data sentiment settings, LLMs still dominate, but NB remains the best choice for resource‑constrained HPC practitioners, and the authors provide a Kubernetes Helm operator to automate model selection.
By Mohammad Firas Sada, Dmitry Mishin, John Graham, Seungmin Kim, Mahidhar Tatineni, Frank W\"urthwein
arXiv:2511.03728v2 Announce Type: replace
Abstract: On-device AI agents offer the potential for personalized, low-latency assistance, but their deployment is fundamentally constrained by limited memo...
By Sanidhya Vijayvargiya, Rahul Lokesh
arXiv:2609.13916v1 Announce Type: new
Abstract: We present North Small Translate, an open-weight, LLM-based machine translation (MT) model with instruction-following capabilities built on the same fo...
By Tom Kocmi, Alexandre B\'erard, Phil Blunsom, Samuel Cahyawijaya, Shaun Cassini, Nicholas Frosst, Ona de Gibert, Aidan Gomez, Nithya Govindarajan, Shun Kiyono, Olivia Lasche, Lawrence Rogers, Kelly Marchisio, Nikita Moghe, Yash More, Camila Moran-Hidalgo, Yiyang Nan, Michael Sachs, Trisha Starostina, Daan van Stigt, Spencer Rarrick, Sebastian Vincent, Ivan Zhang
arXiv:2609.14261v1 Announce Type: cross
Abstract: Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative...
By Prajwal Koirala, Mark Campbell
arXiv:2609.13592v1 Announce Type: cross
Abstract: GPU memory bandwidth and capacity limit throughput in large language model (LLM) inference. The GPU memory system consists of a primary tier of high-...
By Anish Saxena, Jae Hyung Ju, Hritvik Taneja, Po-An Tsai, Aamer Jaleel, Christos Kozyrakis, Moinuddin Qureshi
arXiv:2609.13636v1 Announce Type: cross
Abstract: Privacy-preserving inference via Torus Fully Homomorphic Encryption (TFHE) provides strong protection for sensitive data in outsourced deep learning...
By Mahmoud Y. M. Yassin, Mahmoud AbdelHafeez Sayed, Mostafa Taha