The paper introduces Skeletal Prototypes on Iterative Nerve Expansions (SPINE), a prototype reduction method that represents each class as an embedded 1‑complex rather than a finite set of points. SPINE constructs its initial edge set from a class‑conditional Mapper graph, then refines vertex positions under a classification objective, allowing observations to be assigned to the nearest complex. Evaluated on seventeen benchmark datasets with stratified 10‑fold cross‑validation, SPINE achieves the highest mean accuracy and best average rank among seven competing methods, showing significant improvements over five of them and competitive performance across varying prototype budgets.
By Jordan Eckert, Henry Schenck
The paper "LLM Inference in a Flash!" proposes an integer‑only quantization scheme and a dictionary‑based KV cache compression technique to enable large language model inference on compute‑in‑flash (CIF) devices. By eliminating floating‑point operations and reducing KV cache traffic through sparse dictionary coding, the authors achieve minimal accuracy loss while cutting dynamic KV cache traffic by 15× on Llama‑3.1‑8B and Qwen‑2.5‑7B models.
By Sebastian Zhao, Minseo Kim, Coleman Hooper, Luca Manolache, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami
arXiv:2609.17475v1 Announce Type: new
Abstract: Capable open-weight models make local coding and reasoning attractive, but their context and execution state strain laptop memory. We present JustFit,...
By Yuhua Chen
arXiv:2609.16864v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models have achieved impressive performance in quasi-static manipulation, but struggle in dynamic manipulation tasks bec...
By Zhenyang Feng, Jimin Heo, Erik B. Sudderth, Unnat Jain
arXiv:2609.17193v1 Announce Type: new
Abstract: Large language model (LLM)-powered agentic AI services increasingly demand low-latency inference, motivating the deployment of LLMs across distributed...
By Zhen Li, Jun Cai, Haoran Gao, An Li, Tan Li
arXiv:2609.16245v1 Announce Type: new
Abstract: Long-horizon scientific discovery requires agents to alternate between exploration, disciplined execution, and critical reassessment as evidence change...
By Vincent Karpf, Joseph Reth, Eike Gerhardt, Audrey Wang, Anna Butz, Jiehao Xing, Jialing Song, Larry Callahan
arXiv:2609.17474v1 Announce Type: cross
Abstract: Large language model (LLM) distillation aims to transfer the capabilities of a powerful teacher to a smaller student. Direct imitation, however, can...
By Haichen Hu, Yuheng Zhang, David Simchi-Levi
arXiv:2507.01927v3 Announce Type: replace
Abstract: While CNNs and ViTs dominate vision architectures, all-MLP models offer a structurally simpler alternative whose patch-independent processing is na...
By Zhentan Zheng
arXiv:2609.16065v1 Announce Type: cross
Abstract: Human activity recognition (HAR) is usually framed as gradient-based training of neural networks. Agentic Heuristic Learning (AHL) Studio explores a...
By Siyu Yuan, He Zhang, Sizhen Bian, Bin Guo
arXiv:2609.16874v1 Announce Type: new
Abstract: Beyond model inference, the decoding stage, which converts raw network outputs into task-level representations, constitutes a significant portion of th...
By Carmelo Scribano, Filippo Muzzini, Nedyalko Prisadnikov, Mohammad Mahdi, Yuqian Fu, Giorgia Franchini, Danda Pani Paudel, Marko Bertogna, Luc Van Gool
arXiv:2609.16689v1 Announce Type: new
Abstract: Large-scale vision-language models (VLM) such as CLIP enable strong open-vocabulary reasoning, yet deploying these capabilities on resource-constrained...
By Jinwoo Jeon, GyuYeop Do, Yubin Lim, Nam-Joon Kim, Hyun Gon Ryu, Hyuk-Jae Lee, Byung-Jun Lee
arXiv:2609.16788v1 Announce Type: new
Abstract: Noise2Noise (N2N) trains denoisers on pairs of independently corrupted observations, eliminating clean references. We stress-test two natural conjectur...
By Dingyan Shang, Zhenyu Xu, Youting Wang, Bonan Shen, Bowen Liu
arXiv:2607.09709v2 Announce Type: replace
Abstract: Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact. We study th...
By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv:2609.17284v1 Announce Type: new
Abstract: Statistical heterogeneity limits federated learning when a single global classifier cannot represent client-specific label distributions. In this work,...
By Polycarpo Souza Neto, Jos\'e Mairton Barros da Silva J\'unior, Charles Casimiro Cavalcante
The study investigates how pruning affects large language models (LLMs) used for smart‑home tool calling. Researchers examined four LLMs—dense Transformer, dense hybrid, and mixture‑of‑experts (MoE) architectures—using depth, width, hybrid, and expert pruning, followed by supervised fine‑tuning. They evaluated over 19,500 instances from three smart‑home datasets, analyzing not only overall accuracy but also degradation in action components (operation, device, argument, value) and task complexity, finding that dense models suffer sharp performance drops after a narrow safe pruning range, while MoE models tolerate more pruning; aggressive pruning also leads to over‑refusal and loss of grounded specificity.
By Congjing Zhang, Vashishtha Patil, Henning Lange, Usman Aleem
arXiv:2604.06036v4 Announce Type: replace-cross
Abstract: Continuous inference over concurrent video streams imposes substantial compute and memory demands on vision-language model (VLM) serving. Str...
By Yulin Zou, Wenyan Chen, Yan Chen, Anya Rajan, JooYoung Park, Shivaraman Nitin, Luo Tao, Francisco Romero, Dmitrii Ustiugov
arXiv:2609.16450v1 Announce Type: cross
Abstract: Diffusion large language models (dLLMs) offer a promising parallel decoding paradigm as an alternative to autoregressive generation through iterative...
By Lixuan Wei, Wei Zhou, Jianwen Wu, Yipeng Shen, Meiling Wang, Haoran You
The paper introduces EBL, an Efficient Broad Learning framework designed for distributed adaptive harmonic estimation in power grids affected by electric vehicle charging. It leverages a quantised FPGA implementation to provide high‑accuracy, half‑cycle input harmonic predictions with ultra‑low latency, outperforming existing FPGA methods by 17.4×. The online transfer learning component enables rapid adaptation across multiple charging scenarios, while bespoke quantisation and sparsity reduce resource usage to just 5.9% of the LUTs on a Zynq Ultrascale+ FPGA, compared to 82% of the state‑of‑the‑art accelerator.
By Changhong Li, Georgios Floros, Biswajit Basu, Shreejith Shanker
arXiv:2609.16338v1 Announce Type: new
Abstract: Ternary Large Language Models (LLM) store every weight as one of three symbols $\{-1,0,+1\}$, so the cost of a ternary model is conventionally referenc...
By Evangelos Georganas, Alexander Heinecke, Pradeep Dubey
arXiv:2609.16255v1 Announce Type: cross
Abstract: We present an efficient method to distill reasoning capabilities into compact video-language models (VLMs) for video question answering (VideoQA). Ou...
By Mantek Singh, Jeshwanth Challagundla, Siddharth Raina, Jasmin Jarsania