arXiv:2609.14735v1 Announce Type: cross
Abstract: Deep Learning (DL)-based channel estimation has shown high accuracy and low latency in terrestrial 5G NR, but Low Earth Orbit (LEO) Non-Terrestrial N...
By Miguel Camelo Botero, Nina Slamnik-Krije\v{s}torac, Johann Marquez-Barja
arXiv:2609.13916v1 Announce Type: new
Abstract: We present North Small Translate, an open-weight, LLM-based machine translation (MT) model with instruction-following capabilities built on the same fo...
By Tom Kocmi, Alexandre B\'erard, Phil Blunsom, Samuel Cahyawijaya, Shaun Cassini, Nicholas Frosst, Ona de Gibert, Aidan Gomez, Nithya Govindarajan, Shun Kiyono, Olivia Lasche, Lawrence Rogers, Kelly Marchisio, Nikita Moghe, Yash More, Camila Moran-Hidalgo, Yiyang Nan, Michael Sachs, Trisha Starostina, Daan van Stigt, Spencer Rarrick, Sebastian Vincent, Ivan Zhang
arXiv:2609.14815v1 Announce Type: cross
Abstract: This paper introduces a novel framework for Regularized Multivariate Functional Principal Component Analysis (ReMFPCA) via Functional Singular Value...
By Yue Zhao, Hossein Haghbin, Rebecca Sanders, Mehdi Maadooliat
arXiv:2609.13737v1 Announce Type: new
Abstract: As large language models (LLMs) are increasingly deployed, the generation of harmful content has become a critical safety concern. Existing safeguards...
By Hanling Wang, Chenlong Wei, Ling Xu, Hanyan Niu, Qi Cao, Shizhou Huang, Yang Yang, Xiaohui Zhu, Yao Zhu
arXiv:2609.14246v1 Announce Type: cross
Abstract: In wireless federated learning (FL), data heterogeneity and multiple local updates induce client drift, degrading model convergence. It is further af...
By Changheng Wang, Xianchao Zhang, Zhiqing Wei, Lingzhu Zhao, Zhongming Yang, Zhiyong Feng
arXiv:2609.14146v1 Announce Type: cross
Abstract: Vision-language-action (VLA) deployment can reduce inference latency while changing closed-loop task behavior. We evaluate HuggingFaceVLA/smolvla_lib...
By Rafiqul Islam
arXiv:2609.14261v1 Announce Type: cross
Abstract: Recent robot learning paradigms increasingly rely on large offline datasets of robotic interactions to train control policies. Expressive generative...
By Prajwal Koirala, Mark Campbell
arXiv:2609.13615v1 Announce Type: new
Abstract: For our submission to the WMT26 Creole Language Translation Shared Task, we focus on machine translation (MT) models for Pacific creoles: Tok Pisin, Bi...
By Rapha\"el Merx, Nick Thieberger, Ekaterina Vylomova
arXiv:2609.14648v1 Announce Type: new
Abstract: Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often...
By Ziyi Zhu, Daniel R. Cahn, Thomas D. Hull, Caitlin A. Stamatis, Olivier Tieleman, Guilherme B. Freire, Jinghong Chen
arXiv:2609.13947v1 Announce Type: cross
Abstract: In-sensor computing reduces the cost of transmitting high-resolution image data by performing early-stage processing near the sensor. However, the lo...
By Chengwei Zhou, Abu Masum, Xuming Chen, Mehran Moghadam, Sreetama Sarkar, Arnab Sanyal, Md Abdullah-Al Kaiser, M. Hassan Najafi, Sercan Aygun, Gourav Datta
arXiv:2609.15130v1 Announce Type: cross
Abstract: woma is a real-time foundation model for gastrointestinal endoscopy: a network trained without labels on about a million endoscopy frames, from which...
By Thang Tran, Lan Dang
arXiv:2609.13592v1 Announce Type: cross
Abstract: GPU memory bandwidth and capacity limit throughput in large language model (LLM) inference. The GPU memory system consists of a primary tier of high-...
By Anish Saxena, Jae Hyung Ju, Hritvik Taneja, Po-An Tsai, Aamer Jaleel, Christos Kozyrakis, Moinuddin Qureshi
arXiv:2510.19266v3 Announce Type: replace
Abstract: State-space models (SSMs) have emerged as promising alternatives to Transformers for sequence modeling. However, training competitive SSMs from scr...
By Penghao Wang, Yuhao Zhou, Mengxuan Wu, Panpan Zhang, Zhangyang Wang, Kai Wang
The paper introduces self‑orchestrating language models that annotate semantic dependence—identifying which tokens rely on others—to guide efficient inference. By leveraging these annotations, the authors design runtimes that parallelize autoregressive decoding, evict intermediate context, or determine denoising orders, achieving Pareto‑optimal quality‑efficiency trade‑offs. Three systems—PASTA, TIP, and Planned Diffusion—demonstrate these techniques for parallel decoding, memory‑efficient reasoning, and efficient discrete diffusion, respectively.
By Tian Jin
arXiv:2609.13271v1 Announce Type: cross
Abstract: Medical image segmentation needs diverse training data, but hospitals hold complementary scans they cannot share for privacy and regulatory reasons....
By Armaghan Butt, Shuya Feng, Qing Tian
The paper introduces WaterKron, a method that integrates two-sided GPTQ with row- and column-dependent waterfilling scales and entropy coding for post‑training quantization. It derives a high‑rate distortion measure relative to the full Hessian, introducing a Kronecker‑Hessian mismatch factor Φ that quantifies the distortion penalty of using a Kronecker approximation. Minimizing Φ leads to a Gaussian covariance‑fitting problem solved via classical flip‑flop updates, yielding a FlipFlop Hessian that empirically improves KL divergence and perplexity compared to other Hessian choices.
By Johann Birnick, Rayan Saab
arXiv:2511.08092v2 Announce Type: replace-cross
Abstract: We challenge the conventional view of neural network pruning as solely a compression technique, demonstrating that one-shot magnitude pruning...
By Julian Irigoyen, Arthur S\"ohler, Andreas S{\o}eborg Kirkedal
arXiv:2609.13636v1 Announce Type: cross
Abstract: Privacy-preserving inference via Torus Fully Homomorphic Encryption (TFHE) provides strong protection for sensitive data in outsourced deep learning...
By Mahmoud Y. M. Yassin, Mahmoud AbdelHafeez Sayed, Mostafa Taha
The paper discusses tensorizing neural networks by reshaping dense weight matrices into higher-order tensors and approximating them with low-rank tensor network decompositions. This approach offers promising model compression and introduces bond indices that create new latent spaces, potentially enhancing interpretability. Despite encouraging empirical results, tensorized neural networks remain underused, and the authors call for more research to address practical scaling and adoption challenges.
By Safa Hamreras, Sukhbinder Singh, Rom\'an Or\'us
arXiv:2609.13232v1 Announce Type: new
Abstract: A robot pruning trees needs two facts per pixel: whether it belongs to a tree, and its distance. Both are usually obtained via task heads attached to a...
By Yida Lin, Bing Xue, Mengjie Zhang, Sam Schofield, Richard Green