The paper presents physics‑guided machine‑learning models that predict defect formation energies and zero‑phonon lines (ZPLs) for point defects in semiconductors, aiming to replace costly density‑functional theory (DFT) calculations in the prescreening stage of high‑throughput workflows. Using ridge, kernel ridge, and multilayer perceptron models with three descriptors, the authors achieve mean absolute errors of 0.437 eV for formation energies and 0.202 eV for ZPLs on vacancies and substitutions in 4H‑SiC, while interstitials show larger errors (1.101 eV and 0.230 eV). These results demonstrate that the models can effectively accelerate defect screening, potentially obviating the need for expensive DFT relaxations in many cases.
By Paul Karlsson, Joel Davidsson, Rickard Armiento
arXiv:2609.13692v1 Announce Type: cross
Abstract: LLM serving reuses KV cache by exact prefix match, so when a prompt is assembled from a set of reusable pieces -- retrieved passages, tool definition...
By Rong He
arXiv:2609.14060v1 Announce Type: cross
Abstract: Quantization is one of the default deployment paths for open-weight LLM agents, but it is not behavior-preserving: an adversary can release a full-pr...
By Xiaoqun Liu, Qiben Yan
arXiv:2609.14187v1 Announce Type: cross
Abstract: Memory-efficient feature representations are increasingly important in machine learning settings where storage, transmission cost, bandwidth, or priv...
By John Cartmell, Mihaela Cardei, Ionut Cardei
arXiv:2609.13947v1 Announce Type: cross
Abstract: In-sensor computing reduces the cost of transmitting high-resolution image data by performing early-stage processing near the sensor. However, the lo...
By Chengwei Zhou, Abu Masum, Xuming Chen, Mehran Moghadam, Sreetama Sarkar, Arnab Sanyal, Md Abdullah-Al Kaiser, M. Hassan Najafi, Sercan Aygun, Gourav Datta
The paper introduces ISER, an isolation-based method for unsupervised tabular anomaly detection that uses hypersphere radii to encode local density and maintains linear time and constant space complexity. ISER builds ensemble representations where smaller radii indicate dense regions and larger radii indicate sparse regions, and it employs a similarity-based scoring method that compares these representations to a theoretical anomaly reference pattern. Experiments on 20 real-world datasets show that ISER outperforms 12 state‑of‑the‑art methods, including an enhanced Isolation Forest.
By Yang Cao, Sikun Yang, Hao Tian, Kai He, Lianyong Qi, Ming Liu, Yujiu Yang, Hong-Kun Zhang
arXiv:2609.14715v1 Announce Type: new
Abstract: We scale our conventional sub-150M pretraining recipe from 53.5M to 109.7M parameters, holding the method fixed (Qwen3-style decoder with grouped-query...
By Dushyant Rajput (AltSlate Labs LLP), Nirdesh Chauhan (AltSlate Labs LLP), Siddharth Kosaraju (AltSlate Labs LLP)
arXiv:2512.12767v2 Announce Type: replace-cross
Abstract: Training recurrent neuronal networks consisting of excitatory (E) and inhibitory (I) units with additive noise for working memory computation...
By Thiparat Chotibut, Oleg Evnin, Weerawit Horinouchi
arXiv:2609.13232v1 Announce Type: new
Abstract: A robot pruning trees needs two facts per pixel: whether it belongs to a tree, and its distance. Both are usually obtained via task heads attached to a...
By Yida Lin, Bing Xue, Mengjie Zhang, Sam Schofield, Richard Green
Lightning Weave is a post‑training framework that composes independently learned accuracy and efficiency capabilities of reasoning models into a single student model via on‑policy distillation. It extracts policy shifts from pre‑trained models, aligns log‑ratio shifts at shared token states, and uses Tilted‑Target DOPD to create stable learning targets, allowing anchor pairs to score cached trajectories once without running multiple live models. Experiments on mathematics and code benchmarks show significant gains, such as raising Qwen3.5‑4B’s HMMT 2025 accuracy from 59.2% to 64.0% while reducing response tokens by 10.7%, and improving LiveCodeBench v5 accuracy from 41.7% to 54.2% with 9.6% fewer tokens.
By Yecheng Wu, Song Han, Han Cai
The paper compares Complement Naive Bayes (NB) with zero‑shot and few‑shot large language models (LLMs) across a wide range of model sizes and text classification tasks. NB outperforms LLMs when labeled data is available, achieving comparable accuracy to large LLMs while running thousands of samples per second on a CPU. In zero‑data sentiment settings, LLMs still dominate, but NB remains the best choice for resource‑constrained HPC practitioners, and the authors provide a Kubernetes Helm operator to automate model selection.
By Mohammad Firas Sada, Dmitry Mishin, John Graham, Seungmin Kim, Mahidhar Tatineni, Frank W\"urthwein
The paper discusses tensorizing neural networks by reshaping dense weight matrices into higher-order tensors and approximating them with low-rank tensor network decompositions. This approach offers promising model compression and introduces bond indices that create new latent spaces, potentially enhancing interpretability. Despite encouraging empirical results, tensorized neural networks remain underused, and the authors call for more research to address practical scaling and adoption challenges.
By Safa Hamreras, Sukhbinder Singh, Rom\'an Or\'us
arXiv:2609.15457v1 Announce Type: new
Abstract: Vision-Language Models (VLMs) enable autonomous GUI navigation, but agents still struggle to process and learn from dense, continuous visual histories....
By Shengjie Jin, Zelong Sun, Hengbo Xu, Yanbiao Ma, Zhiwu Lu
The paper predicts single‑sequence llama.cpp throughput from GGUF metadata using roofline‑shaped predictors with quantization‑specific scale factors. Experiments on 318 measurements across 53 host‑file configurations on two Apple M4 Max systems and an NVIDIA RTX 5080 show that an active‑parameter decode model achieves significantly lower mean absolute percentage errors compared to models that use total parameters. The study also finds that low‑bit model ladders alter runtime ordering and that GGUF structure improves predictions, though fitted efficiencies vary across systems.
By Xinyu Qiu, Chuhong Xu, Bo Su, Ziyao Chen, Ruiyang Xu, Shimeng Dai
El Agente Potente is an agentic system that integrates typed execution graphs and a coding mode to facilitate machine‑learning interatomic potential (MLIP) driven atomistic simulations. Typed execution graphs offer structured, provenance‑aware workflows where large language models handle planning and routing while deterministic Python code performs scientific computation and validation. The coding agent builds customized workflows for tasks needing procedural flexibility, invoking existing Potente functions for supported calculations. The system is demonstrated across materials discovery, energy‑landscape exploration, adsorption, and catalytic reaction workflows, with benchmarks on reproducibility and LLM token cost.
By Tsz Wai Ko, Jiaru Bai, Thomas Swanick, Yeonghun Kang, Changhyeok Choi, Angelina Qihong Jiang, Aiwei Yin, Varinia Bernales, Al\'an Aspuru-Guzik
arXiv:2609.14648v1 Announce Type: new
Abstract: Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often...
By Ziyi Zhu, Daniel R. Cahn, Thomas D. Hull, Caitlin A. Stamatis, Olivier Tieleman, Guilherme B. Freire, Jinghong Chen
CrossDistill is a trajectory-level hybrid few-step distillation framework for diffusion models that balances quality and diversity by splitting the sampling trajectory at a crossover point. The high-noise interval uses a trajectory-preserving objective to maintain global mode coverage, while the low-noise interval applies a distribution-matching objective to sharpen local details, with the two stages coupled through the crossover state. This noise-level scheduling policy, demonstrated on text-to-video and image-to-video diffusion models, expands the few-step quality-diversity frontier by preserving seed-level variation while achieving competitive visual fidelity.
By Yuxi Liu, Haoyu Li, Yixiang Cai, Tengxu Sun, Zekun Zhang, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kun Yuan, Kai Zhang
The paper introduces self‑orchestrating language models that annotate semantic dependence—identifying which tokens rely on others—to guide efficient inference. By leveraging these annotations, the authors design runtimes that parallelize autoregressive decoding, evict intermediate context, or determine denoising orders, achieving Pareto‑optimal quality‑efficiency trade‑offs. Three systems—PASTA, TIP, and Planned Diffusion—demonstrate these techniques for parallel decoding, memory‑efficient reasoning, and efficient discrete diffusion, respectively.
By Tian Jin
arXiv:2609.15743v1 Announce Type: new
Abstract: Automatic speech recognition (ASR) systems, trained on paired speech-text data, have been improved by leveraging language models (LMs) trained on text-...
By Hayato Futami, Tatsuya Kawahara
MANE is a distributed inference framework that uses a multi‑path tail architecture to allow dynamic accuracy–throughput trade‑offs during edge onloading of deep neural networks. It introduces a novel multi‑path model, a three‑stage training scheme with Joint Head Network Distillation loss, and a hysteresis‑based scheduler with an equitable device‑fallback policy. The system achieves over 80% SLO satisfaction and 6pp higher accuracy than on‑device alternatives while supporting up to 40 concurrent devices.
By Sokratis Nikolaidis, Stylianos I. Venieris, Leonidas Malachias, Iakovos S. Venieris