arXiv:2509. 21945v2 Announce Type: replace-cross Abstract: To efficiently tune configuration for better software system performance (e.
By Pengzhou Chen, Hongyuan Liang, Tao Chen
The paper introduces a unified evaluation protocol for robust counterfactual explanations (CFE), testing six robust methods and two baselines across four tabular datasets under eight types of model change. It shows that robustness scores vary by change type and that methods designed for one change family may not transfer to others, with RobX performing most consistently. The study emphasizes the need for a common protocol that defines model changes, measures their impact, and separates generation performance from robustness.
By Marcin Kostrzewa, Maciej Zi\k{e}ba
arXiv:2607. 16476v1 Announce Type: cross Abstract: Software configuration tuning is crucial for optimising system performance, and various optimisers have emerged over the last decade.
By Chao Jiang, Yulong Ye, Tao Chen, Miqing Li
arXiv:2609.01244v1 Announce Type: new
Abstract: Every supervised fine-tuning run forces the same chain of decisions, such as learning rate, batch size, LoRA or full fine-tuning, how many epochs, whic...
By Charles O'Neill, Mudith Jayasekara, Harry Partridge
arXiv:2607. 19386v1 Announce Type: new Abstract: Cross-paper comparison of sparse autoencoder (SAE) interpretability often relies on autointerpretability scores.
By Sinie van der Ben, Neele Roch, Anna Hedstr\"om, Mennatallah El-Assady
Every supervised fine-tuning run forces the same chain of decisions, such as learning rate, batch size, LoRA or full fine-tuning, how many epochs, which optimiser, and what data to feed the model. Eac...
The paper investigates how making responsible‑AI evaluations more efficient—through batching, quantization, and benchmark reduction—affects the stability of conclusions drawn about model behavior. By testing three dense and mixture‑of‑experts models on the BBQ and BBQ‑V datasets under seven different conditions, the authors compare accuracy, bias, reasoning quality, subgroup performance, subset‑membership stability, runtime, and GPU energy consumption against a full‑benchmark BF16 baseline. Findings show that larger batching preserves accuracy and reduces energy in most settings, INT8 largely maintains quality but can increase energy use, INT4 introduces larger, context‑dependent changes, and reduced benchmarks save resources but are highly sensitive to which items are retained, underscoring that efficient evaluation must be validated against the benchmark’s intended conclusions.
By Ahmed El Kady, Aravind Narayanan, Rehana Noorani, Yani Ioannou, Shaina Raza
arXiv:2605. 22949v3 Announce Type: replace Abstract: Foundation-model pools are increasingly used as black-box responders in coordinated systems where a coordinator must decide which response to trust.
By Joss Armstrong
arXiv:2606. 00160v1 Announce Type: cross Abstract: Large language models (LLMs) suffer from degraded safety capabilities even when fine-tuned with benign datasets.
By Junbo Zhang, Qianli Zhou, Xinyang Deng, Wen Jiang, Jie Pan, Jinbiao Zhu
arXiv:2601. 17717v3 Announce Type: replace Abstract: Large Language Models (LLMs) have emerged as powerful tools for generating data across various modalities.
By Kaituo Zhang, Mingzhi Hu, Hoang Anh Duy Le, Fariha Kabir Torsha, Zhimeng Jiang, Minh Khai Bui, Chia-Yuan Chang, Yu-Neng Chuang, Zhen Xiong, Ying Lin, Guanchu Wang, Na Zou
arXiv:2603. 28590v3 Announce Type: replace Abstract: Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs.
By Han Wang, Yifan Sun, Brian Ko, Mann Talati, Jiawen Gong, Zimeng Li, Naicheng Yu, Xucheng Yu, Wei Shen, Vedant Jolly, Huan Zhang
arXiv:2609.15122v1 Announce Type: cross
Abstract: Generative CAD models are expected to remain behaviorally correct after parameter edits, so increasing the number of edit checks is often treated as...
By Hongye Yang, Zhihao Xie, Shengjun Xiong, Boxiao Huang