arXiv:2607. 01208v1 Announce Type: cross Abstract: Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering user decisions at scale.
By Shayan Talaei, Abhinav Chinta, Devvrit Khatri, Amin Karbasi, Azalia Mirhoseini, Amin Saberi
arXiv:2605. 00994v2 Announce Type: replace-cross Abstract: Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors.
By Mohammed Abu Baker, Luca Baroni, Dan Wilhelm
arXiv:2607. 22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality.
By Yunting Song, Matthew Watson, Peter Grabowski, Jun Qin
arXiv:2606. 06946v1 Announce Type: cross Abstract: We present LoRA-MINT, a new methodology for Membership Inference Test (MINT) applied to recent Large Language Models (LLMs) fine-tuned for specific Natural Language Processing (NLP) tasks through Low-Rank Adaptation (LoRA).
By Gonzalo Mancera, Daniel DeAlcala, Aythami Morales, Julian Fierrez, Ruben Tolosana, Francisco Jurado
arXiv:2609.22136v1 Announce Type: cross
Abstract: Text anomaly detection, the task of identifying text instances that deviate from normal language patterns, is crucial for language-driven application...
By Yanyu Qian, Pengcheng Weng, Yue Tan, Enguang Zuo, Yu Zheng, Yixin Liu
arXiv:2604. 01904v3 Announce Type: replace-cross Abstract: Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights holders query a target LLM with raw proprietary data and assess whether the model assigns them stronger memorization-based detection signals, e.
By Muxing Li, Zesheng Ye, Sharon Li, Feng Liu
The paper introduces Exemplar Partitioning (EP), an unsupervised technique that constructs interpretable feature dictionaries from large language model activations by clustering streamed activations into Voronoi regions defined by exemplars and their averages. EP allows comparison of dictionaries across layers, checkpoints, and architectures, and demonstrates utility in interpreting model behavior, tracking training dynamics, detecting hidden concepts, and enabling targeted interventions. Experiments on Gemma‑2‑2B and Llama‑3.1‑8B show EP can reveal how instruction tuning reorganizes harmful prompt activations, facilitate interventions that alter model responses, and achieve high concept‑detection performance while requiring far fewer construction tokens than comparable methods.
By Jessica Rumbelow
arXiv:2605. 14568v3 Announce Type: replace-cross Abstract: Context.
By Ali Hassaan Mughal, Noor Fatima, Muhammad Bilal
arXiv:2607. 04222v1 Announce Type: new Abstract: Interpretability methods aim to reveal the features represented inside large language models (LLMs).
By Amit LeVi, Elad David, Max Fomin
J-Miner extracts executable decision knowledge from fine‑tuned language‑model classifiers by aggregating internal signals aligned with named concepts and learning rules over them. The mined rules reproduce up to 98.3% of the original classifier’s decisions and outperform compact rules based on input words by 6.0–29.5 percentage points in behavioral fidelity. Moreover, these rules can be transferred to lightweight student models that use only about 1/24 of the parameters yet retain 99.8% of the source classifier’s accuracy.
By Yunfan Gao, Xinyi Huang, Tao Sheng, Haorui Song, Yun Xiong, Haofen Wang
arXiv:2609.23892v1 Announce Type: new
Abstract: Mechanistic interpretability defines features as the fundamental units of a neural network and circuits as the weighted subgraphs that carry out its co...
By Edward G. Friedman, Xiangchen Song
arXiv:2606. 23989v1 Announce Type: cross Abstract: End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions they offer are typically coarse (whole documents or passages) and generated post hoc, leaving each summary statement hard to verify.
By Shuo Guan