arXiv:2609.16340v1 Announce Type: cross
Abstract: Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system'...
By Rohit Dhaipule, Sukhdeep Singh Kharbanda, Prasanth Bathala, Pradyumna Lanka, Anubhav Shrimal
Large language model unlearning aims to suppress designated undesirable knowledge while preserving benign capabilities. Many unlearning objectives focus on suppressing undesired answers, while recent target-guided variants specify replacement behavior but still leave update locality largely unconstrained.
FishBack introduces a pullback Fisher geometry approach for activation steering in transformers, challenging the common Euclidean assumption of intermediate activation spaces. By deriving a closed‑form steering direction based on the Fisher information metric of the softmax layer, the method achieves target concept changes with minimal off‑target distortion, especially in early and middle layers. Experiments on GPT‑2 Small, Llama‑3‑8B, and Qwen3‑8B demonstrate significant reductions in off‑target KL divergence compared to existing steering baselines.
By Sihan Wang, Jiayi Zhao, Qingyan Cao, Hongbo Yao, Lin Shu
arXiv:2511. 05852v4 Announce Type: replace-cross Abstract: Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs).
By Yinjie Cheng, Paul Youssef, Christin Seifert, J\"org Schl\"otterer, Zhixue Zhao
arXiv:2606. 10989v1 Announce Type: new Abstract: Large language model unlearning aims to suppress designated undesirable knowledge while preserving benign capabilities.
By Bocheng Ju, Jianhua Wang, Chengliang Liu, Xiaolin Chang
arXiv:2607. 14111v1 Announce Type: cross Abstract: Can small language models detect and report on perturbations their own internal activations?
By Ely Hahami, Ishaan Sinha, Lavik Jain
arXiv:2609.13611v1 Announce Type: new
Abstract: The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT...
By Diptesh Kanojia, Chi-kiu Lo, Archchana Sindhujan, Samuel Larkin, Greg Hanneman, Alon Lavie
arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.
By Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
The paper investigates how to efficiently repair stale key-value (KV) caches in retrieval‑augmented generation systems after document edits. It proposes a budgeted in‑place recomputation approach and evaluates training‑free position‑selection policies on a factual RAG benchmark. Across three model families, a contiguous edit‑local window consistently recovers most of the post‑edit answer quality while being 13–21 times faster than a full re‑prefill, though its effectiveness diminishes when answer‑bearing text moves downstream.
By Mingyang Mao, Wyatt Mackey, Xiaomin Lin
arXiv:2609.18720v1 Announce Type: new
Abstract: Learned quality estimation (QE) models such as COMETKiwi are widespread and work well for general machine translation evaluation. However, they are kno...
By Kathy H\"ammerl, Gabriel Bretschner, Joern Wuebker
The paper introduces Latent Space Refusal Anchoring (LSR‑Anchoring), a training‑free technique that extracts a refusal direction from English prompts and applies it to the residual stream of instruction‑tuned models at inference time. The primary variant, Mean‑Activation Steering (MAS), works across several architectures (Llama‑3‑8B, Llama‑3.1‑70B, Mistral‑7B‑Instruct, Qwen2.5‑7B), restoring safety for low‑resource African languages with minimal performance loss, while a refined SAE‑Derived Steering (SDS) further reduces KL divergence without degrading legitimate prompt performance. The method shows positive transfer for Yoruba, Igbo, Igala, and Hausa, but fails for Arabic, suggesting a geometric mismatch rather than a data scarcity issue.
By Godwin Abuh Faruna
arXiv:2608. 16210v1 Announce Type: new Abstract: Aggregate accuracy hides where models succeed and fail.
By Zhi Zhang, Lingfeng Lyu, Yue Kang, Doudou Zhou