The paper introduces a reference‑based bias detection method that audits hidden‑state representations of language models by encoding sentences as similarities to a fixed set of anchor sentences. This relative representation allows comparison across model variants, such as before and after fine‑tuning, and yields a metric called Representational Bias Shift (ΔB). ΔB correlates strongly with output‑level bias changes, can detect bias‑increasing checkpoints with high ROC AUC, and is computationally efficient, requiring only a few minutes and far less compute than traditional benchmarks.
By Marek Jeli\'nski, Jan Dubi\'nski, Maciej Chrabaszcz, Sebastian Cygert
arXiv:2606. 03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks cannot be trusted, their items having likely leaked into pretraining, so scores reflect memorization rather than fitness.
By Alexander Apartsin, Yehudit Aperstein
The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.
By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal
arXiv:2503.05684v2 Announce Type: replace-cross
Abstract: Pre-trained foundation models can be efficiently adapted for specific tasks using Low-Rank Adaptation (LoRA), but the fairness properties of...
By Parameswaran Kamalaruban, Mark Anderson, Stuart Burrell, Maeve Madigan, Piotr Skalski, David Sutton
The paper demonstrates that a single example from the BBQ fairness benchmark can dramatically improve a model’s performance, raising accuracy from 79.9% to 92.9% with Group Relative Policy Optimization and to 99.0% with one-shot in-context learning. This effect is consistent across different model families and is driven by the model’s reasoning traces, which adopt a category‑agnostic "missing evidence" pattern. The authors argue that BBQ-style multiple‑choice abstention tests capture only a single structural cue and therefore do not guarantee true fairness, calling for broader evaluation suites.
By Naihao Deng, Samee Arif, Shuaichen Chang, Yulong Chen, Rada Mihalcea
arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.
By Ning Liu
Swiss-Knife is a framework that extends decode‑time alignment for frozen language models by treating the alignment specification as a runtime object. It introduces hot‑swappable scoring blades, a batch normaliser, a pairwise aggregation operator, and a selection rule, and characterises admissible aggregation operators with a representation theorem. In experiments, Swiss‑Knife paired with DPO‑LoRA blades and an uncertainty‑aware pairwise tournament outperforms six existing decode‑time methods, achieving a higher harmonic F1 score, lower refusal rate, and faster objective reconfiguration.
By Agnibh Karmakar, Mayur Parvatikar, Shreyash Dhoot, Amit Dhanda, Aman Chadha, Kapil Wanaskar, Vinija Jain, Amitava Das
arXiv:2607.27836v2 Announce Type: replace
Abstract: Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially recover...
By Xiangyu Yin, Jiaxu Liu, Zhen Chen, Chih-Hong Cheng
arXiv:2605. 23145v2 Announce Type: replace-cross Abstract: Individual fairness, the notion that "similar individuals should be treated similarly," provides a strong and flexible fairness guarantee for algorithmic decision makers.
By Conlan Olson, Linjun Zhang, Zhun Deng, Pragya Sur
arXiv:2606. 15493v1 Announce Type: new Abstract: Model stealing attacks, where adversaries create high-fidelity surrogate models, are a significant threat to the intellectual property of machine learning services.
By Eliott Baltz, Satoshi Hara, Ulrich A\"ivodji
arXiv:2608. 01023v1 Announce Type: new Abstract: We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision signal used to train a surrogate, and the provable per-input query cost of recovering the clean logits.
By Chi Wang, Hanwen Wang, Yu Xia, Zihan Wang, Guangdong Bai
arXiv:2604. 27733v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with human intent, whether through explicit reward modeling or direct methods such as DPO, fundamentally relies on minimizing a surrogate loss as a proxy for the true pairwise ranking objective.
By Mehryar Mohri, Yutao Zhong