CompOrca: Corpus-Scale Compliance Labelling of Instruction-Tuning Data
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2609.37807v1 Announce Type: cross Abstract: Studying how fine-tuning shapes refusal and noncompliance behaviour requires identifying training examples that refuse, evade or otherwise fail to fu...
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade. The judge is rarely checked.
arXiv:2606. 25487v1 Announce Type: cross Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade.
The paper introduces a diagnostic tool for distinguishing the use of misogynistic slurs from their mention in counter‑speech within code‑mixed Hinglish. It identifies evaluation artifacts in existing corpora, releases a 416‑item minimal‑pair contrast set that decorrelates slur presence and gendered register from labels, and proposes a pair‑consistency metric to assess model performance. Experiments show that even strong baselines struggle to consistently label counter‑speech pairs, while a large language model achieves perfect scores, indicating the benchmark measures genuine capability rather than exploitation of artifacts.
arXiv:2607. 28639v1 Announce Type: cross Abstract: We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias.
arXiv:2609.24885v1 Announce Type: new Abstract: When a language model answers from a curated corpus via graph-based retrieval, a large grounding uplift does not establish reasoning over the retrieved...