Hugging Face Trending Papers

From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data

The paper presents an expert-driven method for turning normative principles—specifically Islamic ethical, theological, and jurisprudential traditions—into alignment data for language models. Over a year, seven experts curated 2.8K supervised fine-tuning examples and 5.4K preference pairs in Arabic-English, then evaluated models trained on these datasets. Experiments show that models trained with the curated SFT data outperform a baseline in expert judgments, while adding preference data yields a smaller, non-significant improvement.

arXiv AI
4d ago

PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.

By Cheng Chang, Yining Mao, Peng Qi
arXiv AI
Sep 1

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.

By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
arXiv AI
Sep 21

Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models

The paper introduces a 12,000-instance dataset of two-option moral dilemmas covering three pairwise value conflicts—Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty—translated into Hindi, Arabic, Spanish, and Chinese to test cross‑lingual behavior. Benchmarking on GPT‑5‑mini shows a consistent preference for Honesty over Autonomy across all languages when no policy is provided, while Llama‑3.2‑1/3B models exhibit a strong first‑option bias that is largely eliminated by plain fine‑tuning or Direct Preference Optimization, raising accuracy above 98%. The authors propose a task vector transfer method that orthogonalizes value preference vectors with respect to general instruction‑following vectors, effectively isolating specific value directions and enabling task arithmetic to flip a model’s stance.

By Utkarsh Agarwal, Monojit Choudhury
Hugging Face Trending Papers
Jul 22

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer.

arXiv AI
Aug 26

Preference Data Selection for Mitigating the Alignment Tax in Large Language Models

The paper introduces BALIGN, a balanced data selection strategy designed to reduce catastrophic forgetting—referred to as the alignment tax—in large language models during preference-based alignment. By analyzing preference optimization gradients, the authors identify three data-centric features that influence parameter drift: the reference model's log-probability margin, token length differences between chosen and rejected responses, and TF‑IDF similarity to general capability corpora. BALIGN aggregates these features into a composite risk score to filter out high-risk preference samples, thereby preserving foundational capabilities while maintaining alignment gains with minimal computational overhead.

By Minsu Kim, Jianxun Lian, Xing Xie, Steven Euijong Whang
Hugging Face Trending Papers
Jul 9

PLURAL: A Global Dataset for Value Alignment

Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse value systems. We introduce PLURAL, a large-scale, value-focused preference dataset grounded in the Integrated Values Survey (IVS), a nationally representative survey spanning 92 countries.