arXiv:2602.01161v2 Announce Type: replace
Abstract: The global deployment of large language models (LLMs) has raised concerns about cultural misalignment, yet the linguistic properties of fine-tuning...
By Reem I. Masoud, Chen Feng, Shunta Asano, Saied Alshahrani, Philip Colin Treleaven, Miguel R. D. Rodrigues
The paper introduces a bias depth score to differentiate between stable model preferences (Deep biases) and prompt‑dependent responses (Shallow biases) in large language models. By analyzing 4,442 opinion prompts across four models, it finds that only about a quarter of concentrated preferences persist after scenario reframing, indicating that most are shallow. The study shows Deep biases are more often inherited from pretraining and harder to remove through fine‑tuning or prompt‑based debiasing, highlighting the need to distinguish learned biases from prompt artifacts.
By An Vo, Vy Tuong Dang, Khai-Nguyen Nguyen, Emilio Villa-Cueva, Thamar Solorio, Anh Totti Nguyen, Daeyoung Kim
arXiv:2604. 17289v2 Announce Type: replace Abstract: Supervised fine-tuning of large language models relies on human-annotated data, yet annotation pipelines routinely involve multiple crowdworkers of heterogeneous expertise.
By Sajjad Ghiasvand, Mark Beliaev, Mahnoosh Alizadeh, Ramtin Pedarsani
arXiv:2609.05899v1 Announce Type: cross
Abstract: Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective...
By Peng Lai, He Zhu, Zhiwen Ruan, Dongdong Zhang, Yun Chen, Peng Li, Furu Wei, Yang Liu, Guanhua Chen
arXiv:2603. 09692v2 Announce Type: replace-cross Abstract: Reinforcement Learning from Human Feedback (RLHF) has become the standard for aligning Large Language Models (LLMs), yet its efficacy is bottlenecked by the high cost of acquiring preference data, especially in low-resource and expert domains.
By Davit Melikidze, Marian Schneider, Jessica Lam, Martin Wertich, Ido Hakimi, Barna P\'asztor, Andreas Krause
arXiv:2606. 09525v1 Announce Type: cross Abstract: During instruction fine-tuning (IFT), large language models (LLMs) learn to follow instructions by using the provided context to answer a query.
By Nadya Yuki Wangsajaya, Haeun Yu, Isabelle Augenstein
The paper introduces CuCu, a multi‑agent LLM framework that converts national social studies curricula into open‑ended, culture‑specific question‑answer pairs for fine‑tuning language models. Using the Korean curriculum, the authors build KCaQA, a dataset of 34.1k QA pairs that cover culture‑specific topics and ground responses in local sociocultural contexts. Experiments show that fine‑tuning with KCaQA improves the model’s cultural alignment and relevance to Korean society.
By Haneul Yoo, Won Ik Cho, Geunhye Kim, Jiyoon Han
arXiv:2608.30902v1 Announce Type: new
Abstract: Adapting large language models to user-specific preferences is often constrained by the cost of human annotation, making preference optimisation imprac...
By Alessio Galatolo, Meriem Beloucif
We’ve obtained state-of-the-art results on a suite of diverse language tasks with a scalable, task-agnostic system, which we’re also releasing. Our approach is a combination of two existing ideas: transformers and unsupervised pre-training.
CHAI for LLMs is a framework that improves large language models’ performance on code‑mixed translation tasks by using LLMs as annotators to create preference data, applying reinforcement learning from AI feedback, incorporating LLM‑generated domain knowledge for iterative refinement, and evaluating on real‑world datasets. The approach yields a 68.45% average win rate over state‑of‑the‑art open‑source models in human‑adjudicated tests. It demonstrates a scalable method to enhance code‑mixed language understanding in open‑source LLMs.
By Wenbo Zhang, Aditya Majumdar, Asif Ekbal, Amulya Yadav
Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to address this limitation is to specialise existing models...
The paper proposes a method for culturally aligning large language models (LLMs) using soft prompt tuning optimized via Differential Evolution (DE). Unlike traditional fine‑tuning or reinforcement learning, this approach keeps model weights frozen and requires no preference data, instead leveraging aggregated survey scores from Hofstede's Value Survey Module (VSM13). Experiments on four countries and four instruction‑tuned models show that DE‑optimized prompts reduce cultural discrepancy, improve agreement with the World Values Survey, and are preferred in blinded pairwise evaluations by LLM judges.
By Reem I. Masoud, Martin Ferianc, Philip Treleaven, Miguel Rodrigues