arXiv AI By Woojin Kim, Sieun Hyeon, Jusang Oh, Jaeyoung Do

VALUEFLOW: Toward Pluralistic and Steerable Value-based Alignment in Large Language Models

Read the original on arXiv AI →

arXiv:2602. 03160v2 Announce Type: replace Abstract: Aligning Large Language Models (LLMs) with the diverse spectrum of human values remains a central challenge: preference-based methods often fail to capture deeper motivational principles.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 21

Cultural Alignment in Large Language Models Using Soft Prompt Tuning

The paper proposes a method for culturally aligning large language models (LLMs) using soft prompt tuning optimized via Differential Evolution (DE). Unlike traditional fine‑tuning or reinforcement learning, this approach keeps model weights frozen and requires no preference data, instead leveraging aggregated survey scores from Hofstede's Value Survey Module (VSM13). Experiments on four countries and four instruction‑tuned models show that DE‑optimized prompts reduce cultural discrepancy, improve agreement with the World Values Survey, and are preferred in blinded pairwise evaluations by LLM judges.

By Reem I. Masoud, Martin Ferianc, Philip Treleaven, Miguel Rodrigues
arXiv AI
Jul 7

Evolutionary Guided Decoding: Iterative Value Refinement for LLMs

arXiv:2503. 02368v4 Announce Type: replace-cross Abstract: While guided decoding, especially value-guided methods, has emerged as a cost-effective alternative for controlling language model outputs without re-training models, its effectiveness is limited by the accuracy of the value function.

By Zhenhua Liu, Lijun Li, Ruizhe Chen, Yuxian Jiang, Tong Zhu, Zhaochen Su, Wenliang Chen, Jing Shao
Hugging Face Trending Papers
Jul 22

D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in daily scenarios involving multiple value conflicts and simplistic evaluation formalisms that fail to assess LLMs' value alignment.

arXiv AI
Aug 21

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

arXiv:2509. 08022v3 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation.

By Yao Liang, Dongcheng Zhao, Feifei Zhao, Guobin Shen, Yuwei Wang, Dongqi Liang, Yi Zeng
arXiv AI
4d ago

PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.

By Cheng Chang, Yining Mao, Peng Qi