arXiv AI

Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment

arXiv AI
6d ago

Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue

The paper examines how conversational AI, specifically ChatGPT, displays aspects of cooperative dialogue such as morality, politeness, and alignment compared to human-human conversations. Using over 26,000 multi‑turn dialogues and mixed‑effects modeling, the authors find that AI mimics the surface features of cooperation—like warmth and hedging—yet lacks the underlying social architecture that drives mutual adaptation. Key findings include a dissociation between AI’s moral output and human negotiation, a decline in linguistic convergence, and a reversal of typical human accommodation mechanisms when interacting with AI.

By Marina Mitiaeva, Lu Xiao
arXiv AI
Jul 17

Align AI to Dynamic Human-AI Workflows

arXiv:2607. 14240v1 Announce Type: new Abstract: Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions.

By Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley, Michael Lee, Tongshuang Wu, Vincent Conitzer, Aarti Singh
arXiv AI
Jun 10

A Comprehensive Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications

arXiv:2410. 15595v4 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical.

By Wenyi Xiao, Zechuan Wang, Leilei Gan, Shuai Zhao, Zongrui Li, Ruirui Lei, Wanggui He, Luu Anh Tuan, Long Chen, Hao Jiang, Zhou Zhao, Fei Wu
arXiv AI
Aug 5

Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation

arXiv:2608. 03044v1 Announce Type: cross Abstract: Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity.

By Seth Grief-Albert, Jessica Bo, Difan Jiao, Ashton Anderson
arXiv Computation and Language
Sep 11

Inverse Turing Bench: Evaluating Language Models as Judges of Human vs. AI Dialogue

The paper introduces Inverse Turing Bench, a benchmark designed to assess how well language models can distinguish between human-only and human-AI dialogues in multi-turn text. It provides paired dialogue transcripts and evaluates models on correctly identifying the type of conversation. Preliminary results show GPTZero, Claude Opus-4.6, and GPT-5.5 achieving the highest accuracies of 89.41%, 77.92%, and 75.94% respectively, highlighting both the strengths and limitations of statistical versus semantic detection approaches.

By William Hager, Ishika Rathi, Masum Hasan, Cameron Jones
arXiv Computation and Language
Aug 31

AI Alignment through a Game-theoretic Lens: A Survey

The article surveys AI alignment from a game-theoretic perspective, focusing on how large language models and AI agents can be aligned with complex human values in high-risk settings. It categorizes recent progress around key game-theoretic elements and addresses three main challenges: preference diversity, alignment priority, and temporal dynamics. The survey clarifies where game theory benefits current alignment methods, where its application is looser, and what remains to be tackled for robust, adaptive, and verifiable AI systems.

By Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue, Yihong Zhang, Shuchao Pang, Wei Xiang
arXiv Computation and Language
6d ago

Cultural Alignment in Large Language Models Using Soft Prompt Tuning

The paper proposes a method for culturally aligning large language models (LLMs) using soft prompt tuning optimized via Differential Evolution (DE). Unlike traditional fine‑tuning or reinforcement learning, this approach keeps model weights frozen and requires no preference data, instead leveraging aggregated survey scores from Hofstede's Value Survey Module (VSM13). Experiments on four countries and four instruction‑tuned models show that DE‑optimized prompts reduce cultural discrepancy, improve agreement with the World Values Survey, and are preferred in blinded pairwise evaluations by LLM judges.

By Reem I. Masoud, Martin Ferianc, Philip Treleaven, Miguel Rodrigues
arXiv AI
Sep 3

TUX: Measuring Human--AI Tacit Understanding

The paper introduces TUX, a Tacit Understanding Index that measures how similarly humans and large language models (LLMs) place concepts along subjective spectra in a task inspired by the game Wavelength. Using 241 human participants and 200 profile-conditioned LLM agents across four models, the study finds that human–agent pairs with similar traits achieve higher TUX scores, indicating that tacit alignment is linked to person-level characteristics. Regression analyses show that richer predictor sets—including individual traits, decision-making styles, and confidence—improve the explainability of TUX beyond simple trait-distance baselines.

By Yueshen Li, Hanyi Min, Vedant Das Swain, Koustuv Saha