arXiv AI

Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning

arXiv:2608. 12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers.

arXiv AI
Jul 17

Align AI to Dynamic Human-AI Workflows

arXiv:2607. 14240v1 Announce Type: new Abstract: Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capture the dynamic, context-dependent nature of real-world human-AI interactions.

By Valerie Chen, Cleotilde Gonzalez, Anita Williams Woolley, Michael Lee, Tongshuang Wu, Vincent Conitzer, Aarti Singh
arXiv Computation and Language
Aug 31

AI Alignment through a Game-theoretic Lens: A Survey

The article surveys AI alignment from a game-theoretic perspective, focusing on how large language models and AI agents can be aligned with complex human values in high-risk settings. It categorizes recent progress around key game-theoretic elements and addresses three main challenges: preference diversity, alignment priority, and temporal dynamics. The survey clarifies where game theory benefits current alignment methods, where its application is looser, and what remains to be tackled for robust, adaptive, and verifiable AI systems.

By Yanan Cai, Zhongrui Zhao, Zhigang Lu, Ickjai Lee, Wei Emma Zhang, Minhui Xue, Yihong Zhang, Shuchao Pang, Wei Xiang
arXiv Machine Learning
Sep 10

Do Reasoning Representations Help Humans Evaluate LLM Outputs?

The paper investigates whether reasoning representations—explanations for large language model outputs—aid humans in evaluating those outputs. A controlled human study tested six reasoning formats across tasks of varying complexity, measuring structural understanding, error detection, and trust calibration. Results revealed a mismatch: participants favored planning- and decomposition-based representations, yet simpler chain-of-thought traces better supported verification, trust, and interpretability, while preferred formats increased calibration risks.

By Jaewoo Lim, Sungbok Shin, Sanghyun Hong
OpenAI Blog
Aug 24, 2022

Our approach to alignment research

We are improving our AI systems’ ability to learn from human feedback and to assist humans at evaluating AI. Our goal is to build a sufficiently aligned AI system that can help us solve all other alignment problems.

arXiv AI
Aug 19

Supporting Calibrated Reliance in Human-AI Collaboration: Different Strategies for Different Tasks

The study investigates how different AI support formats influence human decision-making across two tasks: abstract visual reasoning with RAVEN matrices and deductive logical reasoning with LSAT problems. Findings reveal that in visual reasoning, predictions alone and predicted probabilities best support accuracy and error recovery, while in logical reasoning, LLM explanations outperform other supports. The results suggest that effective human–AI collaboration requires task‑specific support strategies rather than a one‑size‑fits‑all approach.

By Ruth Cohen, Lu Feng, Ayala Bloch, Sarit Kraus
arXiv AI
Sep 3

Thinking effort aligns between humans and reasoning models in abductive reasoning

The study examines how the effort expended by large reasoning models (LRMs) compares to that of humans during abductive reasoning tasks. By analyzing reaction times and reasoning traces, the authors find that LRMs and humans exhibit similar patterns of effort and error types. They also demonstrate that decoding strategies allowing models to explore multiple reasoning paths further align the models’ reasoning costs with human effort.

By Henry Arthur