Interactive Task Alignment as a POMDP
arXiv:2607. 16412v1 Announce Type: new Abstract: Current benchmarks for language models primarily evaluate execution on fully specified tasks.
arXiv:2604. 21827v2 Announce Type: replace Abstract: In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.
arXiv:2607. 16412v1 Announce Type: new Abstract: Current benchmarks for language models primarily evaluate execution on fully specified tasks.
arXiv:2608.29610v1 Announce Type: new Abstract: The current alignment tuning paradigm for Large Language Models (LLMs) prioritizes surface-level behaviors -- fluency, safety, and tonal consistency. W...
arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.
arXiv:2608. 12372v1 Announce Type: new Abstract: AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers.
The paper introduces FISER, a framework that explicitly infers human goals and intentions before planning actions for AI agents to follow natural language instructions in collaborative embodied tasks. It employs Transformer-based models and is evaluated on the HandMeThat benchmark, outperforming end-to-end approaches and strong baselines such as Chain of Thought prompting. FISER achieves state‑of‑the‑art performance on this embodied social reasoning task.
arXiv:2605. 28882v2 Announce Type: replace-cross Abstract: With the rapid advancement of large language models, evaluating human-likeness in open-ended conversation has become increasingly important.
arXiv:2605. 17064v2 Announce Type: replace Abstract: Large language models are optimized for instruction following and agentic tasks remain poorly aligned with the requirements of high-quality creative writing.
arXiv:2608. 01366v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts.
arXiv:2609.01588v1 Announce Type: cross Abstract: Writing involves diverse cognitive activities, from ideation to revision, and writers' needs vary across individuals and moments. Proactive AI promis...
arXiv:2601. 02813v3 Announce Type: replace Abstract: Aligning language models to qualitative behavioral traits, such as human-likeness, remains difficult because they are hard to define, measure, and optimize.
arXiv:2606. 10327v1 Announce Type: cross Abstract: Automated Essay Scoring (AES) systems must judge interdependent discourse elements (e.
We are improving our AI systems’ ability to learn from human feedback and to assist humans at evaluating AI. Our goal is to build a sufficiently aligned AI system that can help us solve all other alignment problems.