Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds
arXiv:2607. 28908v1 Announce Type: new Abstract: Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers.
arXiv:2607. 28908v1 Announce Type: new Abstract: Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers.
The study examines how different vision‑language‑action (VLA) policies execute a manipulation task by comparing the geometry of their end‑effectors across 15,000 closed‑loop LIBERO rollouts. By pairing 3,600 configuration‑matched policy executions, the authors find that when both policies succeed, their end‑effector trajectories are much closer (median DTW distance 0.0120 m) than when only one succeeds (0.0380 m), a pattern consistent across all tasks, policy pairs, and nine representations. Even successful executions remain as far from same‑task demonstrations as the demonstrations are from each other, indicating that task‑associated geometry, rather than training data overlap, drives these differences.
arXiv:2605.27186v2 Announce Type: replace Abstract: Large language models often solve tasks from a fully specified prompt but degrade when the same requirements unfold over multiple turns, known as t...
The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.
arXiv:2602. 24287v2 Announce Type: replace-cross Abstract: In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assistant responses.
arXiv:2608.30198v1 Announce Type: new Abstract: Long-term memory enables large language models (LLMs) to preserve and reuse information across interactions, but it can also turn localized errors into...
The paper introduces the Agent-Editing World Model (AEWM), a new approach that models how reasoning and actions influence future task progress instead of simulating tool responses. AEWM includes an Action Judge that classifies decisions as Critical, Exploratory, or Noisy, and a State Revision mechanism that edits noisy reasoning–action continuations from the same observed history. The integrated system, EditAct, directly updates the underlying state during real execution, leading to significant performance gains across multiple benchmarks and agent backbones.
MIRAGE is a controlled study that examines how multimodal personal agents use historical evidence when conversation state changes. The study keeps evidence, questions, and scoring constant while varying only the conversation state, then checks if agents can determine answerability, recover the correct source, and answer from it. Results across seven multimodal backbones show distinct failure regimes before and after compaction, heavy reliance on context continuity by open-weight models, and mixed effects of retrieval pressure on source attribution.
arXiv:2609.21423v1 Announce Type: new Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...
arXiv:2608.30650v1 Announce Type: new Abstract: LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn...
arXiv:2603. 00270v3 Announce Type: replace-cross Abstract: Large language models can process millions of tokens, yet how they handle conflicting information within context remains poorly understood.
The paper introduces OODA-Tool, a typed closed‑loop policy that separates state preservation from action generation to reduce state‑action competition in multi‑turn tool use. It follows Boyd’s Observe‑Orient‑Decide‑Act cycle, reconstructing task state, deciding on execution, forming admissible actions, and then realizing outputs. Experiments with Qwen3 models show OODA‑Tool consistently improves task success, especially for smaller models and tasks requiring accumulated information.