AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning
arXiv:2607. 21106v1 Announce Type: new Abstract: Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging.
arXiv:2607. 21106v2 Announce Type: replace Abstract: Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging.
arXiv:2607. 21106v1 Announce Type: new Abstract: Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging.
arXiv:2607. 24097v1 Announce Type: new Abstract: Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model.
arXiv:2606. 03329v1 Announce Type: new Abstract: Long-context tasks require LLMs to identify and preserve answer-relevant information from large contexts.
arXiv:2510. 13554v2 Announce Type: replace-cross Abstract: The reasoning pattern of Large language models (LLMs) remains opaque, and reinforcement learning (RL) typically applies uniform credit across an entire generation, blurring the distinction between pivotal and routine steps.
AgenticRag‑R1 is a reinforcement‑learning framework that integrates reasoning, retrieval, and memory through a stack and fine‑grained action space. It uses hierarchical action‑aware rewards and an information‑aware trajectory rejection strategy to support long‑horizon learning. Experiments on multi‑hop, open‑domain, and agentic reasoning benchmarks show that AgenticRag‑R1 outperforms strong baselines and produces robust, interpretable, memory‑aware reasoning behaviors.
arXiv:2606. 10646v1 Announce Type: new Abstract: Token-level credit assignment remains a key obstacle for reinforcement learning (RL) in large language models (LLMs), where RL recipes typically treat all tokens equally, failing to distinguish decisive reasoning steps from routine formatting or fluent filler.
arXiv:2507.21931v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) often produce plausible but poorly-calibrated answers, limiting their reliability on reasoning-intensive tasks....
arXiv:2609.40360v1 Announce Type: cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions...
SIPO (Self‑Instructing Policy Optimization) unifies reinforcement learning with on‑policy self‑distillation by using a contrastive self‑teacher to generate token‑level credit signals. The method samples multiple rollouts per prompt, pairs each with a reference answer and its mistakes, and uses the difference in teacher log‑probabilities to provide dense feedback while still respecting the overall task reward. Experiments on reasoning and code‑generation benchmarks show that SIPO outperforms both RLVR and OPSD baselines without requiring an external teacher or extra generation steps.
arXiv:2609.37119v1 Announce Type: cross Abstract: Recent approaches to reinforcement learning (RL) post-training for large language models increasingly remove the critic to reduce training instabilit...
IterSynth introduces a role-decoupled, iterative synthesis framework for deep search agents, separating planning and synthesis into distinct Planner and Synthesizer modules that maintain a persistent summary state. This design mitigates role coupling and context noise, while the new Role-Decoupled Policy Optimization (RDPO) enhances training by combining outcome rewards with turn-level rubric evaluations. Experiments on five long-horizon benchmarks show IterSynth-8B outperforming prior ≤8B agents by 4.2% and delivering significant zero-shot gains over ReAct on proprietary models.
arXiv:2606. 06787v1 Announce Type: new Abstract: Large Language Models (LLMs) show promise as tool-using agents but remain limited in long-horizon tasks that require remembering, organizing, and reusing knowledge.