Inside VAKRA: Reasoning, Tool Use, and Failure Modes of Agents
Related stories
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
arXiv:2510. 19771v4 Announce Type: replace Abstract: LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, they exercise agency to anticipate user needs and solve them autonomously.
AgentDynEx: Nudging the Mechanics and Dynamics of Multi-Agent Simulations
arXiv:2504. 09662v4 Announce Type: replace-cross Abstract: Multi-agent large language model simulations have the potential to model complex human behaviors and interactions.
AI Agents Are Here. What Now?
Computer-Using Agent
POIROT: Interrogating Agents for Failure Detection in Multi-Agent Systems
arXiv:2606. 02282v1 Announce Type: new Abstract: Orchestrating Large Language Models into Multi-Agent Systems (LLM-MAS) has unlocked remarkable reasoning capabilities, yet emergent failures and hallucinations that resist characterisation block their deployment in safety-critical domains -- a gap made legally untenable by emerging AI regulation.
Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems
arXiv:2607. 07989v1 Announce Type: cross Abstract: Large language model (LLM) based multi-agent systems enable complex problem solving through coordinated reasoning and action, but their distributed structure also introduces new challenges in diagnosing system-level failures.
Agents All the Way Down; A Methodology for Building Custom AI Agents from Substrate to Production
arXiv:2606. 11869v1 Announce Type: cross Abstract: Custom AI agents areagents that live inside their own application, talk to their own data and tools, enforce their own security boundaries, and carry their own brand and audit trail.
Towards the Harness of Embodied Agents
arXiv:2608. 11246v1 Announce Type: new Abstract: The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the infrastructure around it.
Build AI agents with the Mistral Agents API
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP
arXiv:2607. 11098v1 Announce Type: cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work.
LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans
arXiv:2607. 10878v1 Announce Type: new Abstract: AI agents are evolving from answer engines into persistent teams that use tools, delegate work, learn from experience, and modify the artifacts that shape their future behavior.