arXiv AI By Aritra Mazumder, Nusrat jahan Lia

AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP

Read the original on arXiv AI →

arXiv:2607. 11098v1 Announce Type: cross Abstract: Tool-using LLM agents are mostly evaluated assuming all tools work.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Aug 10

Online Monitoring and Corrective Steering of Programming Agents

arXiv:2608. 06701v1 Announce Type: cross Abstract: Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or the issue description lacks the information needed to localize and repair it.

By Shuyang Liu, Saman Dehghan, Ji Young Kim, Jatin Ganhotra, Martin Hirzel, Reyhaneh Jabbarvand