Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents
arXiv:2605. 21347v3 Announce Type: replace Abstract: Diagnosing failures in LLM agents remains largely manual.
arXiv:2605. 21347v3 Announce Type: replace Abstract: Diagnosing failures in LLM agents remains largely manual.
AdaLens is an interactive system designed to monitor and steer long-running, autonomous data analysis workflows powered by large language models. It provides a storyline-based interface that unifies analytical plans, execution progress, intermediate findings, and data-column involvement, enabling analysts to observe evolving reasoning and evidence. The system also offers steering interactions that allow users to redirect low-value directions or deepen promising ones during execution.
arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.
The article introduces AgentActionBench, a benchmark designed to evaluate agent-based experiment reproduction across machine learning and AI4Science papers. It employs an MCP-based Action Recorder to capture agents’ behavior during reproduction and assesses the resulting traces against paper-specific rubrics. The benchmark includes 150 papers, with a human-annotated subset and model-assisted augmentation expanding it to over 10,000 rubric items, revealing that current systems face execution bottlenecks but that model-generated rubrics correlate strongly with human judgments.
arXiv:2603. 14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions.
arXiv:2608. 07346v2 Announce Type: replace Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.
arXiv:2608. 07346v1 Announce Type: new Abstract: With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains.
arXiv:2608. 05179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review.
RefactorPlatform is an open‑source harness that standardizes the evaluation of repository‑scale refactoring agents by fixing the environment and systematically varying design choices such as model backbone, execution regime, and prompt specificity. Each run operates in an isolated workspace, logs detailed telemetry, and verifies changes with AST‑based checks. Experiments on 100 RefactorBench tasks show that AST‑aware chunking improves performance by 25‑30%, a lean retrieval‑augmented single agent outperforms a sub‑agent configuration, and retrieval’s accuracy gains offset its token overhead, keeping cost per successful refactoring unchanged.
The paper introduces a self‑supervised framework called software‑in‑the‑loop reconstruction (SWR) that automatically generates reference outputs and verification targets for terminal agents by leveraging existing scientific software workflows. SWR executes multiple input configurations, partitions cases into public observations and hidden evaluations, and uses a hierarchical verifier to assess agent‑generated programs against hidden workflow outputs. The authors demonstrate the approach on 500 workflows across six domains, achieving significant performance gains on the Terminal‑Bench 2 benchmark with the Qwen3.8‑Max and Qwen3.8‑27B models.
The paper presents EvalAgent, an AI assistant that automates agent evaluation by encoding domain expertise into evaluation skills such as procedural instructions, reusable code, and dynamic API retrieval. EvalAgent constructs a trace-based pipeline that outputs metrics, executable code, and reports, and is evaluated using a new meta-evaluation framework and AgentEvalBench. Results show that EvalAgent improves the Eval@1 metric from 17.5% to 65% and receives 79.5% human expert preference, while ablation studies confirm the importance of evaluation skills.
arXiv:2607. 05456v1 Announce Type: new Abstract: While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whether AI-generated manuscripts meet the quality and rigor required for real-world publication.