Building Agent Harnesses for Scientific Curation from Multimodal Sources
arXiv:2606. 21005v2 Announce Type: replace Abstract: Scientific discovery workflows often depend on structured curation from the literature.
Sci‑MMR is a new benchmark for multi‑step evidence‑grounded scientific reasoning in multimodal agents, featuring 235 multi‑hop tasks across four disciplines and an average of nine figure panels per task. It evaluates not just final answer accuracy but also the recovery of structured evidence from scientific claims, citations, visual data, and supporting regions. Experiments on eight state‑of‑the‑art models show a gap of over 20 points between answer accuracy and complete evidence recovery, highlighting significant challenges in evidence acquisition and integration.
arXiv:2606. 21005v2 Announce Type: replace Abstract: Scientific discovery workflows often depend on structured curation from the literature.
Mr.LHDR is a new benchmark designed to evaluate deep research agents on long‑horizon, multimodal tasks. It presents questions built from hidden Node‑Relation graphs that require an average of 12.1 intermediate conclusions and a mean dependency depth of 10.4 before arriving at a single verifiable answer. The benchmark tests both final answers and the correctness of intermediate conclusions, using metrics such as Overall Accuracy, Strict Accuracy, Checklist Score, and Dependency‑Aware Checklist Score.
arXiv:2606. 13020v1 Announce Type: new Abstract: Three paradigmatic forms of inference recur across scientific reasoning: deduction, induction, and causal abduction.
arXiv:2607. 16131v1 Announce Type: cross Abstract: Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context.
arXiv:2608. 09281v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition, information extraction, or compliance checking, leaving open whether models can combine distributed visual evidence with engineering principles to reach a conclusion.
arXiv:2607. 01131v1 Announce Type: cross Abstract: Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation.
arXiv:2601. 02854v2 Announce Type: replace Abstract: As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to improve answer quality and support complex reasoning.
arXiv:2608. 06931v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex real laboratory science.
arXiv:2608. 14075v1 Announce Type: new Abstract: Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret.
arXiv:2609.05518v1 Announce Type: cross Abstract: Despite the strong capabilities of multimodal large language models (MLLMs), their parametric knowledge remains incomplete and difficult to update, m...
arXiv:2609.06192v1 Announce Type: new Abstract: Scientific coding agents produce interdependent code, results, figures, and claims, yet evaluating final outputs alone does not establish whether the...
arXiv:2608.29088v1 Announce Type: new Abstract: Multimodal question answering remains sensitive to noisy, incomplete, and weakly grounded evidence. Long unstructured contexts can introduce redundancy...