ReactBench is a benchmark designed to evaluate the structural reasoning abilities of multimodal large language models (MLLMs) using chemical reaction diagrams. The dataset contains 1,618 expert‑annotated question‑answer pairs that test reasoning across four hierarchical task dimensions, from simple endpoint counting to complex topological analysis. Evaluation of 24 MLLMs shows a performance gap of more than 30% between anchor‑based tasks and holistic structural reasoning tasks, indicating that current models struggle with reasoning over branching, converging, and cyclic structures.
By Qiang Xu, Shengyuan Bai, Yu Wang, He Cao, Leqing Chen, Yuanyuan Liu, Bin Feng, Zijing Liu, Yu Li
arXiv:2607. 01942v1 Announce Type: new Abstract: LLM-based agents have shown strong potential for solving complex multi-step tasks, yet existing performance improvements often rely on either scaling to larger backbone models or task-specific fine-tuning.
By Yue Zhang, Sihan Chen, Ziwen Huang, Hanyun Cui, Kangye Ji, Zhi Wang
Unified-MAS is a two-stage framework that decouples node implementation from orchestration in Automatic Multi-Agent Systems. It first searches external knowledge to synthesize domain‑specific node blueprints, then uses a perplexity‑guided reward to optimize bottleneck nodes. Experiments across four specialized domains show that adding Unified-MAS to existing baselines improves performance‑cost trade‑offs by up to 14.2% while lowering costs.
By Hehai Lin, Yu Yan, Zixuan Wang, Bo Xu, Sudong Wang, Weiquan Huang, Ruochen Zhao, Minzhi Li, Chengwei Qin
arXiv:2609.14066v1 Announce Type: cross
Abstract: Although existing multi-agent Retrieval-Augmented Generation (RAG) systems have demonstrated promise on complex multimodal reasoning tasks, they rema...
By Zhongyu Wang
arXiv:2607. 08403v1 Announce Type: new Abstract: The application of lightweight Large Language Models in rule-based scientific domains remains severely limited by their tendency to mimic linguistic patterns rather than reproduce axiomatic reasoning, causing frequent hallucinations.
By Runzhe Liu, Biquan Bie, Zihao Wang, Yuchao Ma, Yexin Liu, Xinghai Li, Harry Yang, Wenbo Yang, Jinzhe Cao, Shengyang Tao
arXiv:2607. 07321v1 Announce Type: new Abstract: Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks.
By Haipeng Ding, Yuexiang Xie, Zhewei Wei, Yaliang Li, Bolin Ding