arXiv AI

Schema: Discovering Unknown Environments via Agentic Program Induction

arXiv AI
Jul 7

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

arXiv:2607. 05174v1 Announce Type: new Abstract: Language agents, i.

By Zhiheng Xi, Dingwen Yang, Jiaqi Liu, Jixuan Huang, Honglin Guo, Baodai Huang, Tinggang Chen, Qi Zhang, Zhonghang Lu, Chenyu Liu, Jiajun Sun, Jiazheng Zhang, Dingwei Zhu, Xin Guo, Junzhe Wang, Zhihao Zhang, Yuming Yang, Junjie Ye, Minghe Gao, Dongrui Liu, Jiaming Ji, Guohao Li, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv AI
Aug 24

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

AgentMercury is a scalable framework that synthesizes executable environments from high‑level business scenarios instead of task‑specific benchmarks. It creates a persistent world with entities, services, tools, and invariants, allowing diverse tasks and interaction trajectories to emerge naturally. The authors generated 4,783 environments across 14 industries and 50 countries, and training reinforcement‑learning agents on them improved performance on enterprise workflows and out‑of‑domain benchmarks, while the construction process itself can be learned to increase authoring success.

By Minbyul Jeong, Chanwoong Yoon
arXiv Computation and Language
Sep 17

ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments

arXiv:2609.19134v1 Announce Type: new Abstract: Scientific code repositories encode decades of human knowledge in executable models, methods, and tools. Yet fragmented toolchains, implicit domain con...

By Hejia Geng, Zesen Huang, Haoyang Li, Wenbin Li, Koutian Wu, Zihan Zhou, Yuanbo Pang, Weihao Liu, Zigong Xu, Zhiping Li, Zongzheng Zhang, Chuanfei Dong, Jiankai Sun, Tianzhe Zheng, Fengyu Xie, Yue Ma, Yueheng Shi, Tong Xie, Zonglin Di, Xianrong Liu, Qucheng Gao, Yimin Liu, Jiaming Pan, Sheng Huang, Xiao-Han Ma, Lanqing Yuan, Zhenlin Zhu, Ziang Liu, Ziyang Xu, Junkai Wang, Kangkai Liang, Jiayi Xian, Zehong Zhao, Liuwei Xu, Jingxu Xie, Peijin Zhang, Qiang Gao, Chengyi Xing, Zhe Zhao, Xi Wang, Yaopeng Xing, Xing Meng, Zhenfei Yin, Yingcheng Wu, Ling Yang
arXiv AI
2d ago

LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios

The article surveys LLM-based agentic reasoning frameworks, presenting a unified formal language that categorizes them into single-agent, tool-based, and multi-agent methods. It reviews application scenarios in scientific discovery, healthcare, software engineering, society, economics, and general-purpose tasks, and compares the distinct features and evaluation strategies of each category. The survey highlights the rapid development of complex agentic systems in real-world contexts.

By Bingxi Zhao, Lin Geng Foo, Ping Hu, Christian Theobalt, Hossein Rahmani, Jun Liu
Hugging Face Trending Papers
Sep 24

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

ExplorationBench is a benchmark designed to evaluate AI systems’ ability to conduct scientific exploration in verifiable alien worlds, where rules are executable and can be precisely checked. It consists of two sandboxes—AlienCode and AlienLogic—each offering discovery targets, tasks, flawed manuals, environmental feedback, and tool‑call schemas. The benchmark tests whether systems can generate new hypotheses, design experiments, and iterate on results, rather than merely recalling pre‑trained knowledge, and finds that top performers can acquire and apply unfamiliar rules, though performance varies across exploration trajectories.

arXiv AI
Sep 25

ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds

ExplorationBench is a new benchmark designed to evaluate AI systems’ ability to conduct scientific exploration in verifiable alien worlds. It comprises two sandbox environments—AlienCode and AlienLogic—each containing discovery targets, tasks, flawed manuals, and tool‑call schemas that force systems to formulate hypotheses, design experiments, and iterate on results. Ten AI systems were tested, revealing that while the best performers can learn and apply unfamiliar rules, their progress varies across exploration trajectories and can even regress with continued exploration.

By Ming Zhang, Zhenghao Xiang, Peizhong Gao, Yujiong Shen, Yuhui Wang, Zhonghan Yue, Shihan Dou, Zhangyue Yin, Junjie Ye, Shichun Liu, Weihuang Zheng, Jiahao Chen, Jiayi Chen, Hongzhang Liu, Jiaqi Shao, Tao Gui, Qi Zhang, Xuanjing Huang, Suncong Zheng, Maxm Pan