AI agents

Tool use, function calling, orchestration and the protocols that let models act rather than only answer.

8,054 stories · RSS feed

arXiv AI
Jul 2

Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions

arXiv:2607. 00304v1 Announce Type: cross Abstract: The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (gamma), strategy diversity (H), and small-sample measurement reliability (CV(N)) cannot be simultaneously optimized at fixed sample size N.

By Zewen Liu
arXiv AI
Jul 2

RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation

arXiv:2607. 00147v1 Announce Type: new Abstract: Rare disease differential diagnosis is a critical yet arduous clinical task, requiring physicians to identify precise phenotypes from complex, unstructured patient symptoms and execute intricate reasoning within a vast search space.

By Deyang Jiang, Haoran Wu, Ziyi Wang, Yiming Rong, Yunlong Zhao, Ye Jin, Bo Xu
arXiv AI
Jul 2

BaRA: BFS-and-Reflection Web Data Collection Agent

arXiv:2607. 00007v1 Announce Type: cross Abstract: Large language model (LLM)-based web agents reduce manual scripting for web data collection, yet on live websites, they often miss relevant pages, return incomplete multimodal outputs, or return media URLs that are not directly downloadable.

By Soojeong Lee, Joseph Lee, Yongseong Cho, Sunjae Kim, Youngwoo Moon, Kyungwoo Song
arXiv AI
Jul 2

ASPIRE: Agentic /Skills Discovery for Robotics

arXiv:2607. 00272v1 Announce Type: cross Abstract: Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures.

By Runyu Lu, Yubo Wu, Ethan Kou, Letian Fu, Wenli Xiao, Ajay Mandlekar, Yinzhen Xu, Guanya Shi, Ken Goldberg, Ang Chen, Mosharaf Chowdhury, Yuke Zhu, Linxi "Jim" Fan, Guanzhi Wang
arXiv AI
Jul 2

RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail

arXiv:2607. 00310v1 Announce Type: cross Abstract: Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale generic video leaves them poorly aligned with real-world deployment domains.

By Amirreza Rouhi, Rajat Aggarwal, Parikshit Sakurikar, Anoop M. Namboodiri, Sashi P. Reddi
arXiv AI
Jul 2

EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes

arXiv:2607. 00547v1 Announce Type: cross Abstract: Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to evaluate egocentric perspective itself in isolation.

By Jihyeok Jung (KAIST AI), Jeewu Lee (Sogang University), Sanghyeop Kim (Sogang University), Chanhee Han (Ministry of Science and ICT), Seong Joon Oh (KAIST AI)
arXiv AI
Jul 2

LLM-Guided ODE Discovery and Parameter Inference from Small-Cohort Aggregate Data

arXiv:2607. 00733v1 Announce Type: cross Abstract: Mechanistic modeling via ordinary differential equations (ODEs) provides interpretable descriptions of complex dynamics and enables inference of underlying mechanisms, which is particularly valuable in clinical settings.

By Hanning Yang, Meropi Karakioulaki, Lennart Purucker, Tim Litwin, Cristina Has, Moritz Hess
arXiv AI
Jul 2

Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use

arXiv:2607. 01084v1 Announce Type: new Abstract: While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered by the dynamic nature of user queries, tool sets, and interaction dynamics.

By Song-Lin Lv, Weiming Wu, Rui Zhu, Zi-Jian Cheng, Lan-Zhe Guo