arXiv AI

Agri-SAGE: Simulation-Grounded Multi-Agent LLM for Context-Aware Agricultural Advisory Generation

arXiv:2607. 00454v1 Announce Type: new Abstract: Agricultural advisory systems face a fundamental tension: static agronomic guidelines offer consistent, evidence-based recommendations, yet remain blind to in-season variability and dynamic uncertainties.

arXiv Machine Learning
Jun 29

Non-Linear Model-Based Sequential Decision-Making in Agriculture

arXiv:2509. 01924v4 Announce Type: replace-cross Abstract: Agricultural decision-making faces a dual challenge: sustaining high yields to meet global food security needs while reducing the environmental impacts of input use, including fertilizer losses and other agrochemical applications such as herbicides, insecticides, and fungicides.

By Sakshi Arya, Wentao Lin
arXiv AI
2d ago

Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control

Mimir is a physics‑grounded large language model agent designed for long‑horizon irrigation control. It operates on two timescales: a fast scale that uses a structured physical interface and deterministic simulator to validate and refine LLM proposals before execution, and a slow scale that consolidates recurrent failure patterns into persistent contextual principles. Across multiple sites, crops, and years, Mimir achieves the lowest aggregate control cost and reduces irrigation usage by about 51% compared to historical schedules, while ablation studies confirm the importance of forward simulation, verified revision, and persistent context.

By Yimeng Liu, Mi Zhang, Younsuk Dong, Zhichao Cao
arXiv AI
Sep 2

Deploying and Evaluating a Smart-Agriculture Agentic Engine for Full-Season Soybean Farm Operations

The paper introduces FAIRY, a full-stack smart‑agriculture agent system deployed on a soybean research farm at Harbin Institute of Technology. FAIRY executes and evaluates end‑to‑end agronomic operations—from ridge preparation to storage—using an event‑driven world model that integrates machinery, sensors, drones, satellite data, weather, crop models, and historical yields. The system implements a comprehensive agentic stack and is used to benchmark nine state‑of‑the‑art agent controllers across 100 full‑season soybean scenarios, assessing success, spatiotemporal correctness, token cost, and edge‑device runtime.

By Ao Qu, Panagiotis Michelakis, Linyuan Han, Yiannis Hadjiyianni, Kun Ouyang, Konstantinos Siskos, Feng Li, Ran Meng, Jingchi Jiang, Dimitrios Stamoulis, Jie Liu
Hugging Face Trending Papers
Jun 30

An Agentic AI Framework to Accelerate Scientific Discovery in Plant Phenotyping

High-throughput plant phenotyping now generates image derived datasets far faster than scientists can analyze them. At Oak Ridge National Laboratory's Advanced Plant Phenotyping Laboratory (APPL), automated stations image hundreds of plants daily across multiple remote sensing modalities; yet, trait extraction and interpretation remain manual, expert-bound, and strictly post-hoc, making analysis, not acquisition, the binding constraint on discovery.

arXiv AI
Aug 25

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

arXiv:2608.23525v1 Announce Type: new Abstract: Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards ma...

By Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang, Yuanzhe Hu, Siru Zhong, Weidong Tang, Yuxuan Liang, Weijia Li, Ming Jin, Shirui Pan, Yuhao Kang, Dingyi Zhuang, Jinhua Zhao
arXiv AI
Sep 23

Agentic Explainable Artificial Intelligence (Agentic XAI) Approach To Explore Better Explanation: A Case Study in Decision Support for Rice Cultivation in Japan

arXiv:2512. 21066v4 Announce Type: replace Abstract: Explainable artificial intelligence (XAI) reveals how explanatory variables relate to a response variable, yet communicating XAI outputs to laypersons remains difficult, limiting trust in AI-based predictions.

By Tomoaki Yamaguchi, Yutong Zhou, Masahiro Ryo, Keisuke Katsura
arXiv AI
6d ago

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.

By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth
arXiv AI
2d ago

LLM-based Agentic Reasoning Frameworks: A Survey from Methods to Scenarios

The article surveys LLM-based agentic reasoning frameworks, presenting a unified formal language that categorizes them into single-agent, tool-based, and multi-agent methods. It reviews application scenarios in scientific discovery, healthcare, software engineering, society, economics, and general-purpose tasks, and compares the distinct features and evaluation strategies of each category. The survey highlights the rapid development of complex agentic systems in real-world contexts.

By Bingxi Zhao, Lin Geng Foo, Ping Hu, Christian Theobalt, Hossein Rahmani, Jun Liu
arXiv Machine Learning
Sep 11

Studying Without a Syllabus: Task-Agnostic Environment Preprocessing

The paper investigates whether a language‑model agent can autonomously study an unfamiliar environment without prior task instructions or examples, and decide how to prepare for future tasks. It formalizes task‑agnostic environment preprocessing, where a studying system explores under a budget to produce reusable artifacts for a later solver. Experiments on six diverse benchmarks show that a meta‑agent variant often outperforms fixed methods, though larger budgets do not consistently boost downstream reward, yet the artifacts still reduce test‑time sampling needed to achieve a target score.

By Vinay Samuel, Varun Ursekar, Vijay S. Kalmath, Apaar Shanker, Veronica Chatrath, Yuan Xue