arXiv AI

ASCEND: Personal AI Agents for Autonomous Scientific Computing Across HPC Clusters and GPU Workstations

arXiv AI
5d ago

Can AI Scientists Coordinate at Runtime?

The paper "Can AI Scientists Coordinate at Runtime?" introduces Runtime Agent Coordination (RAC), a system that dynamically selects agents from existing AI‑scientist hosts during execution, assigns scoped work contracts, and provides artifact‑grounded verification. An exploratory evaluation on ResearchClawBench across three hosts—Agent Laboratory, EvoScientist, and ARK—shows that runtime selection improves performance, while adding contracts and verification can reduce scores depending on the host. The study highlights the potential and limitations of runtime coordination under budget constraints.

By Zijian Liu, Yangzhixin Luo, Junyu Lu, Yi Li, Yu Chen, David Xu, William F. Shen, Xinchi Qiu, Xisen Wang
arXiv AI
Sep 15

OpenAl4S: Code as Action, Science as Sessions

OpenAI4S is an open‑source scientific research agent that treats code as action and science as sessions, combining a persistent computing runtime with structured session management. It uses tool calls for orchestration, executes code cells in persistent Python and R kernels, and records an append‑only Action Ledger, per‑cell execution logs, versioned artifacts, environment snapshots, and workspace checkpoints to preserve provenance and enable session recovery, branching, and extension. Evaluated on 36 research scenarios—including retrosynthesis, molecular dynamics, and protein design—OpenAI4S achieved a higher overall score (7.83) than a general‑purpose coding harness, especially on long‑horizon, computation‑intensive workflows, though reproducibility remains an open challenge. whyItMatters":"The system demonstrates that persistent execution coupled with session‑level provenance can enhance the reliability of AI‑assisted scientific workflows, as evidenced by its superior performance across diverse research scenarios."

By Gongbo Zhang, Hao Li, Yu Wang, Mujie Lin, Liuzhenghao Lv, Yicheng Mao, Yimi Wang, Jun Zhu, Minhan Tang, Zhengxiang Jiang, Yusong Wang, Jiayu Yao, Kunpeng Ning, Dawei Pang, Yonghong Tian, OpenAI4S Community, Yuyang Liu, Li Yuan
arXiv AI
Aug 12

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

arXiv:2608. 10366v1 Announce Type: new Abstract: Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments.

By Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince
arXiv AI
Jun 16

The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution

arXiv:2605. 27599v2 Announce Type: replace-cross Abstract: Agentic AI workloads - where a single user goal triggers multi-step orchestration, tool calls, retries, and failure recovery - are being targeted for edge deployment, with NVIDIA, Dell, HP, ASUS, MSI, Acer, and Gigabyte all shipping GB10-based desktop AI systems in 2026.

By Deepak Panigrahy, Aakash Tyagi
arXiv AI
Sep 10

Diamond Agent: Agentic Control of Federated HPC Resources as a Service

arXiv:2609.06181v1 Announce Type: cross Abstract: Efficiently aggregating and orchestrating computing power across heterogeneous clusters for HPC workflows faces four practical challenges: preserving...

By Haotian Xie, Junlin Chen, Mingkai Zheng, Yifan Zhu, Minu Mathew, Max Burnette, Yadu Babuji, Volodymyr Kindratenko, Shivaram Venkataraman, Kyle Chard, Ian Foster, Zhao Zhang
arXiv AI
Sep 17

AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

AutoTuneBench introduces a trustworthy measurement protocol for evaluating how large language model agents auto‑tune GPU kernels and serving engines. The benchmark addresses four failure modes—strawman baselines, machine‑dependent timing, saturated tasks, and infrastructure defects—by enforcing code‑frozen protocols, database validation, anti‑cheat checks, pre‑registered comparisons, and external result anchoring. Using this protocol, the authors demonstrate that previously reported speedups are inflated, revealing more modest improvements across different engines and machines.

By Li Chen
arXiv AI
Sep 4

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

SENTINEL‑RL is an agentic SOC architecture that separates topological reasoning from semantic reasoning. It uses a heterogeneous graph attention encoder to compress a large authentication subgraph into a fixed‑dimensional state, a PPO policy to map that state to constrained investigative actions, and an LLM loop that only consumes policy recommendations and produces analyst‑readable narratives. Experiments on LANL and Indiana University datasets show fast graph ingestion, reliable alerting, high PPO performance, and a median 6.3‑second containment cycle.

By Uday Vallabhaneni, Cassie L. Cagwin, David J. Wild