Hugging Face Trending Papers

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

arXiv Computation and Language
Aug 25

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

LongWoF-Bench is a new benchmark of 778 machine‑verifiable long‑workflow tasks spanning code generation, agent‑environment synthesis, mathematical reasoning, and rule following. The study shows that EvoMap Genes—structured representations of verifier‑confirmed execution trajectories—outperform the Skill baseline by 8.7–15.5 percentage points across seven models, and for Claude Opus they enable 39 additional task completions while cutting token consumption by 9.9%. The results demonstrate that verified execution experience can be externalized and reused, improving long‑workflow completion without repeatedly discovering new strategies.

By Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang
arXiv AI
Jun 2

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

arXiv:2606. 01286v1 Announce Type: cross Abstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differentiate model capabilities or provide useful training signal.

By Yangzhen Wu, Aaron J. Li, Wenjie Ma, Li Cao, Ziheng Zhou, Mert Cemri, Shu Liu, Yuran Xiu, Chenxiao Yan, Haikun Zhao, Bin Yu, Ion Stoica, Dawn Song
arXiv AI
Jul 24

Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills

arXiv:2607. 20999v1 Announce Type: new Abstract: Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not jointly resolve where a failure occurs in a workflow, which mechanism caused it, and how relevant knowledge from third-party Skills should be reused locally.

By Zibin Lin, Shengli Zhang, Taotao Wang, Yihan Xia, Deen Ma, Guofu Liao
arXiv AI
Jul 21

DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines

arXiv:2607. 16617v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable platform artifacts.

By Runming He, Zhen Hao Wong, Hao Liang, Zimo Meng, Chengyu Shen, Xiaochen Ma, Wentao Zhang
arXiv AI
Sep 15

OpenAl4S: Code as Action, Science as Sessions

OpenAI4S is an open‑source scientific research agent that treats code as action and science as sessions, combining a persistent computing runtime with structured session management. It uses tool calls for orchestration, executes code cells in persistent Python and R kernels, and records an append‑only Action Ledger, per‑cell execution logs, versioned artifacts, environment snapshots, and workspace checkpoints to preserve provenance and enable session recovery, branching, and extension. Evaluated on 36 research scenarios—including retrosynthesis, molecular dynamics, and protein design—OpenAI4S achieved a higher overall score (7.83) than a general‑purpose coding harness, especially on long‑horizon, computation‑intensive workflows, though reproducibility remains an open challenge. whyItMatters":"The system demonstrates that persistent execution coupled with session‑level provenance can enhance the reliability of AI‑assisted scientific workflows, as evidenced by its superior performance across diverse research scenarios."

By Gongbo Zhang, Hao Li, Yu Wang, Mujie Lin, Liuzhenghao Lv, Yicheng Mao, Yimi Wang, Jun Zhu, Minhan Tang, Zhengxiang Jiang, Yusong Wang, Jiayu Yao, Kunpeng Ning, Dawei Pang, Yonghong Tian, OpenAI4S Community, Yuyang Liu, Li Yuan
arXiv AI
Jun 9

SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?

arXiv:2606. 07682v1 Announce Type: cross Abstract: AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments.

By Rishi Desai, Jesse Hu, Joan Cabezas, Neel Harsola, Pratyush Shukla, Roey Ben Chaim, Adnan El Assadi, Omkaar Mukund Kamath, Fenil Faldu, Prannay Hebbar, Jiankai Sun, Yiyuan Li, Pramod Srinivasan, Ishan Gupta, Christopher Settles, Daniel Wang, Derek Chen, Pranav Raja, Albert Liu, Marek \v{S}uppa, Nevasini Sasikumar, Luyang Kong, Erik Quintanilla, Xiangyi Li, Ivan Bercovich, Steven Dillmann