arXiv AI

An Agentic Framework for Neuro-Symbolic Programming

The paper introduces AgenticDomiKnowS (ADS), a framework that converts free‑form task descriptions into fully functional DomiKnowS neuro‑symbolic programs. ADS employs an agentic workflow that builds and tests each DomiKnowS component independently, optionally allowing human‑in‑the‑loop refinement. The authors demonstrate that ADS enables both experienced and novice users to create complete neuro‑symbolic programs in 10–15 minutes, compared to the hour required for manual coding.

Hugging Face Trending Papers
Jul 5

Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming

Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's context window. The prevailing route to improve such reasoning is test-time scaling, which trains models to search over long chains of thought; but the resulting capability is entangled in model weights, is not verifiable step-by-step, and is costly at inference.

arXiv AI
2d ago

Auto-Formalizing Neuro-Symbolic Predictors

arXiv:2610.01519v1 Announce Type: cross Abstract: Neuro-Symbolic (NeSy) predictors incorporate prior knowledge into the prediction process of neural networks, ensuring that outputs satisfy specified...

By Samuele Bortolotti, Weixin Chen, Han Zhao, Andrea Passerini, Stefano Teso, Antonio Vergari
arXiv AI
Sep 2

UI-Venus-2 Technical Report

UI‑Venus‑2 is a general‑purpose foundation GUI agent that operates across mobile, web, and desktop environments using a unified closed‑loop reasoning‑action framework. The report details how the system expands environment coverage to over 170 multilingual mobile apps and native desktop OSes, scales task generation through a deep‑research pipeline, and enhances verification with trace‑level and sample‑level evaluators that use visual keypoints and multi‑model voting. Safety‑aware mechanisms are also incorporated to control consequential actions, positioning UI‑Venus‑2 as an efficient, open‑source tool for more generalizable, verifiable, and self‑reflective agents in real‑world applications.

By Venus Team, Zhuohan Cai, Haoxing Chen, Jiaxuan Chen, Weizhi Chen, Changlong Gao, Zhangxuan Gu, Yuan Guo, Yusong Hu, Jianrong Jiang, Jianguo Li, Runze Li, Jinzhen Lin, Zhenyu Ma, Changhua Meng, Han Peng, Xinyu Qiu, Shuheng Shen, Zhongyi Shui, Weiqiang Wang, Ming Wen, Zhuoer Xu, Hang Yan, Kaiwen Yang, Ruilin Yao, Nanjun Yu, Zhengwen Zeng, Lianrui Zhang, Yunzhu Zhang, Zhe Zhao, Beitong Zhou
arXiv AI
Sep 1

WebXSkill: Skill Learning for Autonomous Web Agents

arXiv:2604.13318v2 Announce Type: replace Abstract: Autonomous web agents powered by large language models (LLMs) remain brittle on long-horizon browser workflows. A key bottleneck is a grounding gap...

By Zhaoyang Wang, Qianhui Wu, Xuchao Zhang, Chaoyun Zhang, Wenlin Yao, Fazle Elahi Faisal, Baolin Peng, Si Qin, Suman Nath, Qingwei Lin, Chetan Bansal, Dongmei Zhang, Saravan Rajmohan, Jianfeng Gao, Huaxiu Yao
arXiv AI
6d ago

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.

By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth