arXiv Machine Learning

UniK: Universal Knowledge Perception for Digital and Physical AI

The paper introduces UniK, a universal knowledge perception platform designed to serve both digital AI—such as chatbots and agent workflows—and physical AI, which controls robots and autonomous systems. UniK handles the entire knowledge lifecycle—ingestion, enrichment, indexing, retrieval, and continuous evaluation—across diverse modalities including text, video, molecular data, and sensor telemetry, without task‑specific fine‑tuning. In five digital AI domains, UniK paired with a 70‑billion‑parameter model matches or surpasses larger proprietary LLMs, achieving high retrieval‑augmented generation accuracy on government data, medical QA, and chemistry tasks, and it also addresses similar data challenges in physical AI world‑model training.

arXiv AI
Jun 2

AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

arXiv:2603. 19005v2 Announce Type: replace-cross Abstract: Data science plays a critical role in transforming complex data into actionable insights across numerous domains.

By An Luo, Jin Du, Xun Xian, Robert Specht, Fangqiao Tian, Ganghua Wang, Xuan Bi, Charles Fleming, Ashish Kundu, Jayanth Srinivasa, Mingyi Hong, Rui Zhang, Tianxi Li, Galin Jones, Jie Ding
arXiv AI
Jul 8

SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation

arXiv:2607. 05943v1 Announce Type: new Abstract: Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing pipelines construct training data, search environments, and reward signals independently, causing synthesized structural metadata to be discarded, environments to rely on irreproducible external engines, and RL rewards to remain sparse at the trajectory level.

By Zhengbo Jiao, Yiming Cheng, Yilei Jiang, Kaituo Feng, Rui Huang, Tianyi Jiang, Juanxi Tian, Jiapeng li, Qunzhong Wang, Tailai Chen, Qianshan Wei, Chuan Xiao, Shanyu Rong, Yangfu Li, Yanhan Zhou, Yunpu Ma, Yifan Zhang, Xiangyu Yue
arXiv Computation and Language
Sep 22

MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine

arXiv:2603.00842v2 Announce Type: replace Abstract: Biomedical multimodal assistants have the potential to unify radiology, pathology, and clinical-text reasoning, yet a critical deployment gap remai...

By Kai Zhang, Zhengqing Yuan, Cheng Peng, Songlin Zhao, Mengxian Lyu, Ziyi Chen, Yanfang Ye, Wei Liu, Ying Zhang, Kaleb E Smith, Lifang He, Lichao Sun, Yonghui Wu
arXiv AI
Sep 23

ORDER: A Fictitious-World Benchmark for Domain-Adaptive Embodied AI

The paper introduces ORDER, a fictitious-world benchmark designed to evaluate domain-adaptive embodied AI. ORDER consists of a synthetic 342,069-token corpus defining a self-consistent physics, a 500-question knowledge test (ORDER‑BENCH), and a compositional spatial task (ORDER‑SPATIAL) that requires ordering objects for safe manipulation. The benchmark demonstrates that models like GPT‑4.1 perform poorly without adaptation, while small models improve significantly after continual pre‑training, and that performance on ORDER‑SPATIAL better predicts real plan quality than knowledge-test accuracy.

By Sai Krishna Reddy Sathi, Anuj Tiwari
arXiv AI
Sep 12

From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development

The paper introduces a modular agentic-AI platform that transforms heterogeneous CMC process-development documents into a dual-layer knowledge graph. The base layer creates a lexical Document‑Section‑Chunk hierarchy, while the intelligence layer extracts ontology‑aligned entities and links cross‑document concepts, all anchored by provenance. LLM agents navigate these layers to answer queries, and a novel three‑tier evaluation protocol demonstrates high retrieval‑augmented generation performance on proprietary data from a Sanofi program.

By Reza Amirmoshiri, Faryad Sahneh, Yasser Jangjou
arXiv AI
Aug 20

SkillNet: Create, Evaluate, and Connect AI Skills

SkillNet is an open infrastructure that enables the creation, evaluation, and organization of AI skills at scale. It structures skills within a unified ontology, supports multi‑dimensional evaluation (Safety, Completeness, Executability, Maintainability, Cost‑awareness), and integrates a repository of over 600,000 skills, an interactive platform, and a Python toolkit. Experiments on ALFWorld, WebShop, and ScienceWorld demonstrate a 40% increase in average rewards and a 30% reduction in execution steps across multiple backbone models.

By Yuan Liang, Ruobin Zhong, Haoming Xu, Chen Jiang, Yi Zhong, Runnan Fang, Jia-Chen Gu, Shumin Deng, Yunzhi Yao, Mengru Wang, Shuofei Qiao, Yida Xue, Xin Xu, Tongtong Wu, Kun Wang, Yang Liu, Zhen Bi, Jungang Lou, Yuchen Eleanor Jiang, Hangcheng Zhu, Gang Yu, Haiwen Hong, Longtao Huang, Hui Xue, Chenxi Wang, Yijun Wang, Zifei Shan, Xi Chen, Zhaopeng Tu, Feiyu Xiong, Xin Xie, Peng Zhang, Zhengke Gui, Lei Liang, Jun Zhou, Chiyu Wu, Jin Shang, Yu Gong, Junyu Lin, Changliang Xu, Hongjie Deng, Wen Zhang, Keyan Ding, Qiang Zhang, Fei Huang, Ningyu Zhang, Jeff Z. Pan, Guilin Qi, Haofen Wang, Huajun Chen
arXiv Computation and Language
Sep 10

Benchmarking Hybrid Deep Research Across Database Querying and Web Search

arXiv:2609.09410v1 Announce Type: new Abstract: While autonomous agents have made significant strides in "deep research" by iteratively navigating the open web to synthesize information, real-world p...

By Ruofan Wu, Peiran Xu, Xiaolong Li, Fan Shu, Soyoung Yoon, Yite Wang, Xiaodong Yu, Boyi Liu, Feng Yan, Debiao Li, Yuxiong He, Zhewei Yao
arXiv AI
Sep 15

ClinAgent: A ReAct-Based Agent for Conversational Access to Clinical Trial Information

ClinAgent is a conversational system that uses a ReAct-based LLM agent to retrieve and synthesize clinical trial information from multiple sources such as ClinicalTrials.gov, PubMed, and a local dataset. The agent iteratively reasons over user queries, selects appropriate tools, and refines its actions to provide grounded, up-to-date responses in natural language across multi-turn interactions. Evaluation across three phases shows that DeepSeek (thinking mode) excels in planning quality while Gemini 3.0 Flash delivers the highest overall performance and expert ratings, demonstrating the promise of agentic AI for improving clinical trial data access.

By Antonino Vaccarella, Riccardo Cantini, Domenico Talia, Paolo Trunfio, Marianna Talia, Rosamaria Lappano, Marcello Maggiolini