arXiv Computation and Language

Evaluation of OpenAI o1: Opportunities and Challenges of AGI

The study evaluates OpenAI’s o1‑preview large language model on a wide range of complex reasoning tasks across domains such as computer science, mathematics, natural sciences, medicine, linguistics, and social sciences. It reports high success rates, including 83.3% on competitive programming, 100% on high‑school math reasoning, and superior performance in radiology reporting, chip design, anthropology, geology, quantitative investing, and social media analysis. While the model excels at intricate reasoning and knowledge integration, it still shows occasional errors on simpler problems and struggles with some highly specialized concepts.

arXiv AI
Jul 20

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

arXiv:2607. 15686v1 Announce Type: new Abstract: We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation.

By Jiahao Zhao, Junyi Liu, Lifeng Xu, Nan Xu, Qingli Wang, Qingxiao Li, Tianle Chen, Xiaoyu Wu, Yawen Zheng, Zikai Wang, Guanming Liu, Hequn Zhou, Jingyi Wang, Jingyuan Shu, Keqi Wang, Li He, Songyang Diao, Wenhui Xu, Xinyu Ren, Yaqin Fan, Yujin Zhou, Zhanao Yao
OpenAI Blog
Dec 11, 2025

Ten years

OpenAI reflects on ten years of progress, from early research breakthroughs to widely used AI systems that reshaped what’s possible. We share lessons from the past decade and why we remain optimistic about building AGI that benefits all of humanity.

arXiv AI
Jun 2

AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

arXiv:2603. 19005v2 Announce Type: replace-cross Abstract: Data science plays a critical role in transforming complex data into actionable insights across numerous domains.

By An Luo, Jin Du, Xun Xian, Robert Specht, Fangqiao Tian, Ganghua Wang, Xuan Bi, Charles Fleming, Ashish Kundu, Jayanth Srinivasa, Mingyi Hong, Rui Zhang, Tianxi Li, Galin Jones, Jie Ding
arXiv Machine Learning
Sep 22

UniK: Universal Knowledge Perception for Digital and Physical AI

The paper introduces UniK, a universal knowledge perception platform designed to serve both digital AI—such as chatbots and agent workflows—and physical AI, which controls robots and autonomous systems. UniK handles the entire knowledge lifecycle—ingestion, enrichment, indexing, retrieval, and continuous evaluation—across diverse modalities including text, video, molecular data, and sensor telemetry, without task‑specific fine‑tuning. In five digital AI domains, UniK paired with a 70‑billion‑parameter model matches or surpasses larger proprietary LLMs, achieving high retrieval‑augmented generation accuracy on government data, medical QA, and chemistry tasks, and it also addresses similar data challenges in physical AI world‑model training.

By Nirmit Desai, Kunal Sawarkar, Aditya Mahakali, Dongkon Lee, Kevin Park, Eric Song
arXiv AI
Sep 10

Open Tabular Insight Extraction: Where Do We Stand, and Where Should We Go?

The paper introduces Open Tabular Insight Extraction (OpenTI), a unified framework aimed at democratizing access to insights from large table corpora. It highlights how current research is fragmented across domains like table QA, text‑to‑SQL, and data analysis agents, and shows that existing systems and benchmarks fall short of covering the full end‑to‑end scope of OpenTI. The authors propose a consolidated terminology, conduct a systematic review, and outline a research agenda for developing comprehensive OpenTI systems, evaluation methods, and interaction paradigms.

By Daniel Gomm, Maarten de Rijke, Madelon Hulsebos
arXiv AI
Aug 18

DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

arXiv:2505. 14107v5 Announce Type: replace-cross Abstract: The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios.

By Yakun Zhu, Zhongzhen Huang, Linjie Mu, Yutong Huang, Wei Nie, Jiaji Liu, Shaoting Zhang, Pengfei Liu, Xiaofan Zhang