UniDataAgent (UniDataAgent) is an ontology‑grounded system designed to automate enterprise question‑to‑report tasks while preserving organization‑specific semantics. It separates semantic acquisition from online execution, with an Ontology Acquisition and Validation (OAV) stage that builds versioned ontologies from metadata, business knowledge, and expert input, and a Question‑to‑Report Execution (QRE) stage that retrieves semantic contracts, coordinates skills and data tools, validates results, and produces evidence‑linked reports. In a deployment across 27 enterprise tables and thousands of metric types, ontology construction took a few hours versus a week manually, and report generation took minutes versus several working days, achieving 95.0% strict accuracy on real business questions compared to 72.5% for document RAG.
By Yutai Duan, Yahui Zhao, Zhangti Li, Yu Ma, Zhenfeng Qi, Shaoyang Yuan, Jing Fan, Jie Liu
arXiv:2608.28594v1 Announce Type: new
Abstract: Conversational analytics systems assume the user already has a well-formed question, leaving a non-expert facing a blank query box on an unfamiliar ent...
By Harmohit Singh, Rahul Sharma
arXiv:2606. 09809v1 Announce Type: new Abstract: AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs.
By Avijit Ghosh, Anka Reuel, Jenny Chim, Wm. Matthew Kennedy, Srishti Yadav, Jennifer Mickel, Yanan Long, Andrew Tran, Anastassia Kornilova, Damian Stachura, Kevin Klyman, Felix Friedrich, Jeba Sania, Max Lamparth, Jan Batzner, Anoop Mishra, Eliya Habba, Yixiong Hao, Nathan Heath, Shalaleh Rismani, Usman Gohar, Andrea Loehr, David Manheim, Ruchira Dhar, Sree Harsha Nelaturu, Aarush Sinha, Leshem Choshen, Drishti Sharma, Ishan Khire, Amit Saha, Subramanyam Sahoo, Michael Hardy, Michael Alexander Riegler, Kabir Manghnani, Michelle Lin, Yanan Jiang, Yilin Huang, Asaf Yehudai, Jessica Ji, Aris Hofmann, Mubashara Akhtar, Nuno Moniz, Yacine Jernite, Stella Biderman, Zeerak Talat, Sanmi Koyejo, Mykel Kochenderfer, Irene Solaiman
Benchmark Radar is a living database and search engine that aggregates AI benchmark papers, datasets, code, and score histories. It automatically discovers new benchmark resources from 37 sources, maintains a catalog of 1,283 records with 12,916 numeric observations, and provides tools such as a web dashboard, CLI, and downloadable evidence for researchers. The system also offers visualizations like a Pareto frontier and trend views to help users assess benchmark saturation and adoption.
By Koutian Wu, Junjie Zhou, Ergan Shang, Jiayu Wang, Pengqian Han, Junkai Wang, Wanghan Xu
OpenAI for Law brings frontier intelligence for law, custom firm workflows, connected legal data sources, and legal-grade controls for confidential client work.
arXiv:2606. 23533v2 Announce Type: replace Abstract: Recent large language models (LLMs) are good at general text generation, but it is still hard to use them for domain-specific data generation because the output must follow strict formatting and structural rules.
By Hung Phan, Aniroop Naladala, Dubey Avanindra, Supryia Chinthavali, Lunga Dalton, Ali Jannesari
arXiv:2607. 25042v1 Announce Type: new Abstract: The evolution of customer support systems is rapidly advancing with agentic chatbots, yet these systems face significant limitations when accessing enterprise data without predefined API endpoints.
By Bhanu Teja Rangaraju, Chandan Kumar
OpenAI built a system to extract contract data quickly, cutting turnaround times and making it easier for teams to access the details they need.
arXiv:2608.22817v1 Announce Type: new
Abstract: Industrial technical reports contain high-value knowledge for maintenance, troubleshooting, and product engineering, but their heterogeneous structure...
By Parsa Bakhtiari, Hassan Bashiri, Alireza Khalilipour, Masoud Nasiripour, Moharram Challenger
arXiv:2511. 07322v3 Announce Type: replace-cross Abstract: While LLMs have shown great success in financial tasks like stock prediction and question answering, their application in fully automating Equity Research Report generation remains uncharted territory.
By Song Jin, Shuqi Li, Shukun Zhang, Rui Yan
arXiv:2606. 23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language models on financial tasks.
By Mostapha Benhenda