arXiv AI By Ajay Patel, Kartik Hosanagar, Ramayya Krishnan, Chris Callison-Burch, Karim Lakhani, Mitch Weiss

Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

Read the original on arXiv AI →

arXiv:2607. 16057v1 Announce Type: cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 18

DataCanvas-EDU: An Agentic Framework for Instructor-Guided Synthetic Data Generation in Business Analytics Education

DataCanvas-EDU is an agentic framework that lets instructors guide the creation of synthetic datasets for business analytics courses. Instructors set teaching goals and desired patterns via conversation, and an AI agent writes generation code, verifies the data, and produces assignments, reference solutions, and rubrics. The process is organized into four phases—Plan, Create, Verify/Test Analysis, and Evaluate—to streamline case preparation and enable students to explore new patterns with AI.

By Bang An, Maria Hamdani, Joseph Fox
arXiv AI
Aug 24

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.

By Zhicheng Lin