arXiv:2609. 12366v1 Announce Type: new Abstract: We present ORQA, a method for testing occupation-level knowledge in large language models.
By Shreyas Krishnan, Serina Chang, Abhishek Nagaraj
arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao
DataCanvas-EDU is an agentic framework that lets instructors guide the creation of synthetic datasets for business analytics courses. Instructors set teaching goals and desired patterns via conversation, and an AI agent writes generation code, verifies the data, and produces assignments, reference solutions, and rubrics. The process is organized into four phases—Plan, Create, Verify/Test Analysis, and Evaluate—to streamline case preparation and enable students to explore new patterns with AI.
By Bang An, Maria Hamdani, Joseph Fox
The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.
By Zhicheng Lin
arXiv:2608. 14212v1 Announce Type: new Abstract: As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses.
By Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou, Xiaoyu Shi, Jiansheng Cai, Bo Zhang, Zhe Li, Xu-Yao Zhang