ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge
arXiv:2609. 12366v1 Announce Type: new Abstract: We present ORQA, a method for testing occupation-level knowledge in large language models.
arXiv:2607. 16057v1 Announce Type: cross Abstract: Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such as factual recall, narrow question answering, mathematical problem-solving, and coding and agentic tool-use.
arXiv:2609. 12366v1 Announce Type: new Abstract: We present ORQA, a method for testing occupation-level knowledge in large language models.
arXiv:2609.21841v1 Announce Type: new Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
DataCanvas-EDU is an agentic framework that lets instructors guide the creation of synthetic datasets for business analytics courses. Instructors set teaching goals and desired patterns via conversation, and an AI agent writes generation code, verifies the data, and produces assignments, reference solutions, and rubrics. The process is organized into four phases—Plan, Create, Verify/Test Analysis, and Evaluate—to streamline case preparation and enable students to explore new patterns with AI.
The article presents a minimal working model for large language model (LLM) systems, emphasizing four key distinctions—pretraining vs. deployment, distribution vs. samples, types of memory, and task competence vs. agency. Using this framework, it diagnoses six common misconceptions about LLMs (next‑token prediction, regression to the mean, training‑data regurgitation, model memory, alignment, and understanding), explaining what each misconception captures correctly, where it conflates distinctions, and the implications for evaluation, design, and governance. The model is applied to AI policy language, illustrating how policy can misrepresent these distinctions and offering a diagnostic toolkit to correct such errors.
arXiv:2608. 14212v1 Announce Type: new Abstract: As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses.
arXiv:2606. 16206v1 Announce Type: new Abstract: Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learning support.
arXiv:2608. 04549v1 Announce Type: cross Abstract: Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on.
arXiv:2602. 07840v3 Announce Type: replace-cross Abstract: Evaluating relevance in large-scale search systems is fundamentally constrained by the governance gap between nuanced, resource-constrained human oversight and the high-throughput requirements of production systems.
arXiv:2609.39846v1 Announce Type: cross Abstract: We investigate the problem of role-capability leakage (RCL), in which a role-prompted reasoning model generates convincing in-role text while continu...
The paper introduces the Grounded Integration Measure (GIM), a benchmark of 820 expert‑authored problems designed to test models on tasks that integrate multiple cognitive operations such as constraint satisfaction, state tracking, epistemic vigilance, and audience calibration. GIM emphasizes realistic, broadly accessible knowledge rather than specialized expertise, and uses a judge‑aware 2‑parameter logistic IRT model to produce robust ability estimates across 53 model‑thinking‑level configurations. The authors provide a comprehensive leaderboard of 22 models and 47 test configurations, and conduct an extensive study on how test‑time compute affects model capability, finding that configuration choices like thinking budget and quantization can be as influential as model selection itself. whyItMatters":"By focusing on integration of multiple cognitive domains, GIM offers a more realistic assessment of model reasoning capabilities than benchmarks that either overemphasize memorization or abstract reasoning alone."
arXiv:2606.16206v2 Announce Type: replace Abstract: Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learni...