KnowSim introduces an evaluation framework that uses a user simulator with explicit knowledge states to assess how well large language models calibrate information to users. The simulator represents knowledge as a graph of Information Units with prerequisite relationships and updates these states based on learning theory. KnowSim computes Knowledge Gain, Delivery Calibration, and Cognitive Overload metrics, and its rankings align with human judgments, outperforming baseline simulators and revealing model performance differences across user knowledge levels.
By Yoonjoo Lee, Hyoungwook Jin, Tae Soo Kim, Shaoyang Zhang, Philippe Laban, Q. Vera Liao
arXiv:2606. 15225v1 Announce Type: cross Abstract: Large-scale learner-task interaction data are crucial for intelligent educational systems but are costly to collect and constrained by privacy and learner engagement.
By Weibo Gao, Qi Liu, Linan Yue, Zheng Zhang, Yichao Du, Fangzhou Yao, Ao Yu, Zhenya Huang, Shijin Wang
arXiv:2607. 21306v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems.
By Verona Teo, Raghav Jain, Tobias Gerstenberg, Max Kleiman-Weiner
arXiv:2608. 00007v1 Announce Type: cross Abstract: Equipping Large Language Models (LLMs) with human-like personas is crucial for agentic applications, such as role-play and user simulation.
By Bohan Tang, Yiwen Guo
arXiv:2608. 10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them.
By Rose Niousha, Minwoo Kang, Narges Norouzi
arXiv:2609.00608v1 Announce Type: new
Abstract: LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, the...
By Daeheon Jeong, Yoonjoo Lee, Eugene Choi, Sinie van der Ben, Juho Kim
The paper evaluates Large Language Models for automatically analyzing responses in digital teacher simulations. Experiments compare DeBERTaV3 and Llama 3 across zero‑shot, few‑shot, and fine‑tuning settings, revealing that performance varies by characteristic and that Llama 3 consistently outperforms DeBERTaV3, especially when new characteristics must be identified. The findings suggest Llama 3 is preferable for dynamic simulation environments where teacher educators introduce new evaluation criteria.
By David de-Fitero-Dominguez, Mariano Albaladejo-Gonz\'alez, Antonio Garcia-Cabot, Eva Garcia-Lopez, Antonio Moreno-Cediel, Erin Barno, Justin Reich
The paper introduces SIC-Agents, a self‑improving framework designed to enhance simulation for pediatric serious illness communication (SIC) training. It presents two new benchmark suites—PitfallBench and DialogueBench—that assess simulators at both turn‑level and full‑dialogue levels, specifically addressing the unique challenges of multi‑party interactions and parental distress. Experiments demonstrate that SIC‑Agents surpasses static expert prompting, and the authors release the benchmarks for broader research use.
By Zihan Wang, Anita Marie Slominska, Rennie Bimman, Elizabeth Di Flumeri, Amanda Mayappo-Neeposh, Conall Francoeur, Tamara Ellen Carver, Xiao-Wen Chang, Doina Precup, Esin Darici Haritaoglu, Ismail Haritaoglu, Akshatha Arodi, Naomi Goloff
arXiv:2606. 06546v1 Announce Type: new Abstract: Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know.
By Tao Liu, Ye Lu, Ruohua Zhang, Siyu Song, Wentao Liu, Aimin Zhou, Hao Hao
arXiv:2605.30051v2 Announce Type: replace
Abstract: A key part of developing large language model (LLM)-powered, automated tutoring tools is student simulation, i.e., using LLMs to role-play as stude...
By Zhangqi Duan, Shuyan Huang, Alexander Scarlatos, Jaewook Lee, Simon Woodhead, Andrew Lan
The article reports that large language models can predict and collaboratively modulate human memory search during a semantic fluency task. By tracking and forecasting participants’ semantic retrieval patterns, the models outperform other humans in following these mental trajectories. This suggests that AI can serve as a cognitive tool to extend human abilities in open‑ended conceptual exploration and creative ideation.
By Eric Lacosse, Mariana Duarte, Graham Todd, Peter M. Todd, Daniel C. McNamee
arXiv:2605. 24828v2 Announce Type: replace Abstract: With the continuous advancement of Large Language Models (LLMs), intelligent agents are becoming increasingly vital.
By Wentong Chen, Xin Cong, Zhong Zhang, Yaxi Lu, Siyuan Zhao, Yesai Wu, Qinyu Luo, Haotian Chen, Yankai Lin, Zhiyuan Liu, Maosong Sun