arXiv:2606. 07594v1 Announce Type: new Abstract: Personal AI agents must increasingly operate across APIs, shells, web surfaces, and desktop GUIs, yet many systems remain tuned to a single interface and offer limited support for user teaching and auditability.
By Bo Zhang, Borui Zhang, Chenghao Jiang, Minglei Shi, Xiaofeng Wang, Zheng Zhu, Jie Zhou, Jiwen Lu
arXiv:2609. 29672v1 Announce Type: new Abstract: Artificial intelligence helps education most where an essential provision has been rationed by cost.
By Qiming Guo, Jinwen Tang, Xingran Huang, Hung-Yu Lin, Yafu Zhong, Xiatian Zhuang
arXiv:2606. 16428v1 Announce Type: cross Abstract: Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but also dynamically adapt their instruction to diverse learners.
By Jaward Sesay, Yue Yu, Siwei Dong, Yemin Shi, Guangyao Chen, B\"orje F. Karlsson
arXiv:2606. 15766v1 Announce Type: new Abstract: A central pedagogical value evaluated in AI tutor benchmarks is scaffolding: guiding students through graduated steps toward a solution.
By Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson
arXiv:2602. 15707v2 Announce Type: replace-cross Abstract: Real-time conversational assistants for procedural manual tasks often depend on video input, which can be computationally expensive and compromise user privacy.
By Rehana Mahfuz, Yinyi Guo, Erik Visser, Phanidhar Chinchili
The technical report introduces Gander, an end‑to‑end model that integrates omni perception, real‑time interaction, and agentic capabilities into a single framework. Unlike traditional turn‑based systems, Gander continuously processes streaming inputs from video, speech, and text, enabling natural full‑duplex interaction in both everyday conversations and workflow‑oriented scenarios. Its architecture features a Cerebellum‑Brain collaboration—where the Cerebellum handles real‑time interaction and omni conversational tasks while the Brain manages complex reasoning—and a streaming Thinker‑Talker design that flattens inputs and outputs into an ordered token stream for low‑latency, continuous dialogue. Evaluations across conversational ability, omni understanding, interactive capability, and agentic intelligence show that Gander matches state‑of‑the‑art open‑source models in spoken dialogue while maintaining robust performance in noisy, multi‑party, and backchannel environments.
By Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddy Sun, Steve Yves, Zhou Zhao
arXiv:2602. 12873v5 Announce Type: replace-cross Abstract: Generative social robots (GSRs) powered by large language models enable adaptive, conversational tutoring but also introduce risks such as misinformation, overreliance, and privacy violations.
By Stephan Vonschallen, Dominique Oberle, Theresa Schmiedel, Friederike Eyssel
The survey "From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning" reviews over 300 works on efficient multimodal learning (EML), proposing a structured taxonomy that spans model, algorithm, and system layers. It synthesizes how cross‑layer co‑design addresses the Efficiency‑Utility‑Privacy trade‑off and illustrates this through a case study of multimodal large language models. The paper also offers optimization blueprints for various domains, discusses a shift toward self‑regulating intelligence, and outlines open challenges for future EML research.
By Pan Wang, Siwei Song, Hui Ji, Siqi Cao, Heng Yu, Zhijian Liu, Huanrui Yang, Yingyan Celine Lin, Beidi Chen, Mohit Bansal, Xiaoming Liu, Pengfei Zhou, Ming-Hsuan Yang, Tianlong Chen, Jingtong Hu
arXiv:2606. 13722v1 Announce Type: new Abstract: This paper introduces YeasierAgent, an application-building paradigm based on symbiotic agents, narrative worlds, and scene-aware interaction.
By Jory He
arXiv:2607. 22651v1 Announce Type: new Abstract: Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time interactive environments remains a significant challenge.
By Luka Borozan, Domagoj Matijevi\'c
ChatDev 2.0, also called DevAll, is a no-code platform that lets users build, run, and inspect heterogeneous multi‑agent systems (MAS) powered by large language models. It combines a declarative executable graph abstraction with a cycle‑aware execution engine, enabling representation and execution of dynamic, cyclic interactions among diverse agents. The integrated visual interface allows users to author, monitor, and inspect MAS—including human‑in‑the‑loop steps—without writing code, and experiments show it matches state‑of‑the‑art MAS performance across three tasks.
By Yufan Dang, Shu Yao, Bowen Lai, Chenting Xu, Ruijie Shi, Wai-Shing Leung, Huatao Li, Chen Qian, Zhiyuan Liu
arXiv:2606. 11210v1 Announce Type: cross Abstract: Model Construction is a foundational practice in science learning that relies on visualization and interactivity.
By John Kos, Rudra Singh, Ashok Goel