arXiv:2607. 02563v1 Announce Type: cross Abstract: Diffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semantic structure remain difficult to interpret.
By Yiran Xiao, George Legrady
arXiv:2608. 08617v1 Announce Type: new Abstract: Group discussion-based teaching is widely used to foster collaborative learning, yet teachers in physical classrooms often struggle to simultaneously monitor multiple groups and quickly diagnose a target group before intervening.
By Yiping Sun, Ziyao Kang, Wei Zeng, Minli Wu, Jiazhi Xia
arXiv:2609.36852v1 Announce Type: new
Abstract: Trajectory prediction is a key component for understanding human behavior patterns in dynamic scenes. Researchers have devoted substantial efforts to m...
By Ziqian Zou, Conghao Wong, Qinmu Peng, Xinge You
arXiv:2604. 03401v4 Announce Type: replace-cross Abstract: Understanding student engagement usually requires time-consuming manual observation or invasive recording that raises privacy concerns.
By Nolan Platt, Sehrish Nizamani, Alp Tural, Elif Tural, Saad Nizamani, Andrew Katz, Yoonje Lee, Nada Basit
arXiv:2601. 14569v2 Announce Type: replace-cross Abstract: Social understanding abilities are crucial for multimodal large language models (MLLMs) to interpret human social interactions.
By Leena Mathur, Bhaavanaa Thumu, Youssouf Kebe, Louis-Philippe Morency
arXiv:2603.18480v2 Announce Type: replace-cross
Abstract: Inferring human engagement from gameplay video is important for game design and player-experience research, yet it remains unclear whether vi...
By Ziyi Wang, Qizan Guo, Rishitosh Singh, Xiyang Hu
arXiv:2607. 04501v1 Announce Type: cross Abstract: The ability to automatically infer analytic intent from user interaction histories could enable interactive AI systems to proactively assist users during exploratory data analysis.
By Steffen Holter, Tobias St\"ahle, Arpit Narechania, Mennatallah El-Assady
arXiv:2606. 06388v1 Announce Type: new Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human collaborators.
By Jiaju Chen, Yuxuan Lu, Jiayi Su, Chaoran Chen, Songlin Xiao, Zheng Zhang, Yun Wang, Yunyao Li, Jian Zhao, Tongshuang Wu, Toby Jia-Jun Li, Dakuo Wang, Bingsheng Yao
arXiv:2608. 20157v1 Announce Type: new Abstract: Egocentric action understanding is often addressed using large video models pretrained on extensive exocentric datasets.
By Marko Haralovi\'c, Akash Ramakrishnan, Estefania Talavera Martinez
arXiv:2606. 01897v1 Announce Type: new Abstract: Traditional Video Quality Assessment (VQA) focuses narrowly on aesthetic fidelity, overlooking the complex social dynamics that define quality in User-Generated Content (UGC).
By Tianjiao Li, Kai Zhao, Xiang Li, Yang Liu, Huyang Sun
arXiv:2606. 24635v1 Announce Type: cross Abstract: Traditional visual data storytelling relies on binary graphics that depict two simplified groups in conflict.
By Lisa Schirch, Beth Goldberg
arXiv:2608. 10195v1 Announce Type: cross Abstract: Human vision organizes what it sees into wholes: same-colored points group into series, similar marks cohere into categories, and shapes complete into recognizable objects.
By Sudhanva Manjunath Athreya, Sai Phani Kumar Malladi