arXiv:2608. 13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks.
By Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi
arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.
By Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language mode...
arXiv:2607. 11523v1 Announce Type: cross Abstract: When should an intelligent assistant speak up without being asked?
By Gong Sitong, Tianyu Yan, Caixin Kang, Bo Zheng, Xiang Ruan, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang
NextMe-800 is an approximately 800‑hour first‑person video dataset collected from a single volunteer over 126 days, featuring 1 Hz images, gaze, and audio. The data are captioned at five hierarchical abstraction levels—from atomic actions to major activities—enabling personalized action anticipation as an open‑vocabulary K‑step sequence prediction task. The authors also introduce NextAct, a 1,500‑point benchmark that combines NextMe‑800 with the multi‑person EgoLife dataset, and evaluate models using an embedding‑based soft edit distance to assess how well personal behavior can be anticipated across abstraction levels and prediction horizons.
By Zhaoxu Meng, Yiming Sun, Mingyuan Gao, Jiachang Zhang, Zhuhan Dai, Yipeng Du, Zheng Lian, Jian-Qiao Zhu
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive.
arXiv:2608. 10765v1 Announce Type: new Abstract: Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a semantic hierarchy.
By Farnaz Soleimani (LISSI), Abdelghani Chibani (LISSI), Yacine Amirat (LISSI), Ghazaleh Khodabandelou (LISSI)
The paper investigates when vision‑language models (VLMs) can independently analyze human‑centered video and when human oversight is still needed. By reviewing 1,702 CHI 2026 papers, the authors develop a five‑dimensional taxonomy of video annotation tasks and build a benchmark of 15 representative tasks. Experiments show that VLMs alone achieve near‑human accuracy (HNS = 97.0), while human verification of VLM outputs yields the highest accuracy (HNS = 121.5) and significantly reduces annotation time and cost.
By Xiyuan Shen, Jiuyang Lyu, Seokhyun Hwang, Huanfen Yao, Shwetak Patel, Zhihan Zhang, Jacob O. Wobbrock
arXiv:2606.03371v4 Announce Type: replace
Abstract: Reliable proactive agents must choose an action and judge whether current evidence is sufficient to act. We study retail service from sparse third-...
By Honghui Zhang, Anna Min, Chenmeinian Guo, Yujia Zhang, Yichen Yu, Zezhou Zhang, Guanyu Liu, Yongming Qin, Chongguo Song, Mengyue Yang, Lei Yu, Tianyu Shi
arXiv:2604. 00767v2 Announce Type: replace Abstract: Wearable human activity recognition (HAR) has made steady progress, yet much of this progress remains grounded in fixed-window, closed-set classification benchmarks.
By Lala Shakti Swarup Ray, Mengxi Liu, Alcina Pinto, Deepika Gurung, Daniel Geissler, Paul Lukowoicz, Bo Zhou
arXiv:2603. 09731v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether they can reliably reason about the long-term physical consequences of actions from an egocentric viewpoint.
By Chengjun Yu, Xuhan Zhu, Chaoqun Du, Pengfei Yu, Wei Zhai, Yang Cao, Zheng-Jun Zha
arXiv:2608.27562v1 Announce Type: new
Abstract: Translating continuous, noisy egocentric video streams into discrete, temporally ordered action steps is fraught with visual challenges. Heavy ego-moti...
By Anubhav Gupta, Archit Kambhamettu, Vatsal Agarwal, Pulkit Kumar, Abhinav Shrivastava