TransMeme introduces a multi‑agent framework for cross‑cultural meme transcreation, addressing the unique challenges of preserving intent, adapting cultural meaning, and maintaining multimodal consistency. The system coordinates specialized agents for cultural adaptation, text rewriting, revision, and visual adjustment, and is evaluated on Chinese‑English meme pairs. Human and LLM‑based evaluations show that TransMeme outperforms baselines, achieving a 33.1% average improvement in human scores and a 60% Top‑1 ranking rate in LLM judgments.
By Jingyi Zheng, Yule Liu, Zifan Peng, Tianyi Hu, Yuemeng Zhao, Xinhu Zheng, Xinlei He
Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate the content of the verbal channel. However, huma...
arXiv:2608.22731v1 Announce Type: new
Abstract: Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate...
By Parisa Ghanad Torshizi, Stacy Marsella
arXiv:2510. 08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.
By Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan
Internet memes are a pervasive form of multimodal online communication; however, such communication often involves users from diverse linguistic and cultural backgrounds. Therefore, adapting memes acr...
arXiv:2609.00802v1 Announce Type: new
Abstract: Multi-party interaction is a central setting for human communication and a necessary target for human-agent interaction systems that must participate i...
By Taiga Mori, Koji Inoue, Mikey Elmers, Divesh Lala, Tatsuya Kawahara
arXiv:2609.37853v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly deployed as social agents, yet credible human-like interaction requires more than fluent responses or per...
By Wentao Liu, Xi Chen, Siyu Song, Biao Yuan, Yu Zhang, Zhou Zhuotong, Jingying Zhou, Guohao Feng, Shasha Hu, Tianfu Wang, Shangshang Yang, Haoyang Liu, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang
HUG‑VIS is a unified multimodal benchmark for human‑centered visual intelligence, comprising 8,400 half‑body videos of 30 professional actors performing 280 emotion‑action prompts in Mandarin. The dataset provides synchronized video, audio, text, and alpha mattes for four tasks—human emotion recognition, video generation, voice cloning, and video matting—allowing evaluation of both open‑ and closed‑source models under a zero‑shot protocol. Results reveal that linguistic cues dominate emotion recognition, visual affect is weakest, and that automatic metrics and human judgments diverge in generation and cloning tasks, while motion‑related boundary fidelity remains a key challenge for matting.
By Fei Ma, Zebang Cheng, Minghui Li, Hongbo Xu, Yuyong Tan, Yihua Shao, Hanling Wang, Zhou Liu, Yuqing Gao, Dong Wang, Long Ma, Laizhong Cui, Nicu Sebe, Qi Tian
arXiv:2606. 28769v1 Announce Type: new Abstract: Emotional body motion expressions are an essential element of non-verbal communication.
By Huakun Liu, Miao Cheng, Xin Wei, Felix Dollack, Victor Schneider, Hideaki Uchiyama, Chia-huei Tseng, Yoshifumi Kitamura, Monica Perusquia-Hernandez
User interface (UI) and user experience (UX) evaluation is central to product development, yet reliable feedback still relies on recruiting human participants or running online A/B tests, making early-stage iteration slow and costly. In light of this, recent work has explored Multimodal Large Language Models as proxy evaluators.
WorldBench is a new multilingual benchmark that tests large language model agents on culturally grounded everyday workflows, offering 1,600 tasks in seven languages and eight cultures. The benchmark evaluates agents through structured sandbox actions and introduces Constrained Task Success (CTS), a metric that assesses task completion, minimal modification, and other complementary aspects via deterministic and LLM-as-a-Judge evaluations. Experiments show that even leading models achieve only 49.2% CTS, revealing significant gaps in correctness and state preservation across languages and cultures.
By Leonardo Ranaldi, Sherrie Shen, Jushi Kai, Alexandra Birch
arXiv:2610.08513v1 Announce Type: cross
Abstract: LLMs are increasingly deployed as autonomous agents in social environments, making it critical to study their ability to faithfully simulate human in...
By Dennis Fucci, Andrea Bacciu, Dong Liu, Weronika {\L}ajewska, Saab Mansour