arXiv:2606.26403v2 Announce Type: replace
Abstract: Foundation-model research increasingly needs data about people: user state, personal histories, relationships, contact-like fields, documents, and...
By Sriram Selvam, Anneswa Ghosh
ICM-Bench is a new benchmark for evaluating identity-centric reasoning in multimodal agents with long-term memory. It consists of 839 synthetic video clips totaling 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. The benchmark isolates the ability to maintain recurring person identities and reason over their cross-time relations, and compares various baseline systems, showing that while Gemini 3.1 Pro performs well overall, its accuracy drops on questions requiring long-term identity profiles.
By Shidu Ren, Yunze Liu, Xing Liu, Chi-Hao Wu, Enmin Zhou, Junxiao Shen
arXiv:2602.19001v2 Announce Type: replace
Abstract: As large language models increasingly power personal assistants, users expect them to reason over multimodal life histories, from recognizing peopl...
By Xia Hu, Honglei Zhuang, Brian Potetz, Alireza Fathi, Bo Hu, Babak Samari, Howard Zhou
arXiv:2605. 06142v2 Announce Type: replace-cross Abstract: When people recount personal memories, they often refer to people, places, and events indirectly, relying on con-textual cues rather than explicit names.
By Yehudit Aperstein, Eden Moran, Alexander Apartsin
TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.
By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
GraphProfiler is an auditable LLM-based profiler that infers sensitive attributes such as age, income, and occupation from user-generated content by aggregating indirect cues across many posts. It represents each user's post history as a source-linked personal knowledge graph, allowing predictions to be traced back to specific posts, concepts, and relationships. The system achieves an 86.7% success rate on the SynthPAI benchmark and 84.6% on PANDORA, citing supporting evidence for over 98% of predictions and demonstrating that cited posts significantly contribute to attack success.
By Ahmed Sohair Khan, Estrid He, Chenglong Ma, Monica Wachowicz, Elham Naghizade