arXiv AI

PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums

arXiv:2606. 16175v1 Announce Type: new Abstract: Longitudinal personal albums are weak-schema multimodal databases: noisy perceptual records whose key facts require joins across faces, text, timestamps, locations, and repeated events.

arXiv Computer Vision
Sep 7

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

ICM-Bench is a new benchmark for evaluating identity-centric reasoning in multimodal agents with long-term memory. It consists of 839 synthetic video clips totaling 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. The benchmark isolates the ability to maintain recurring person identities and reason over their cross-time relations, and compares various baseline systems, showing that while Gemini 3.1 Pro performs well overall, its accuracy drops on questions requiring long-term identity profiles.

By Shidu Ren, Yunze Liu, Xing Liu, Chi-Hao Wu, Enmin Zhou, Junxiao Shen
arXiv AI
Sep 7

TruthInsightBench: An Evidence-Grounded Benchmark for Automated Evaluation of Open-Ended Scientific Discovery Agents

TruthInsightBench is a new benchmark designed to evaluate automated scientific discovery agents by presenting them with 40 blind tasks drawn from peer‑reviewed studies across ten domains. Each task provides only a neutral objective and frozen data, withholding source conclusions, expected values, and analysis paths, forcing agents to determine which claim the data support. A fixed LLM‑based judge scores agents on evidentiary maturity across six dimensions, using 29 artifact‑grounded items, enabling fully automated, repeatable evaluation without human grading.

By Zhibo Yang, Chen Zhang, Yuewei Zhang, Hao Wang
arXiv Computation and Language
Sep 14

GraphProfiler: Source-Linked Sensitive Attribute Inference via Personal Knowledge Graphs

GraphProfiler is an auditable LLM-based profiler that infers sensitive attributes such as age, income, and occupation from user-generated content by aggregating indirect cues across many posts. It represents each user's post history as a source-linked personal knowledge graph, allowing predictions to be traced back to specific posts, concepts, and relationships. The system achieves an 86.7% success rate on the SynthPAI benchmark and 84.6% on PANDORA, citing supporting evidence for over 98% of predictions and demonstrating that cited posts significantly contribute to attack success.

By Ahmed Sohair Khan, Estrid He, Chenglong Ma, Monica Wachowicz, Elham Naghizade
arXiv AI
Aug 26

Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes

The paper audits a 366‑day autobiographical book generated by a large language model (LLM) against an independent verification corpus. Using a four‑level rubric, 354 of the 366 days (96.7%) failed verification, with only 12 days containing corroborated scenes and 19 days containing actively contradicted claims. Regenerating the same days with current models yielded 100% verification failure, while grounding the generation in the subject’s own corpus improved the rate to 83.3% but still left substantial residual failure.

By Heather Renze
arXiv AI
Aug 24

PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval

PhotoBench is a new benchmark built from authentic personal photo albums that moves beyond simple visual matching to focus on personalized, intent-driven retrieval. It incorporates a multi-source profiling framework that combines visual semantics, spatial‑temporal metadata, social identity, and temporal events to generate complex queries reflecting users’ life trajectories. Evaluation on PhotoBench reveals two key limitations: a modality gap where unified embedding models fail on non‑visual constraints, and a source fusion paradox where agentic systems struggle with tool orchestration.

By Tianyi Xu, Rong Shan, Junjie Wu, Jiadeng Huang, Teng Wang, Jiachen Zhu, Wenteng Chen, Minxin Tu, Quantao Dou, Zhaoxiang Wang, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin
arXiv AI
4d ago

A Benchmark & Dataset for Detecting AI-Manipulated Visual Evidence in the Court System

arXiv:2609.37783v1 Announce Type: cross Abstract: Photographic evidence is becoming increasingly vulnerable to forms of alteration and fabrication that existing legal and technical workflows are not...

By Kelly McConvey, Sajad Ebrahimi, Nima Jamali, Jalehsadat Mahdavimoghaddam, Matina Mahdizadeh Sani, Maksym Taranukhin, Wentao Zhang, Jacquelyn Burkell, Yuntian Deng, Karen Eltis, Maura R. Grossman, Vered Shwartz, Ebrahim Bagheri