arXiv AI

Signals of Provenance: Practices & Challenges of Navigating Indicators in AI-Generated Media for Sighted and Blind Individuals

arXiv:2505. 16057v2 Announce Type: replace-cross Abstract: AI-Generated (AIG) content has become increasingly widespread by recent advances in generative models and the easy-to-use tools that have significantly lowered the technical barriers for producing highly realistic audio, images, and videos through simple natural language prompts.

arXiv AI
Aug 11

Towards Expert-level Medical AI for Real-time Video Consultations

arXiv:2608. 09861v1 Announce Type: new Abstract: Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues.

By Mahvish Nagda, Jihyeon Lee, Matthew Thompson, Chunjong Park, Tim Strother, Valentin Li\'evin, Roma Ruparel, Akshay Goel, Teya Bergamaschi, Suhana Bedi, Meet Shah, Pavel Dubov, Liviu Panait, Toshiyuki Fukuzawa, Sam Schmidgall, Craig Schiff, Joseph Xu, Aliya Rysbek, Yana Lunts, Jan Freyberg, Rebecca Hemengway, Sunny Virmani, David Racz, Carey Radebaugh, Jo\"elle Barral, Kavi Goel, Dale R. Webster, Katherine Chou, Avinatan Hassidim, Yossi Matias, James Manyika, Gregory Wayne, Tao Tu, Yun Liu, Ethan Goh, Christina Chen, Ryutaro Tanno, Po-Hsuan Cameron Chen, Mike Schaekermann, Anil Palepu
arXiv Machine Learning
Sep 10

I Don't Miss You, but I Do: Self-Explanation Faithfulness of Modality Missingness in Vision-Language Models

The paper introduces an interventional protocol to assess how vision‑language models (VLMs) explain the impact of missing modalities on their predictions. By comparing the models’ self‑explanations with actual changes observed after restoring missing inputs, the study finds that VLMs routinely overstate the sufficiency of available evidence and underestimate the effect of adding back missing modalities. Across eight open‑weight VLMs and four tasks, the discrepancy between predicted and realized changes is substantial, revealing systematic mischaracterization of modality dependence.

By Aydin Javadov, Daniel Schoess, Florian von Wangenheim
arXiv Computer Vision
Aug 26

From Seeing to Acting: Smart Glasses as First-Person Intelligence Platforms

The paper surveys the evolution of smart glasses from simple capture devices to first‑person intelligence platforms that integrate human perception, context, and action. It introduces a unified framework that formalizes data flow, hardware capabilities, and seven foundational capabilities, and presents an L0‑L5 hierarchy for capture to embodied action. The study also maps nine application scenes, proposes a nine‑dimensional deployment framework, and outlines an evidence ladder for evaluation and trustworthiness.

By Jiangning Zhang, Haojun Chen, Yong Liu
arXiv AI
Sep 12

Exploring Multimodal Prompt for Visualization Authoring with Large Language Models

The paper investigates how large language models (LLMs) interpret ambiguous or incomplete text prompts for visualization authoring and introduces visual prompts as a complementary modality to improve precision. An empirical study informs the design of VisPilot, a system that allows users to create visualizations using text, sketches, and direct manipulation. A controlled user study and expert evaluation show that multimodal prompts help users convey spatial constraints, local references, and design preferences while maintaining task efficiency comparable to text-only prompting.

By Zhen Wen, Luoxuan Weng, Yinghao Tang, Runjin Zhang, Yuxin Liu, Bo Pan, Minfeng Zhu, Wei Chen