arXiv AI

When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence

arXiv Machine Learning
Sep 11

New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models

The paper introduces a new benchmark for vision‑language models that tests their ability to decide whether to answer a physics question immediately or to request additional experimental evidence. Each problem presents one measurement image and four possible physical worlds defined by two masses and two values of another property; the model must either stop and answer or choose the cheapest experiment that resolves the question. Across six open models and 144 parameter sets, the models almost always repeat the same action even when the optimal choice changes, and only a single model gets both decisions correct on 5.9% of cases.

By Sourajit Saha, Shubhashis Roy Dipta, Nobin Sarwar, Shaswati Saha, Yuxuan Jiang, Siyuan Li, Qiheng Wang
arXiv AI
2d ago

Beyond Answer Confidence: A Controlled Audit of Self-Knowledge in a Black-Box Decision Model

The paper investigates whether confidence scores from a black-box decision model, Jev, truly reflect missing knowledge. Using over 15 public datasets and 6 synthetic task families, the authors find that while Jev’s confidence is calibrated on familiar closed-choice tasks, it fails to indicate when the model lacks relevant information—assigning high confidence to salient options even without answer-relevant data and overestimating accuracy on news beyond its knowledge boundary. Targeted yes/no questions about whether an outcome is settled or whether evidence suffices provide sharper indicators of knowledge gaps, but only when surface cues are controlled.

By Sharath M Shankaranarayana, Davor Runje, Jan Jannink