When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2606. 03134v1 Announce Type: cross Abstract: Imitation-learning policies for robot manipulation inherit the quality of the success labels attached to their training episodes, and those labels are usually produced by the robot's own success check.
The paper introduces a new benchmark for vision‑language models that tests their ability to decide whether to answer a physics question immediately or to request additional experimental evidence. Each problem presents one measurement image and four possible physical worlds defined by two masses and two values of another property; the model must either stop and answer or choose the cheapest experiment that resolves the question. Across six open models and 144 parameter sets, the models almost always repeat the same action even when the optimal choice changes, and only a single model gets both decisions correct on 5.9% of cases.
arXiv:2603.12717v2 Announce Type: replace-cross Abstract: Vision-language-action policies map camera images and natural-language instructions to a robot's motor actions. Some of these policies are de...
arXiv:2609.08123v1 Announce Type: cross Abstract: A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it still cannot do, and then ask for exactly t...
The paper investigates whether confidence scores from a black-box decision model, Jev, truly reflect missing knowledge. Using over 15 public datasets and 6 synthetic task families, the authors find that while Jev’s confidence is calibrated on familiar closed-choice tasks, it fails to indicate when the model lacks relevant information—assigning high confidence to salient options even without answer-relevant data and overestimating accuracy on news beyond its knowledge boundary. Targeted yes/no questions about whether an outcome is settled or whether evidence suffices provide sharper indicators of knowledge gaps, but only when surface cues are controlled.
arXiv:2609.39971v1 Announce Type: cross Abstract: Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action,...