When Does Test-Time Physical Diagnosis Pay? A Frozen Policy Buys Evidence It Never Reads
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.21942v1 Announce Type: cross Abstract: A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt...
CoreSense is a robot‑system integration architecture that traces episodic evidence and uses a conflict‑aware belief gate to decide whether to proceed, re‑observe, abstain, or escalates. The gate evaluates scope, provenance, time, contradiction, and support before making a recommendation. Evaluation on public robot datasets, simulations, and a live cloud deployment shows that belief gating can eliminate protocol‑defined unsafe proceeds while maintaining auditability.
arXiv:2606. 03134v1 Announce Type: cross Abstract: Imitation-learning policies for robot manipulation inherit the quality of the success labels attached to their training episodes, and those labels are usually produced by the robot's own success check.
arXiv:2607. 12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes.
arXiv:2609.09250v1 Announce Type: cross Abstract: A verifier for robot policies reads a candidate behavior and returns a score for how well it did, used both to evaluate vision-language-action polici...
The study investigates how language models equipped with tools can still produce unsupported final claims, even when a single tool call could resolve the uncertainty. It defines two metrics—occurrence (how often unsupported claims arise) and conditional repair (how often they are fixed when evidence is provided). Experiments on Qwen3-32B and Gemma 4 show that providing the missing evidence consistently repairs all unsupported claims in the Qwen3-32B setup, while the Gemma 4 model never produced unsupported claims under the tested conditions.