AnchorGUI: Asymmetric Memory for Dual-Scale Learning in GUI Navigation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2604. 17473v3 Announce Type: replace-cross Abstract: Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions.
AVERT-VLN introduces a closed‑loop framework for vision‑and‑language navigation that incorporates an abstention‑aware Monitor to detect instruction‑execution inconsistencies. The Monitor is trained on a new LOSTNAV dataset of 20K counterfactual risk trajectories and fine‑tuned on 40K normal trajectories to recognize semantic deviations. During deployment, the Monitor can suspend autonomous navigation and request human guidance, while offline preference learning uses deviation‑associated failures to improve the policy, achieving 76.2% and 66.3% success on R2R‑CE and RxR‑CE unseen splits.
arXiv:2607. 10383v1 Announce Type: cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks.
arXiv:2606. 20244v1 Announce Type: cross Abstract: Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and easy to overlook, leading to failures in evidence readout even when high-level reasoning is intact.
arXiv:2606. 03598v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved remarkable success in language-conditioned robotic manipulation.
arXiv:2606.20092v3 Announce Type: replace Abstract: Memory remains a critical bottleneck for long-horizon robotic manipulation, as standard Vision-Language-Action (VLA) policies often fail when task-...