TEMPO: Learning Temporal Context for Dynamic Robot Manipulation
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2608. 14379v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently achieved promising performance in robotic manipulation.
StreamPI introduces a streaming multimodal temporal modeling framework that enhances Vision‑Language‑Action models by adding temporal reasoning without extra parameters. It anchors each visual observation and language instruction pair as a temporal unit, using bidirectional attention for cross‑modal fusion and causal attention for autoregressive streaming inference. The method employs random‑interval streaming training to improve robustness and leverages the LLM backbone’s length extrapolation to inherit pretrained weights, achieving superior performance over pi0.5 on real‑robot and simulation tasks.
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to re...
arXiv:2607. 15275v1 Announce Type: cross Abstract: Recent robot foundation models operate with single-step or short-history visuomotor context.
arXiv:2606. 09811v1 Announce Type: cross Abstract: World-action models have emerged as a promising paradigm for robot manipulation, jointly modeling visual scene dynamics and actions to inject physical priors into policy learning.
arXiv:2609.18207v1 Announce Type: cross Abstract: Reinforcement learning fine-tuning on top of large, pretrained Vision-Language-Action (VLA) models offers promise for highly reliable robot deploymen...