arXiv AI

Toward Controlling Biology with Language:Offline Learning of Prompt-Conditioned Interventions for Cells, Organoids, and Biobots

The paper demonstrates that a natural-language interface can be trained offline to control a xenobot—a synthetic multicellular construct—by using a vision‑language model to judge whether archived intervention outcomes match language descriptions. By treating existing intervention–outcome data as a fixed dataset, the authors train a language‑to‑intervention mapping without new experiments, achieving 80% accuracy on held‑out data compared to a 66.7% baseline. This approach shows that language‑driven control of living systems can be learned purely from archival data and automated visual assessment.

arXiv AI
Jun 12

LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories

arXiv:2606. 13578v1 Announce Type: cross Abstract: Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely outside their reach.

By Baochang Ren, Xinjie Liu, Xi Chen, Yanshuo Liu, Chenxi Li, Daqi Gao, Zeqin Su, Jintao Xing, Zirui Xue, Rui Li, Xiangyu Zhao, Shuofei Qiao, Minting Pan, Wangmeng Zuo, Lei Bai, Dongzhan Zhou, Ningyu Zhang, Huajun Chen
arXiv AI
Sep 17

WetRobo: A Reproducible Robot Kit for Coding Agents in Biological Laboratories

WetRobo is a reproducible robot kit designed to enable wet‑lab researchers to delegate tasks to coding agents without teleoperation or neural‑network training. The kit includes a robot arm, essential lab equipment, pre‑recorded teleoperation demos, and a skill file, allowing a coding agent to observe the lab, write, and execute programs using external tools. Experiments with OpenAI Codex on tasks such as lifting a Petri dish lid, removing a bottle cap, and opening an incubator door demonstrated successful performance in two different laboratories, outperforming a fine‑tuned vision‑language‑action policy that failed to transfer.

By Yuna Oikawa, Kei Endo, Takanori Uzawa, Yunzhe Zhang, Manan Anjaria, Lerrel Pinto, Sherry Yang, Koji Tsuda
arXiv AI
Aug 19

Teach and Grow: An Agent-Centered Architecture for General Robot Learning

Teach-and-Grow Learning (TGL) is an agent-centered architecture that transforms a few successful demonstrations into reusable Skill Blocks, enabling a robot to compose, execute, and revise behaviors in new scenes without task-specific policy retraining. The system maintains a Skill Library and structured Experience Memory to capture successes, failures, and repairs, allowing persistent reuse and agent-directed adaptation. Evaluation on the LIBERO benchmark shows state-of-the-art performance, and the authors propose a scaling-law hypothesis suggesting that accumulated reusable experience reduces future-task error and teaching demand following a power-law trend.

By Chang Nie, Zhe Liu, Hesheng Wang
arXiv Machine Learning
Jun 11

Learning What to Say to Your VLA: Mostly Harmless Vision Language Action Model Steering

arXiv:2606. 12299v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone.

By Hyun Joe Jeong, Gokul Swamy, Andrea Bajcsy
Hugging Face Trending Papers
Sep 28

Shockingly Simple Self-retrospection Improves Agentic Models Without RL

The paper explores Retrospection-Only Fine-Tuning (ROFT), a method where a language-model agent improves its behavior by generating and training on explanations of its own experiences, without external teachers or reward signals. In software‑engineering tasks with Qwen3.5‑4B, ROFT achieves comparable or better solve rates than GRPO while requiring fewer updates and training time, and can learn from failures alone. Behavioral analysis shows ROFT indirectly assigns credit to actions and can produce shorter, more direct solutions when prompted to focus on direct solutions.

arXiv AI
Sep 23

ORDER: A Fictitious-World Benchmark for Domain-Adaptive Embodied AI

The paper introduces ORDER, a fictitious-world benchmark designed to evaluate domain-adaptive embodied AI. ORDER consists of a synthetic 342,069-token corpus defining a self-consistent physics, a 500-question knowledge test (ORDER‑BENCH), and a compositional spatial task (ORDER‑SPATIAL) that requires ordering objects for safe manipulation. The benchmark demonstrates that models like GPT‑4.1 perform poorly without adaptation, while small models improve significantly after continual pre‑training, and that performance on ORDER‑SPATIAL better predicts real plan quality than knowledge-test accuracy.

By Sai Krishna Reddy Sathi, Anuj Tiwari
arXiv AI
Sep 25

Do World Models Make Better Robots? A Survey of Evaluation Benchmarks for Predictive Embodied Intelligence

The paper surveys 160 benchmarks from 2017‑2026 that evaluate predictive embodied intelligence, categorising them into policy suites, embodied agents, world‑model evaluation, and prediction‑to‑action bridges. It finds that most benchmarks are model‑agnostic, rarely compare Vision‑Language‑Action policies to world models, and seldom turn predictions into executed actions. The authors argue that the lack of benchmarks designed to directly test the closed‑loop advantage of world models prevents the field from answering whether such models truly improve robotic performance.

By Gaytri Jena, Kapil Wanaskar, Vinija Jain, Aman Chadha, Vasu Sharma, Amitava Das