WebWorld: The Browser as a World Model for Self-Improving Web Code
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2608. 06474v1 Announce Type: new Abstract: Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remaining functional gap.
arXiv:2610.03036v1 Announce Type: cross Abstract: We present WebFovea, a vision-based web agent that placed 2nd in the WebRetriever Challenge 2026 with a final score of 57.0 out of 100. The challenge...
arXiv:2606. 03220v1 Announce Type: cross Abstract: Existing benchmarks for MLLM-generated web artifacts assess interaction through local evidence and miss the requirement-induced states and transitions that determine whether a page works.
Wuying-Browser-Agent is a unified framework designed to improve long-horizon browser agents by aligning execution, supervision, optimization, and evaluation. It introduces a structured browser harness, reflection and UI-specialized Curriculum SFT (RUIC‑SFT) for recovery and complex UI interactions, and Divergence‑Aware Online GRPO (DAO‑GRPO) for better credit assignment. The framework is evaluated on BrowserBench—a bilingual real‑web benchmark of 350 tasks—and achieves state‑of‑the‑art results on multiple browser‑use benchmarks, while also transferring well to other agentic tasks.
arXiv:2606. 17645v1 Announce Type: new Abstract: Large language model (LLM) web agents are usually deployed as tool callers: each turn, the model reads a fresh page observation and emits one structured tool action.
WebUIProof is a new execution‑oriented benchmark for WebUI code generation that supplies structured specifications and dense, executable interaction tests across general WebUIs and 3D interactive simulations. It employs a UI‑agent harness that runs tests in a headless browser using a plan–act–observe loop to locate DOM elements, perform actions, observe changes, and verify assertions. Evaluations on eight commercial LLMs reveal frequent interaction‑based failures, especially on 3D interfaces, and demonstrate that training compact models with RL rewards from these tests improves functional completion and reduces build failures.