Shot‑scraper 1.12 adds WebP support, allowing users to capture web page screenshots in WebP format with an optional quality setting. The new --quality flag controls compression, while omitting it produces lossless images. WebP screenshots are reported to be significantly smaller than JPEG or PNG equivalents.
Our latest image generation model is now available in the API via ‘gpt-image-1’—enabling developers and businesses to build professional-grade, customizable visuals directly into their own tools and platforms.
arXiv:2606. 01022v1 Announce Type: cross Abstract: Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value for domains such as marketing, advertising, and E-commerce.
By Zhihong Liu, Siqi Kou, Zheng Li, Ye Ma, Quan Chen, Peng Jiang, Kai Yu, Zhijie Deng
arXiv:2606. 28593v1 Announce Type: cross Abstract: While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code for webpages, charts, or SVGs, it remains unclear whether they can recover temporal dynamics when motion is present.
By Anya Ji, Abhijith Varma Mudunuri, David M. Chan, Alane Suhr
The paper introduces a platform that renders web pages using FitLayout and captures their visual and structural properties in an RDF-based representation. The system offers a REST API for pipeline control, SPARQL queries for data retrieval, and a Python client to integrate with machine learning workflows. It demonstrates how rendered pages can be converted into graph representations to train graph neural networks for key content recognition, highlighting reproducibility and dataset sharing.
By Radek Burget, Radek Hranick\'y
arXiv:2606. 28344v1 Announce Type: cross Abstract: Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipelines that linearize HTML and discard layout, visual structure, and formatting.
By Yichuan Wang, Zhifei Li, Zirui Wang, Paul Teiletche, Lesheng Jin, Matei Zaharia, Joseph E. Gonzalez, Sewon Min
Image inputs and structured outputs with Gemma 4 and Ollama The post Building Multimodal Workflows with a Local LLM appeared first on Towards Data Science .
By Shuai Guo
Learn how to apply coding agents to verify work in your browser. The post How to Use Claude Code in Your Browser appeared first on Towards Data Science .
By Eivind Kjosbakken
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts that impose substantial demands on users and offer limited expressivity for page layout and cross-page visual coherence.
arXiv:2607. 06306v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated growing competence in web page generation.
By Grace Man Chen, Litao Guo, Yifan Wu, Yiyu Chen, Yenchi Tseng, Sicheng Liu, Yuyu Luo, Ying-Cong Chen