arXiv AI By Peiyu Li, Xiaobao Huang, Ting Hua, Nitesh V. Chawla

CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?

Read the original on arXiv AI →

arXiv:2511. 09483v3 Announce Type: replace Abstract: While multimodal large language models can describe visual content, their ability to generate executable procedures remains underexplored.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.