arXiv AI By Itay Itzhak, Eliya Habba, Gabriel Stanovsky, Yonatan Belinkov

From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs

Read the original on arXiv AI →

arXiv:2604. 14137v3 Announce Type: replace-cross Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 2

Vibe-driven model-based engineering

arXiv:2604. 10645v2 Announce Type: replace-cross Abstract: There is a pressing need for better development methods and tools to keep up with the growing demand and increasing complexity of new software systems.

By Jordi Cabot
arXiv AI
Sep 23

WebCraftBench: Evaluating Web Application Generation from a Software Testing Perspective

arXiv:2609.15387v3 Announce Type: replace-cross Abstract: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automat...

By Chenxu Liu, Zilu Zou, Peizhong Gao, Jiawen Tao, Zhexin Zhang, Guang Chen, Haowei Lin, Ying Zhou, Tianyi Bai, Dolly Deng, Suncong Zheng, Maxm Pan