arXiv AI By Itay Itzhak, Eliya Habba, Gabriel Stanovsky, Yonatan Belinkov

From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs

Read the original on arXiv AI →

arXiv:2604. 14137v3 Announce Type: replace-cross Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
Jun 2

Vibe-driven model-based engineering

arXiv:2604. 10645v2 Announce Type: replace-cross Abstract: There is a pressing need for better development methods and tools to keep up with the growing demand and increasing complexity of new software systems.

By Jordi Cabot
arXiv AI
2d ago

Don't Claim Benchmark-Oriented Optimization Improves General Coding Capability -- Diverse Evaluation Is Required

arXiv:2608. 13566v1 Announce Type: cross Abstract: Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.

By Egor Shibaev, Vera Kudrevskaia, Timur Galimzyanov, Mikhail Evtikhiev, Ana Terna, Rastislav Rabatin, Timur Kudashev, Timofey Bryksin, Arina Puchkova, Patrik Bartak, Egor Bogomolov, Sergey Titov