Vibe-driven model-based engineering
arXiv:2604. 10645v2 Announce Type: replace-cross Abstract: There is a pressing need for better development methods and tools to keep up with the growing demand and increasing complexity of new software systems.
arXiv:2604. 14137v3 Announce Type: replace-cross Abstract: Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness.
arXiv:2604. 10645v2 Announce Type: replace-cross Abstract: There is a pressing need for better development methods and tools to keep up with the growing demand and increasing complexity of new software systems.
arXiv:2609.15387v3 Announce Type: replace-cross Abstract: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automat...
arXiv:2601. 02430v3 Announce Type: replace-cross Abstract: Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and commercial potential.
arXiv:2607. 27816v2 Announce Type: replace-cross Abstract: Role-playing agents (RPAs) have become one of the most important consumer applications of large language models.
arXiv:2609.15387v1 Announce Type: cross Abstract: Human evaluation provides a direct measure of the quality of LLM-generated web applications. However, fitting human judgments through automated evalu...
arXiv:2606. 18293v1 Announce Type: cross Abstract: Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forever.
arXiv:2608.27831v1 Announce Type: cross Abstract: Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues--long, structured, a...
arXiv:2608. 11493v1 Announce Type: new Abstract: Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale.
arXiv:2608. 13566v1 Announce Type: cross Abstract: Post-training papers, model cards, and blog posts often treat scores on a small set of coding benchmarks (e.
The paper presents EvalAgent, an AI assistant that automates agent evaluation by encoding domain expertise into evaluation skills such as procedural instructions, reusable code, and dynamic API retrieval. EvalAgent constructs a trace-based pipeline that outputs metrics, executable code, and reports, and is evaluated using a new meta-evaluation framework and AgentEvalBench. Results show that EvalAgent improves the Eval@1 metric from 17.5% to 65% and receives 79.5% human expert preference, while ablation studies confirm the importance of evaluation skills.
The paper introduces VIBE‑Bench, a new benchmark designed to test personalized large language models (PLLMs) in a regime where user profile cues and query‑specific preferences do not share the same conceptual space, a situation termed profile‑preference conceptual misalignment (PRCM). VIBE‑Bench contains 3,504 personas, 12,239 dialogues, and a manually verified gold test set, and includes two psychology‑grounded tasks that require cross‑concept preference reasoning beyond surface semantic overlap. Experiments show that existing PLLMs largely depend on shallow semantic correlations and struggle to learn robust cross‑concept mappings, highlighting PRCM as a distinct failure mode for personalization models.
arXiv:2510. 07315v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have catalyzed vibe coding, where users leverage LLMs to generate and iteratively refine code through natural language interactions until it passes their vibe check.