Mistral AI

Upgrading agentic coding capabilities with the new Devstral models

arXiv AI
Jun 17

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering

arXiv:2606. 17799v1 Announce Type: cross Abstract: Coding agents have become a major mode of software engineering, but the benchmarks we use to compare them were designed in a pre-agent era: they collapse model, harness, and environment into a single end-to-end score, typically computed against one reference solution, with no component-level signal for iteration.

By Maria I. Gorinova, Macey Baker, Amy Heineike, Maksim Shaposhnikov, Rob Willoughby, Dru Knox
arXiv AI
1d ago

Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025

The paper titled "Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025" examines how experienced developers employ AI agents in software development. Through field observations and surveys, it finds that developers value agents for productivity but maintain control over design and implementation to ensure quality. They use agents as collaborative tools rather than full delegation, selecting tasks based on suitability and leveraging their expertise to guide agent behavior.

By Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, Brian Hempel