Introducing smolagents: simple agents that write actions in code.
Related stories
Tiny Agents: an MCP-powered agent in 50 lines of code
Tiny Agents in Python: a MCP-powered agent in ~70 lines of code
Agents All the Way Down; A Methodology for Building Custom AI Agents from Substrate to Production
arXiv:2606. 11869v1 Announce Type: cross Abstract: Custom AI agents areagents that live inside their own application, talk to their own data and tools, enforce their own security boundaries, and carry their own brand and audit trail.
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
arXiv:2604. 14228v2 Announce Type: replace-cross Abstract: Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user.
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
arXiv:2604. 11556v2 Announce Type: replace-cross Abstract: LLM-assisted software development has become increasingly prevalent, and can generate large-scale systems, such as compilers.
Inside VAKRA: Reasoning, Tool Use, and Failure Modes of Agents
Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity
The paper introduces JAZ, a minimalist LLM agent framework that centers on a single primitive called “invoke”, which allows an LLM to write and execute arbitrary code, including recursive calls, while treating all inputs and interaction history as variables in the code environment. JAZ provides built‑in hooks for constraints and monitoring but relies solely on prompting, without external tools, memory systems, or file‑system access. Experiments show that JAZ “invoke” outperforms specialized external harnesses such as Letta (MemGPT) and ACE on long‑horizon recall tasks and continual self‑improvement, achieving higher accuracy at lower cost.
MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs
The paper introduces MAGS, a multi-agent framework that automatically generates executable programs with formal safety guarantees. MAGS translates LLM-generated code into the verification-aware language Dafny, repairs any safety violations using verifier feedback, and then compiles the verified code back into executable form. Evaluations on 220 diverse examples—including CUDA kernels, terminal scripts, and robotic-arm tasks—show a 100% success rate in producing programs that meet frozen safety specifications, with additional safety and functional tests confirming strong performance across domains.
Computer-Using Agent
AgentArmor: A Framework, Evaluation, \& Mitigation of Coding Agent Failures
arXiv:2606. 19380v1 Announce Type: cross Abstract: Software engineering and deployment are increasingly being delegated to AI coding agents.
Quasar: A Programming Language Specialized for LLM Code Actions
Quasar is a new programming language designed to improve large language model (LLM) code actions by separating internal program logic from external tool calls. It allows developers to annotate external calls with effect information and modify internal execution to track these effects, enabling easier implementation of new features. The authors demonstrate Quasar’s utility by adding access control, autoparallelization, and conformal prediction for uncertainty quantification.