Hugging Face Blog

Introducing smolagents: simple agents that write actions in code.

arXiv AI
Sep 24

Harness as a Language: A Minimalist Agent Framework With Maximal Expressivity

The paper introduces JAZ, a minimalist LLM agent framework that centers on a single primitive called “invoke”, which allows an LLM to write and execute arbitrary code, including recursive calls, while treating all inputs and interaction history as variables in the code environment. JAZ provides built‑in hooks for constraints and monitoring but relies solely on prompting, without external tools, memory systems, or file‑system access. Experiments show that JAZ “invoke” outperforms specialized external harnesses such as Letta (MemGPT) and ACE on long‑horizon recall tasks and continual self‑improvement, achieving higher accuracy at lower cost.

By Zhening Li, Joshua Liu, Mateja Vukelic, Nicole Shen, Supriya Lall, Amitayush Thakur, Alex Zhang, Omar Khattab, Jonathan Light, Armando Solar-Lezama
arXiv AI
Sep 18

MAGS: Multi-agent Auto-formalization Guarantees Safety for Agentic Outputs

The paper introduces MAGS, a multi-agent framework that automatically generates executable programs with formal safety guarantees. MAGS translates LLM-generated code into the verification-aware language Dafny, repairs any safety violations using verifier feedback, and then compiles the verified code back into executable form. Evaluations on 220 diverse examples—including CUDA kernels, terminal scripts, and robotic-arm tasks—show a 100% success rate in producing programs that meet frozen safety specifications, with additional safety and functional tests confirming strong performance across domains.

By Albert Wu, Nicholas Roberts, Tzu-Heng Huang, Haoran Lin, Gil Friedman, Sungjun Cho, Gabriel Orlanski, Frederic Sala
arXiv AI
Aug 26

Quasar: A Programming Language Specialized for LLM Code Actions

Quasar is a new programming language designed to improve large language model (LLM) code actions by separating internal program logic from external tool calls. It allows developers to annotate external calls with effect information and modify internal execution to track these effects, enabling easier implementation of new features. The authors demonstrate Quasar’s utility by adding access control, autoparallelization, and conformal prediction for uncertainty quantification.

By Stephen Mell, Botong Zhang, David Mell, Shuo Li, Ramya Ramalingam, Nathan Yu, Stephan Zdancewic, Osbert Bastani