APEX-Accounting
arXiv:2607. 27189v2 Announce Type: cross Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants.
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports.
arXiv:2607. 27189v2 Announce Type: cross Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants.
arXiv:2609.14811v1 Announce Type: new Abstract: We introduce Crypto Accounting Bench (CAB), a benchmark for assessing whether frontier and open-weight language models can reconstruct the complete jou...
The paper evaluates a manager‑worker scaffold that uses a shared filesystem workspace to orchestrate multi‑agent large language model (LLM) coding tasks without training or tuning. Across nine models—including five open‑weight and four closed‑weight systems—the scaffold yields statistically significant accuracy gains for some models (e.g., Qwen3.8‑27B, GPT‑5.6‑Luna, GPT‑5.6‑Terra, Kimi‑K3, Minimax‑M3) while producing null or negative effects for others (e.g., Qwen3.6‑35B). The study shows that the manager can triple token usage but still achieves higher accuracy at a fraction of the cost compared to larger single‑pass models, with key mechanisms identified as context management and problem decomposition.
arXiv:2607. 24889v1 Announce Type: cross Abstract: Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations.
arXiv:2606. 05104v1 Announce Type: new Abstract: Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets.
arXiv:2609.40190v1 Announce Type: cross Abstract: Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is report...
arXiv:2607. 01740v1 Announce Type: new Abstract: Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services work: a model that leads on MMLU-Pro may underperform on document-grounded compliance reasoning, and a coding leader may handle multi-turn customer interactions poorly.
arXiv:2607. 22585v1 Announce Type: new Abstract: Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that issues tools, manages context, and decides when to stop) is often under-specified.
arXiv:2609.39229v1 Announce Type: cross Abstract: Automatic evaluation of faithfulness increasingly relies on a large language model acting as a judge, yet the most reliable judges are proprietary fr...
Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway.
Adding inference structure to a language model lets it search, verify, and revise, but these actions consume the very budget they are supposed to use well. In this paper, we investigate whether there...
FM‑Bench is a new benchmark that tests large language model agents on long‑horizon decision making by having them run a football club for 20 in‑game years. The agent must manage a squad, trade players, negotiate contracts, invest in facilities and youth, set lineups, and respond to a board that can fire it, all using 26 tools and roughly 340–400 decision stops, with a deterministic engine producing a final score without human or LLM judges. The benchmark includes a solo track where each of 15 frontier models competes against a frozen scripted world, and an Arena track where the same models plus a scripted anchor share one 20‑year world, allowing the first head‑to‑head evaluation at this scale. whyItMatters":"FM‑Bench provides a rigorous, large‑scale test of sustained, cumulative decision‑making in language‑model agents, revealing that managerial strategy—not computational scale or vendor—drives performance over long horizons."