arXiv AI

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

arXiv:2607. 24889v1 Announce Type: cross Abstract: Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations.

Hugging Face Trending Papers
Aug 12

FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents

AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand.

arXiv Computation and Language
Sep 11

Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models

The paper argues that verbalized confidence—once viewed as overconfident and coarse—has become the preferred soft‑scoring method for LLM‑as‑a‑Judge on top‑tier proprietary models released after 2025. Experiments on SummEval, AggreFact, and HelpSteer2 across up to 18 LLMs show that log‑probabilities are no longer the best signal, and that adding an overconfidence advisory and self‑debate further improves calibration and robustness. The authors note that these enhancements incur little accuracy loss on post‑2025 models but do affect pre‑2025 ones, highlighting a compatibility shift in how confidence should be measured.

By Yu-Chung Hsiao
arXiv Machine Learning
Sep 24

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.

By Md Rezwanul Islam, Wael Mohammed
arXiv AI
3d ago

Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking

The paper evaluates an agentic harness that combines a 27B language model, financial calculations, narrative templates, and validation checks to produce corporate and investment banking presentation decks. Using a panel of five judges, the system consistently scores higher than a baseline model that generates directly from a short prompt, with scores ranging from 20.4 to 33.6 out of 95. However, judge variability and changes in grading criteria make it challenging to discern small improvements in deck quality.

By Ludovic Gibert, Matis Despujols, Andre-Louis Rochet
arXiv Computation and Language
Aug 27

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.

By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang