arXiv AI By Ludovic Gibert, Matis Despujols, Andre-Louis Rochet

Better Deck or Different Judge? Evaluating Agentic Harness Gains in Corporate and Investment Banking

Read the original on arXiv AI →

The paper evaluates an agentic harness that combines a 27B language model, financial calculations, narrative templates, and validation checks to produce corporate and investment banking presentation decks. Using a panel of five judges, the system consistently scores higher than a baseline model that generates directly from a short prompt, with scores ranging from 20.4 to 33.6 out of 95. However, judge variability and changes in grading criteria make it challenging to discern small improvements in deck quality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 21

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

JudgeSense is a benchmark comprising 880 items from human‑labelled corpora, each presented under two differently worded instructions that ask the same question. The study evaluates 25 judges from six providers across four tasks, measuring how rewording affects agreement with the judge’s own verdicts. Results show that rewording reduces agreement on all tasks, with significant effects on two, and that stability varies across tasks and is not predicted by parameter count.

By Rohith Reddy Bellibatlu, Edward Raff, Wenbin Zhang
arXiv Computation and Language
Aug 27

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.

By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang