arXiv AI

BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs

BENCHCOMPASS is a new payment‑domain benchmark that transforms typed evidence packs into scenario‑grounded tasks, applies LLM‑based quality checks, generates attack variants, and reserves final item admission for domain experts. It includes an expert‑reviewed Pro benchmark covering payment knowledge, context‑grounded scenario reasoning, and attacked open robustness, plus a lower‑assurance Normal pool. Across 16 model variants, BENCHCOMPASS reveals distinct failure modes—missing payment knowledge, incomplete reasoning, and failure to reject invalid workflows—while the best model scores 89.6% on open context‑grounded reasoning and 81.7% under attacked inputs. "whyItMatters":"The benchmark provides a structured way to isolate and evaluate specific weaknesses in LLMs for payment operations, a critical financial infrastructure where rules change rapidly and decisions depend on complex contextual factors."

arXiv AI
Jul 17

Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents

arXiv:2607. 14573v1 Announce Type: new Abstract: Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-server flows, verify payment outcomes, and preserve consistency between transaction and business states.

By Shiyu Ying, Xuejie Cao, Yingfan Ma, Yuanhao Dong, Wenyu Chen, Bowen Song, Lin Zhu
arXiv AI
Jun 29

DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain

arXiv:2504. 16116v4 Announce Type: replace-cross Abstract: The Web3 ecosystem, underpinned by cryptographic primitives and decentralized consensus, represents a high-stakes environment where software vulnerabilities and incentive misalignments translate directly into financial loss.

By Enhao Huang, Pengyu Sun, Shuxun Wang, Zixin Lin, Alex Chen, Kaichun Hu, Joey Ouyang, Frank Li, Zhiyu Zhang, Haobo Wang, Yiming Li, Zhan Qin, James Yi, Gang Zhao, Ziang Ling, Lowes Yang
arXiv AI
Aug 20

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

FinRCA-Bench is a deterministic synthetic benchmark comprising 2,250 accounts‑payable‑to‑bank reconciliation cases that span 14 operational tables and include 1,500 injected failures across 15 causal categories. The benchmark hides root‑cause labels and record‑level evidence contracts from models, enabling independent evaluation of evidence retrieval versus reasoning accuracy. Experiments show that retrieval architecture dramatically influences performance, with retrieval improvements raising macro‑required‑record recall from 0.83% to 77.70% and exact 16‑class accuracy from 2.05% to 72.44%.

By Pratik Ghawate
arXiv Computation and Language
Aug 27

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.

By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang
Hugging Face Trending Papers
6d ago

RGDT-Bench: Benchmarking LLM Reasoning for Rule-Governed Decisions and Their Justifications

RGDT-Bench is a new benchmark that evaluates large language models on Rule‑Governed Decision Tasks, where models must apply external rules to facts, justify decisions, and provide checkable justifications. The benchmark offers 202.1K condition‑level supervision slots across four task tracks and eight task‑probe combinations, and it labels warrant completeness through label‑blind extraction and deterministic checks. Evaluation shows that among correct responses, 40.2% of warrants are incomplete, and existing evaluators struggle to detect this, prompting the authors to train a reward model that improves AUROC to 69.24% and outperforms outcome‑supervised baselines.

arXiv Machine Learning
Aug 31

VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning

The paper introduces VICT, a method that leverages the internal structure of verifiable tasks to perform fine‑grained credit assignment for long‑horizon LLM agents. VICT exposes executable or evidence‑backed atoms from a task’s terminal verifier and traces them back to actions via dependency‑valid proof edges, redistributing advantage only along these edges. This approach improves performance on ALFWorld and WebShop compared to outcome‑only training and matches recent fine‑grained credit methods without requiring additional critics, labels, or inference‑time verifier access.

By Pengcheng Li, Zhengyang Zhang, Dongxu Zhang, Sui Huang, Shaohua Ma
arXiv AI
Aug 5

ZK-SR117: A Chunked Zero-Knowledge Attestation Design for Aggregated Fair-Lending Metrics, with a Control Mapping toward Full SR 11-7 Coverage

arXiv:2608. 02664v1 Announce Type: cross Abstract: Deploying ML models in regulated decision-making (credit underwriting, fraud detection, loan approval) requires demonstrating fairness and robustness to auditors without exposing model weights or customer data.

By Mohammad Nasir Uddin, Rahnuma Tabassum Orpita, Eklachur Rahman Bhuiyan, Asaduzzaman Anik