BENCHCOMPASS: From Scores to Signals for Training and Harness Decisions in Payment-Domain LLMs
Read the original on arXiv AI →BENCHCOMPASS is a new payment‑domain benchmark that transforms typed evidence packs into scenario‑grounded tasks, applies LLM‑based quality checks, generates attack variants, and reserves final item admission for domain experts. It includes an expert‑reviewed Pro benchmark covering payment knowledge, context‑grounded scenario reasoning, and attacked open robustness, plus a lower‑assurance Normal pool. Across 16 model variants, BENCHCOMPASS reveals distinct failure modes—missing payment knowledge, incomplete reasoning, and failure to reject invalid workflows—while the best model scores 89.6% on open context‑grounded reasoning and 81.7% under attacked inputs. "whyItMatters":"The benchmark provides a structured way to isolate and evaluate specific weaknesses in LLMs for payment operations, a critical financial infrastructure where rules change rapidly and decisions depend on complex contextual factors."
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.