arXiv Computation and Language

TxSum: User-Centered Ethereum Transaction Understanding with Micro-Level Semantic Grounding

TxSum introduces a user-centered approach to understanding Ethereum transactions by providing structured, risk-aware explanations grounded at the token‑flow level. The authors built a dataset of 187 complex transactions with 2,375 token‑flow annotations and transaction‑level summaries, and developed MATEX, a multi‑agent framework that retrieves external knowledge and audits explanations for factual consistency. MATEX outperforms existing baselines, improving user comprehension from 52.9% to 76.5% and increasing malicious‑transaction rejection from 36.0% to 88.0% while keeping false‑rejection rates low.

arXiv AI
Jul 21

Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification

arXiv:2607. 17586v1 Announce Type: cross Abstract: Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous nature of transactional and behavioural data.

By Yuge Zhang, Yuanxing Zhang, Yichao Jin, Khairul Amsyar Mohd Razis, Nicholas Qi An Choo, Kai Yin Anders Wong, Xinyan Tang, Kenneth Zhu Ke, Wee Keong Dennis Lee, Jingyuan Zhao
arXiv Computation and Language
Aug 25

GRACE: Step-Level Benchmark for Faithful Reasoning over Context

GRACE is a step‑level benchmark for evaluating the faithfulness of chain‑of‑thought reasoning over context. It provides human annotations for each step in CoT traces from 10 models across 4 datasets, labeling faithfulness, error category, and natural‑language explanations. The benchmark introduces a data‑driven taxonomy that splits errors into GRACE‑Inference (deductive) and GRACE‑Grounding (factual) tracks, each with four categories, and demonstrates that incorporating step‑level faithfulness signals can improve downstream accuracy and reasoning reliability.

By Hoang Pham, Dong Le, Anh Tuan Luu
arXiv AI
Jul 28

Traceable LLM Reasoning for Fake-Order Fraud Detection

arXiv:2607. 23075v1 Announce Type: cross Abstract: Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approaches often rely on expert-designed features, produce black-box decisions, and provide limited interpretability.

By Siqi You, Bingsong Xu, Zhixian Zheng, Xinjian Peng, Yang Xie, Ying Wang, Jiarong Xu