arXiv Machine Learning

ETH-TraceBench: A Large-Scale Event-Stream Benchmark for Ethereum DeFi under Temporal, Protocol, and Contract Shift

arXiv AI
Sep 7

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

FinalityBench is an executable benchmark that tests how agents decide on shipping, re‑capturing, refunding, or waiting when a merchant’s payment processor, ledger, ERP, and bank feed receive delayed, duplicated, dropped, or reordered messages, causing contradictory beliefs about an order. The benchmark uses a hidden canonical event log and faulted delivery streams to generate system views, scoring each episode by the merchant’s terminal economic position relative to a privileged reference. It contains 321 tasks, including 45 twin pairs where all four views are identical yet the correct disposition differs, and evaluates nine programmatic policies, revealing that a ship‑on‑first‑sign policy performs best by accuracy but worst by paired loss, while a runtime‑gated irreversible‑action policy achieves 85.4% accuracy without losing money.

By Abhishek Sharma
arXiv Machine Learning
Aug 19

Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal

The paper audits the impact of temporal leakage on financial-news direction prediction across 49,799 articles and 16 feature-model combinations, including TF‑IDF, MiniLM, FinBERT, and fine‑tuned RoBERTa‑large / DeBERTa‑v3‑large, as well as zero/few‑shot and LoRA probes of Llama‑3 and Qwen2.5. Random train‑test splits inflate MCC scores by 1.1× to 6.5×, with larger models and richer features showing greater gains, while end‑to‑end FinBERT fine‑tuning actually increases the gap. Only the mergers and acquisitions (M&A) category shows a positive locked‑test signal under near‑temporal chronological evaluation, with the signal localized to 2024‑2025 European‑tilted M&A semantics and not transferring to a 2009‑2020 U.S. corpus.

By Chenhao Xue, Raslen Guesmi, Siwei Feng, Yucheng Gong, Jacob Xavier Sundram, Jordan Pang, Lan Wang, Julian Kaljuvee
arXiv Machine Learning
Sep 4

Population-Calibrated Graph Screening at 835-Million-Address Scale, with Label-Free Transfer to New Chains

The paper presents a deployed system that scores blockchain addresses using their position in a massive multi‑chain transaction graph instead of relying on sanctions lists. The system operates on a single graph of 835 million addresses and 15.8 billion edges across five EVM chains, employing a shared inductive encoder with per‑chain normalization and two scoring heads. It demonstrates label‑free transfer, achieving high recall on held‑out positives for Base, Arbitrum, and Gnosis at a very low alert rate, and shows significant lead time over external registry events, while maintaining fast, reproducible serving performance.

By Yury Korolev
Hugging Face Trending Papers
Sep 2

Population-Calibrated Graph Screening at 835-Million-Address Scale, with Label-Free Transfer to New Chains

The paper presents a compliance screening system that evaluates blockchain addresses by their position in a large multi‑chain transaction graph instead of relying on sanctions lists. Using a single graph of 835 million addresses and 15.8 billion edges across five EVM chains, the system employs a shared inductive encoder with per‑chain normalization and two scoring heads, with decision thresholds set as exact quantiles of the score distribution. The authors demonstrate label‑free transfer, achieving high recall on held‑out positives for Base, Arbitrum, and Gnosis, and report significant lead‑time in flagging external registry events, efficient serving latency, and robustness checks against adversarial behavior.

arXiv Machine Learning
Sep 23

What Does 99% Accuracy Measure? A Reproducible Audit of Shortcut Learning in a Widely Used Fake News Corpus

The paper audits the widely used ISOT/Kaggle Fake and Real News corpus and finds that extremely high reported accuracies (≈0.98) are largely due to shortcut signals rather than genuine veracity detection. A simple TF‑IDF linear classifier achieves perfect F1 when using only subject metadata, and even after removing metadata, newswire tags, and duplicate documents, the F1 drops only modestly, indicating that editorial style rather than specific tokens drives performance. Under topic‑disjoint and temporal transfer tests, performance collapses, and models transfer poorly to the independent LIAR benchmark, showing that within‑corpus scores reflect source and topic separability, not truth verification. whyItMatters:"The study demonstrates that current high accuracy metrics on this fake‑news dataset are misleading, highlighting the need for more robust evaluation protocols that guard against shortcut learning."

By Yuvraj Verma