arXiv Machine Learning

Population-Calibrated Graph Screening at 835-Million-Address Scale, with Label-Free Transfer to New Chains

The paper presents a deployed system that scores blockchain addresses using their position in a massive multi‑chain transaction graph instead of relying on sanctions lists. The system operates on a single graph of 835 million addresses and 15.8 billion edges across five EVM chains, employing a shared inductive encoder with per‑chain normalization and two scoring heads. It demonstrates label‑free transfer, achieving high recall on held‑out positives for Base, Arbitrum, and Gnosis at a very low alert rate, and shows significant lead time over external registry events, while maintaining fast, reproducible serving performance.

Hugging Face Trending Papers
Sep 2

Population-Calibrated Graph Screening at 835-Million-Address Scale, with Label-Free Transfer to New Chains

The paper presents a compliance screening system that evaluates blockchain addresses by their position in a large multi‑chain transaction graph instead of relying on sanctions lists. Using a single graph of 835 million addresses and 15.8 billion edges across five EVM chains, the system employs a shared inductive encoder with per‑chain normalization and two scoring heads, with decision thresholds set as exact quantiles of the score distribution. The authors demonstrate label‑free transfer, achieving high recall on held‑out positives for Base, Arbitrum, and Gnosis, and report significant lead‑time in flagging external registry events, efficient serving latency, and robustness checks against adversarial behavior.

arXiv AI
Sep 17

Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds

The paper introduces DISCERN, a two-tier protocol for certifying that updates to production models do not increase risk. It first uses unlabeled data to detect benign updates based on disagreement rates, then selectively labels only disagreements through an anytime-valid confidence sequence. The method achieves finite-sample validity with label-complexity bounds of order ρ²/ε², demonstrating significant label savings and strong empirical performance across 14,000+ audit streams.

By Vishnu Bindu Balachandran
arXiv AI
Sep 3

Coverage, Not Targeting: A Structural Regime in Multi-Turn Agent Credit Assignment

The paper argues that in multi‑turn agentic reinforcement learning, credit assignment should be viewed as a coverage problem rather than a targeting problem. It introduces verifier information density (V_d) as a structural metric, showing that terminal‑state verifiers operate in a low‑V_d regime where targeting fails. Experiments on tau^2‑bench, BFCL, and ToolACE‑2‑8B demonstrate that uniformly distributing reward across all turns outperforms sparse, targeted rewards, and that full chain coverage is necessary for optimal performance.

By Chenyu Zhou, Qiliang Jiang, Shuning Wu, Xu Zhou
arXiv Machine Learning
Sep 24

From Reasoning Strings to Partial Orders: Verifier-Certified Rule Transport through Quotient Policy Optimization

The paper introduces Verifier-Certified Rule Transport (VCRT), a method that uses native verifiers to replay adjacent operation pairs and identify commutation certificates or anti-diamonds, thereby distinguishing true logical dependencies from mere serialization choices in reinforcement learning with verifiable rewards. VCRT assigns policy credit based on the total probability mass of each certified orbit and imposes constraints on post-swap consistency, source retention, and policy drift. In leave-one-environment-out transfer experiments across ProofWriter, CLRS, and Lean, VCRT achieves a 77.60% macro pass rate, outperforming the strongest baseline by 13.06 points, with the largest gains observed in Lean.

By Bang Xie, Hao Liu, Zhiyuan Peng, Xin Yin, Chenhao Ying, Yuan Luo, Senjian Zhang, Wei Chen
arXiv AI
Aug 5

BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests

arXiv:2608. 02685v1 Announce Type: cross Abstract: Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcome or a fixed change sequence.

By Zetong Xiong, Qiao Zhao, Jun Zhang, Xueying Lyu, Zhi Li, Yixiang Tu, Xiaowen Yang, Yunjie Zhang, Yufeng Wang, Zhe Zhang, Kaize Yu, Hanwen Du, Zhongkai Sun, Zhuoxin Liu, Zekun Lin, Jianwen Yang, Ruining Chen, Ying Zhang, Tingxuan Pan, Ke Chen, Shubin Han, Chuanhao Sun, Yehua Yang