arXiv AI

LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior

arXiv:2607. 19300v1 Announce Type: new Abstract: As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection tools and through heuristics based on language patterns.

arXiv AI
Aug 20

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.

By Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang
arXiv Computation and Language
Aug 27

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

The study investigates how prior scores influence large language model (LLM) judgments in the LLM-as-a-Judge paradigm. By testing three prompt conditions—no metadata, revision framing, and anchored metadata containing prior scores—the authors find that prior scores systematically bias evaluations, shifting ratings toward those scores across 192,000 attempts. The bias also affects categorical decisions, blocking 48% of error corrections and flipping 10.18% of correct judgments, and is not mitigated by Chain-of-Thought or a warning, underscoring the need for careful context engineering.

By Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic
arXiv Machine Learning
4d ago

RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals

The paper introduces RAISE, a diagnostic framework that tests whether a costly large language model (LLM) signal provides enough pre-call information to justify selective use. It identifies the failure mode of acquisition collapse, where an LLM appears useful overall but lacks actionable evidence for individual decisions. The authors demonstrate RAISE with Structured Hypothesis Embeddings (SHE) and evaluate it across multiple study designs, showing that predictable incremental benefit, rather than average lift, indicates recoverable selective value.

By Ying Yuan, Yu Wang, Yize Cheng, Xuyang Wu
arXiv AI
Jul 21

Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification

arXiv:2607. 17586v1 Announce Type: cross Abstract: Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous nature of transactional and behavioural data.

By Yuge Zhang, Yuanxing Zhang, Yichao Jin, Khairul Amsyar Mohd Razis, Nicholas Qi An Choo, Kai Yin Anders Wong, Xinyan Tang, Kenneth Zhu Ke, Wee Keong Dennis Lee, Jingyuan Zhao
arXiv AI
Aug 20

When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators

The paper evaluates large language models (LLMs) as data quality annotators on two e-commerce tasks: entity matching and brand mislabeling. In entity matching, a simple rule-based baseline matched the LLM’s zero-shot performance (F1≈0.95), and a few-shot prompt actually lowered performance, highlighting the risk of small-sample prompt tuning. For brand mislabeling, the LLM outperformed a naive rule baseline (F1 0.833 vs 0.721) by leveraging background knowledge, and demonstrated high consistency across repeated runs (99.7% agreement).

By Praphulla Lal Shrestha