arXiv AI

Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)

The paper presents a validated protocol for adapting drone‑based crowd‑counting models to the extreme conditions expected at the 2034 FIFA World Cup in Saudi Arabia. Using 525 controlled runs and a full‑resolution corpus, the authors demonstrate that label‑free adaptation can recover 31‑49% of shift‑induced error across multiple corruptions and severities, achieving a 41.8 MAE improvement over a frozen source model. They also introduce a severity law, a stability budget, and a flux‑based risk module that detects real congestion episodes, culminating in a six‑point deployment protocol for safe aerial crowd monitoring.

arXiv Computer Vision
Sep 1

SNF-Bench: Separating Static Drift from Natural Flow in Long-Horizon Fixed-Camera Video Generation

SNF-Bench is an evaluation framework for long‑horizon fixed‑camera video generation that separates static background fidelity from dynamic flow persistence and drift leakage. It reports these three factors independently, using controlled injections of translation, rotation, scale drift, and progressive freezing to validate each metric’s sensitivity. Auditing public checkpoints shows that whole‑frame motion metrics can mislead, while SNF‑Bench reveals the true trade‑offs between motion quality and background stability.

By Matiur Rahman Minar, Seunghun Oh, Ganghyeon Jeong, Unsang Park
arXiv Machine Learning
Sep 7

A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision

The paper audits the Duckworth‑Lewis‑Stern (DLS) method, the standard for revising cricket scores after rain, using 8,150 international matches to generate 233,550 synthetic interruption scenarios. It finds two structured biases: a 137‑run prediction error range across match‑state buckets and a gender‑differential bias in ODIs, with women’s scores over‑predicted by an average of +7.63 runs versus +1.51 runs for men. The authors benchmark DLS against five modern machine‑learning models and introduce DLS‑Cal, a lightweight calibration layer that reduces overall bias by 31% in ODIs and 19% in T20Is, and a gender‑aware variant that nearly eliminates the residual bias for women’s ODIs.

By Soumyadeep Roy
arXiv Machine Learning
Sep 11

How Much Velocity Does Off-Ball Space Value Need? A Broadcast-Viewport Benchmark

The paper investigates how different velocity regimes affect broadcast‑viewport basketball analytics. Using a calibrated off‑screen imputation protocol, it compares four velocity settings—none, viewport‑legal observed, true‑for‑visible, and true‑for‑all—against a velocity‑aware ground truth across three analytical layers. Results show that velocity is largely useless for imputation, modestly useful for the control surface, and minimally useful for final team verdicts, with the visible channel providing the most benefit.

By Seongjin Choi
arXiv Machine Learning
Aug 4

Real-Time Detection and Repair of LLM Agent Failures

arXiv:2608. 02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the standard remedy, judging every step with a second LLM, costs more than the agent itself.

By Sunny Dubey