arXiv Machine Learning By Soumyadeep Roy

A Fairness Audit of the Duckworth-Lewis-Stern Method: Format-Specific and Gender-Differential Bias, with an Interpretable Calibration Layer for Cricket Target Revision

Read the original on arXiv Machine Learning →

The paper audits the Duckworth‑Lewis‑Stern (DLS) method, the standard for revising cricket scores after rain, using 8,150 international matches to generate 233,550 synthetic interruption scenarios. It finds two structured biases: a 137‑run prediction error range across match‑state buckets and a gender‑differential bias in ODIs, with women’s scores over‑predicted by an average of +7.63 runs versus +1.51 runs for men. The authors benchmark DLS against five modern machine‑learning models and introduce DLS‑Cal, a lightweight calibration layer that reduces overall bias by 31% in ODIs and 19% in T20Is, and a gender‑aware variant that nearly eliminates the residual bias for women’s ODIs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 19

Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)

The paper presents a validated protocol for adapting drone‑based crowd‑counting models to the extreme conditions expected at the 2034 FIFA World Cup in Saudi Arabia. Using 525 controlled runs and a full‑resolution corpus, the authors demonstrate that label‑free adaptation can recover 31‑49% of shift‑induced error across multiple corruptions and severities, achieving a 41.8 MAE improvement over a frozen source model. They also introduce a severity law, a stability budget, and a flux‑based risk module that detects real congestion episodes, culminating in a six‑point deployment protocol for safe aerial crowd monitoring.

By AlAnoud AllGhayth, AlJawharh AlOtaibi, Jude AlSubaie
arXiv Machine Learning
Sep 14

Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration

The paper investigates how large language models (LLMs) used as judges in absolute scoring tasks exhibit systematic biases that compromise reliability. It shows that a judge’s task accuracy strongly predicts both its judging accuracy and its directional bias, yet more capable examinee models consistently receive more lenient judgments. To mitigate these biases, the authors propose a calibrated weighted majority voting (WMV) ensemble that estimates judges’ error rates from inter-judge agreement patterns, achieving near-oracle performance without labeled data and improving both accuracy and fairness.

By Gemma Zhang, Prachi Badarayani, Asmi Kumar, Sadid Hasan, Sulaiman Vesal