Hugging Face Trending Papers

Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence

Read the original on Hugging Face Trending Papers →

The paper introduces the Profiling, Investigation, and Judgment (PIJ) benchmark, which contains 2,500 real homicide cases from five countries to evaluate large language models (LLMs) on pre‑arrest criminal investigation tasks. It assesses LLMs across criminal profiling, crime process reconstruction, and sentence prediction, revealing that performance drops as tasks require more implicit reasoning about unknown suspect profiles. The study finds that LLMs lag behind human experts, especially in inferential categories like motivation and victim‑offender relationships, and exhibit biases in gender, age, and motive attribution.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 18

Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence

The paper introduces the Profiling, Investigation, and Judgment (PIJ) benchmark, which contains 2,500 real homicide cases from five countries to evaluate large language models (LLMs) on pre‑arrest criminal investigation tasks. It assesses LLMs across criminal profiling, crime process reconstruction, and sentence prediction, revealing that performance drops as tasks require more implicit reasoning about unknown suspect profiles. The study finds that LLMs lag behind human experts, especially on inferential tasks like motivation and victim‑offender relationships, and exhibit biases in gender, age, and motive attribution.

By Yutong Yao, Yanjie Cao, Guanhua Chen, Xu Yang, Junchao Wu, Zeyu Wu, Lidia S. Chao, Derek F. Wong
arXiv AI
Sep 3

OBJECTION! Lawyer Agents Mitigate Guilty Bias in Legal Judgment Prediction

The paper introduces OBJECTION, an inference-time pipeline that adds an Adversarial Lawyer Agent to each of the three reasoning steps—offense, unlawfulness, and culpability—in legal judgment prediction models. By actively injecting defense arguments, the agent challenges the model’s default assumption of guilt, which is common in datasets biased toward guilty outcomes. Using a new Natural Innocent dataset of 3.4k real cases, OBJECTION reduces the False Guilty Rate from 82.93% to 16.69%, demonstrating significant improvement in substantive legal reasoning.

By Jaehoon Jeong, Jay-Yoon Lee
arXiv Computation and Language
Sep 14

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

The paper introduces Tasks over Application Manuals (TAM), a benchmark designed to test long‑horizon procedural reasoning in large language models. TAM uses real‑world tasks from ICD‑10‑CM clinical coding and U.S. federal sentencing, requiring models to follow extensive, rule‑based manuals and perform interdependent steps to produce exact answers. Experiments with GPT‑5 and various prompting strategies show very low exact‑match accuracy—1% for coding and 15.5% for sentencing—highlighting a gap between current benchmarks and the ability to reliably follow complex procedures.

By Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie, Sachin Chandrasekhar, Eugene Wen
arXiv AI
Jul 7

Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal

arXiv:2607. 04261v1 Announce Type: new Abstract: Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that models perform retrospective classification rather than true forecasting.

By Joe Watson, Joana Ribeiro de Faria, Marcus Tomalin, M{\aa}ns Magnusson, Huiyuan Xie, Hao Tian Yeung, Felix Steffek
arXiv AI
Aug 18

When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning

arXiv:2608. 14610v1 Announce Type: new Abstract: Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a case -- a capability we term temporal applicable-law determination.

By Yiqian Huang, Shuyuan Zheng, Qianying Liu, Shaowen Peng, Yuntao Kong, Kotaro Funakoshi, Chuan Xiao, Manabu Okumura, Yang Cao