arXiv AI By Julie Yu, Rock Yuren Pang, Jevan Hutson, Katharina Reinecke

Visible to the Court: How AI Is (and Isn't) Litigated in U.S. Federal Court Opinions

Read the original on arXiv AI →

arXiv:2607. 23888v1 Announce Type: cross Abstract: In the United States, artificial intelligence (AI) is rapidly deployed amid limited federal regulation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 17

Legal LLM Hallucination Should Be Evaluated as Failure of Legal Warrant

The paper argues that hallucinations by legal language models should be judged as failures of legal warrant rather than mere factual or citation errors. It defines claim-authority warrant as a context-sensitive relationship between a legal claim and applicable, current authority, and proposes that evaluating warrant can uncover failures missed by traditional accuracy or citation metrics. The authors outline a pilot study, benchmark specifications, and a research agenda to assess whether legal AI systems’ claims are properly licensed by law.

By Maksym Taranukhin, Vered Shwartz
arXiv AI
Sep 10

Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?

The paper investigates whether large language model (LLM) chatbots can emulate human legal judgments of reasonableness. By comparing responses from 26 LLMs to those of human participants across 25 legal scenarios, the study finds that chatbots generally track human answers but tend to produce more homogeneous, government‑ and corporation‑friendly responses and align more closely with white, male, older, and more educated respondents. The authors note that these patterns warrant further systematic research.

By Nirav Patel, Emily Wenger, Christopher Buccafusco
arXiv AI
Jun 17

The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act

arXiv:2606. 18158v1 Announce Type: cross Abstract: Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate whether they perform doctrinal legal reasoning, which forms the interpretive core of legal work, rather than the ancillary, paralegal tasks that most current legal-AI evaluations measure.

By Mich\`ele Finck