Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

18,779 stories · RSS feed

arXiv AI
Jul 28

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

arXiv:2607. 22658v1 Announce Type: new Abstract: Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited.

By Yuzhe Wang (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Thomas Thebaud (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Jennifer Hu (Department of Cognitive Science, Johns Hopkins University, Baltimore, USA), Jes\'us Villalba-Lopez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Venkatesh Ravichandran (Amazon AGI, USA), Georgi Tinchev (Amazon Research, UK), Najim Dehak (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Laureano Moro-Vel\'azquez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA)
arXiv AI
Jul 28

PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis

arXiv:2607. 23794v1 Announce Type: cross Abstract: Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellular morphology at higher magnification.

By Chi Phan, Tianyi Zhang, Yufeng Wu, Qiaochu Xue, Jiajie Zhang, Linghan Cai, Zeyu Liu, Sudong Wang, Yueming Jin, Dan Hu
arXiv AI
Jul 28

Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models

arXiv:2601. 21003v3 Announce Type: replace Abstract: Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is especially severe when fine-tuned on small datasets due to the inherent tendency toward miscalibration.

By Moule Lin, Shuhao Guan, Andrea Patane, David Gregg, Goetz Botterweck
arXiv Machine Learning
Jul 28

When Low CER is Not Enough: An Analysis of Hallucinations in Vision-Language OCR Systems on Historical Uruguayan Documents

arXiv:2607. 24077v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) is a key component in the digitization of historical archives.

By Marina Gardella (CB), Camilo Mari{\~n}o (UDELAR, CB), Diego Belzarena (UDELAR, CB), Ignacio Ram{\'i}rez (UDELAR), Gregory Randall (UDELAR), Jean-Michel Morel (LU - Hong Kong)