Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

19,431 stories · RSS feed

arXiv AI
Jul 28

A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility

arXiv:2607. 24663v1 Announce Type: cross Abstract: Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenance records, and live control-system data.

By Rajat Sainju, Dariusz Jarosz, Hairong Shang, Michael Prince, Ryan M. Aydelott, Mathew J. Cherukara, Yine Sun, Michael D. Borland
arXiv AI
Jul 28

Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty

arXiv:2508. 08992v4 Announce Type: replace Abstract: Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provides a classic framework for modeling human behavior under such uncertainty.

By Rui Wang, Qihan Lin, Jiayu Liu, Qing Zong, Tianshi Zheng, Dadi Guo, Haochen Shi, Peixuan Han, Weiqi Wang, Yangqiu Song
arXiv AI
Jul 28

GFLAN: Generative Functional Layouts

arXiv:2512. 16275v2 Announce Type: replace-cross Abstract: Automated floor plan generation lies at the intersection of combinatorial search, geometric constraint satisfaction, and functional design requirements -- a confluence that has historically resisted a unified computational treatment.

By Mohamed Abouagour, Eleftherios Garyfallidis
arXiv AI
Jul 28

SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving

arXiv:2607. 23933v1 Announce Type: cross Abstract: As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment introduces a fundamental tension between resource utilization and interactive tail latency.

By Yihui Zhang (Beihang University), Tianyu Wo (Beihang University), Jinghao Wang (Beihang University), Xiaoyang Sun (University of Leeds), Menghao Zhang (Beihang University), Cangzhou Yuan (Beihang University), Li Li (Beihang University), Chunming Hu (Beihang University), Albert Y. Zomaya (The University of Sydney), Renyu Yang (Beihang University)
arXiv AI
Jul 28

StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech

arXiv:2607. 22658v1 Announce Type: new Abstract: Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these cues remain limited.

By Yuzhe Wang (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Thomas Thebaud (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Jennifer Hu (Department of Cognitive Science, Johns Hopkins University, Baltimore, USA), Jes\'us Villalba-Lopez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Venkatesh Ravichandran (Amazon AGI, USA), Georgi Tinchev (Amazon Research, UK), Najim Dehak (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA), Laureano Moro-Vel\'azquez (Electrical and Computer Engineering Department, Johns Hopkins University, Baltimore, USA)