Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

14,166 stories · RSS feed

arXiv AI
3d ago

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies

arXiv:2604. 09860v4 Announce Type: replace-cross Abstract: The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing.

By Jenai Xuning Yang, Rishit Dagli, Alex Zook, Hugo Hadfield, Ankit Goyal, Stan Birchfield, Fabio Ramos, Jonathan Tremblay
arXiv Machine Learning
3d ago

Macroeconomic Forecasting with Large Language Models

arXiv:2407. 00890v5 Announce Type: replace-cross Abstract: This paper presents a comparative analysis evaluating the accuracy of Large Language Models (LLMs) against traditional macro time series forecasting approaches.

By Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar
arXiv AI
3d ago

Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152

arXiv:2608. 16394v1 Announce Type: new Abstract: Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards.

By Vahid Zolfaghari, Nenad Petrovic, Andr\'E Schamschurko, Alois Knoll