The State Of LLMs 2025: Progress, Problems, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
A large language model trained on synthetic limit order book data can generate valid sequences of LOB events with near‑perfect accuracy, yet its internal world model does not capture the true state of the book. This shortfall results in biased estimates and misleading predictability when the model is used to forecast future LOB events. The study introduces new tests for an LLM’s world model, extending previous deterministic analyses to the stochastic dynamics inherent in limit order books.
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
The paper demonstrates that deep limit order book forecasting models can be repurposed to quantify scenario-conditioned market impact without retraining. By injecting counterfactual order‑book messages into a trained Transformer forecaster, the authors compare predictive distributions before and after the injection, defining a short‑horizon model‑implied market impact. The approach achieves a Spearman correlation of 0.99 and 97.2% directional agreement with historical outcomes for non‑neutral scenarios, and captures incremental sequence‑dependent variation beyond scenario identity and pre‑event forecasts.
Deep Limit Order Book forecasting models capture nonlinear market dynamics, but their ability to quantify the effects of counterfactual order book messages has not been systematically validated. We in...
arXiv:2604. 06543v2 Announce Type: replace-cross Abstract: In this work, we demonstrate that reliable stochastic sampling is a fundamental yet unfulfilled requirement for Large Language Models (LLMs) operating as agents.
arXiv:2608. 11215v1 Announce Type: new Abstract: Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent.
arXiv:2606. 07624v1 Announce Type: new Abstract: This discussion argues that sequential statistical inference can naturally contribute to LLM trustworthiness.
arXiv:2606. 03685v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) improves end-to-end classical planning in large language models (LLMs), but do these models also learn to represent and reason about the planning problems they are solving?
Text-conditioned time-series forecasting predicts a series from both its numerical history and natural-language context, allowing forecasts to account for events and constraints that the past alone cannot reveal. This requires both reliable numerical forecasting and the ability to interpret contextual information.
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem.
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social events unfold, has barely been measured.
arXiv:2602. 12756v2 Announce Type: replace Abstract: Large Language Models (LLMs) have recently shown exceptional potential in time series forecasting (TSF), leveraging their inherent sequential reasoning capabilities to model complex temporal dynamics.
arXiv:2607. 24892v1 Announce Type: cross Abstract: Text-conditioned time-series forecasting predicts a series from both its numerical history and natural-language context, allowing forecasts to account for events and constraints that the past alone cannot reveal.