Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,035 stories · RSS feed

arXiv Computation and Language
1d ago

Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge

arXiv:2610.08675v1 Announce Type: new Abstract: Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while cit...

By Chuhong Xu (Sofia University), Bo Su (Indiana University), Ziyao Chen (University of California, San Diego), Ruiyang Xu (Northeastern University), Shimeng Dai (Michigan State University), Xinyu Qiu (Northeastern University)
arXiv Machine Learning
1d ago

Machine Learning for German Redispatch Forecasting under Data Delays and Temporal Distribution Shift

The study evaluates probabilistic machine‑learning models for forecasting German grid redispatch volumes under data‑delay constraints. Using 48,242 records from 2021‑2024, boosted‑tree LightGBM with rolling calibration achieved the best performance (nWIS 0.7767), outperforming autoregressive and seasonal baselines. Neural models with zero‑censored outputs performed similarly but revealed undercoverage during high‑volume events.

By Faraz Shamim (KIST Medical College and Teaching Hospital, Nepal), Faris Shamim (OTH Regensburg)
arXiv Machine Learning
1d ago

Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling

The paper introduces Hierarchical Continuous Diffusion Language Models (H-CDLMs), a framework that jointly diffuses multiple token modalities—individual tokens and coarser clusters of token embeddings—to enhance continuous diffusion language models. Applied to the CoBit architecture, the resulting H-CoBit achieves significant empirical gains, improving MAUVE scores and achieving lower generative perplexity on LM1B and OWT, while also outperforming prior continuous diffusion models on GSM8K. The approach generalizes to other continuous generative paradigms, as shown by consistent improvements when applied to the flow matching model FLM.

By Mathias Ollu, Nikos Komodakis
arXiv Computation and Language
1d ago

Readout Stability in Prefill-Only Decision Models:Zero-Label Prediction and Inference-Time Compute Allocation

Prefill‑only decision models evaluate every candidate in a menu in a single forward pass, avoiding decoding and reducing cost by one to two orders of magnitude compared to generative language models. The paper demonstrates that when only the candidate menu changes, the model’s post‑intervention accuracy can be predicted solely from the cached first‑pass distribution using a simple estimator that renormalizes and selects the argmax, without any labels or second pass. Across seven model families, ten datasets, and two task types, this menu‑only intervention prediction is within 4.2 points of actual accuracy, and in one family it is exact, whereas a probability‑level variant fails by 21 points, indicating the property resides in ranking rather than calibrated probabilities.

By Ran Li, Lei Chen
arXiv Computation and Language
1d ago

Evaluating Large Language Model Raters for German Open-Response Clinical Questions: A Physician-Annotated Benchmark Study of Agreement, Evaluator Bias, and Abstention

The study introduces MedQADE, a German open‑response clinical benchmark with 3,800 question‑answer pairs and physician reference annotations. It evaluates large language models (LLMs) as judges, finding that while some LLMs (e.g., Gemini 3 Flash) achieve physician‑level agreement on correctness, they exhibit self‑bias and low abstention rates. Physicians showed moderate agreement on correctness but limited agreement on difficulty, and their abstention increased with perceived difficulty.

By William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian L\"ons, Ronald B\"ock, Sebastian Fudickar
arXiv Computer Vision
1d ago

Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning

The paper investigates how vision‑language models acquire and specialize semantic capabilities during fine‑tuning, proposing a trajectory‑based framework that separates acquisition, optima, and specialization phases. It introduces Structured Semantic Routing (SSR) to analyze how supervision representation affects what is learned, showing that SSR improves name‑free attribute‑profile retrieval and that different capabilities peak at different training stages. The study reveals that continued optimization can preserve target‑class retrieval while diminishing transferable semantic knowledge, highlighting a trade‑off between specialization and generalization.

By Suguru Onda, Matthew Bailey, Ryan Farrell
arXiv Computer Vision
1d ago

Latent-Action-Guided Video-Language Feature Learning for Surgical Instrument-Tissue Interaction Recognition

The paper introduces Latent-Action-Guided Video-Language Feature Learning (LAG-VLFL) for recognizing surgical instrument–tissue interactions. By compressing frame-to-frame feature changes into latent actions and predicting next‑frame features, the method aligns video and textual action descriptions without requiring extra spatial or motion annotations. Experiments show that LAG-VLFL improves interaction grounding, temporal‑direction sensitivity, and achieves competitive recognition with faster inference and lower INT4 accuracy loss compared to V‑JEPA2/2.1.

By Jiajun Cheng, Sainan Liu, Subarna Tripathi, Xiaofan Yu, Shan Lin