Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,737 stories · RSS feed

arXiv Computation and Language
1d ago

Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge

arXiv:2610.08675v1 Announce Type: new Abstract: Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while cit...

By Chuhong Xu (Sofia University), Bo Su (Indiana University), Ziyao Chen (University of California, San Diego), Ruiyang Xu (Northeastern University), Shimeng Dai (Michigan State University), Xinyu Qiu (Northeastern University)
arXiv Machine Learning
1d ago

Machine Learning for German Redispatch Forecasting under Data Delays and Temporal Distribution Shift

The study evaluates probabilistic machine‑learning models for forecasting German grid redispatch volumes under data‑delay constraints. Using 48,242 records from 2021‑2024, boosted‑tree LightGBM with rolling calibration achieved the best performance (nWIS 0.7767), outperforming autoregressive and seasonal baselines. Neural models with zero‑censored outputs performed similarly but revealed undercoverage during high‑volume events.

By Faraz Shamim (KIST Medical College and Teaching Hospital, Nepal), Faris Shamim (OTH Regensburg)
arXiv Machine Learning
1d ago

Denoising Hierarchical Representations: Joint Continuous Diffusion for Language Modeling

The paper introduces Hierarchical Continuous Diffusion Language Models (H-CDLMs), a framework that jointly diffuses multiple token modalities—individual tokens and coarser clusters of token embeddings—to enhance continuous diffusion language models. Applied to the CoBit architecture, the resulting H-CoBit achieves significant empirical gains, improving MAUVE scores and achieving lower generative perplexity on LM1B and OWT, while also outperforming prior continuous diffusion models on GSM8K. The approach generalizes to other continuous generative paradigms, as shown by consistent improvements when applied to the flow matching model FLM.

By Mathias Ollu, Nikos Komodakis
arXiv Computation and Language
1d ago

Readout Stability in Prefill-Only Decision Models:Zero-Label Prediction and Inference-Time Compute Allocation

Prefill‑only decision models evaluate every candidate in a menu in a single forward pass, avoiding decoding and reducing cost by one to two orders of magnitude compared to generative language models. The paper demonstrates that when only the candidate menu changes, the model’s post‑intervention accuracy can be predicted solely from the cached first‑pass distribution using a simple estimator that renormalizes and selects the argmax, without any labels or second pass. Across seven model families, ten datasets, and two task types, this menu‑only intervention prediction is within 4.2 points of actual accuracy, and in one family it is exact, whereas a probability‑level variant fails by 21 points, indicating the property resides in ranking rather than calibrated probabilities.

By Ran Li, Lei Chen
arXiv Computation and Language
1d ago

Evaluating Large Language Model Raters for German Open-Response Clinical Questions: A Physician-Annotated Benchmark Study of Agreement, Evaluator Bias, and Abstention

The study introduces MedQADE, a German open‑response clinical benchmark with 3,800 question‑answer pairs and physician reference annotations. It evaluates large language models (LLMs) as judges, finding that while some LLMs (e.g., Gemini 3 Flash) achieve physician‑level agreement on correctness, they exhibit self‑bias and low abstention rates. Physicians showed moderate agreement on correctness but limited agreement on difficulty, and their abstention increased with perceived difficulty.

By William Philipp, Finn Fassbender, Daniel Fister, Thorsten Langer, Martje G. Pauly, Rebecca Herzog, Markus A. Hobert, Theresa Paulus, Alexander Baumann, Chi Wang Ip, Lukas L. Goede, Johanna Reimer, Sebastian L\"ons, Ronald B\"ock, Sebastian Fudickar
arXiv Computer Vision
1d ago

Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning

The paper investigates how vision‑language models acquire and specialize semantic capabilities during fine‑tuning, proposing a trajectory‑based framework that separates acquisition, optima, and specialization phases. It introduces Structured Semantic Routing (SSR) to analyze how supervision representation affects what is learned, showing that SSR improves name‑free attribute‑profile retrieval and that different capabilities peak at different training stages. The study reveals that continued optimization can preserve target‑class retrieval while diminishing transferable semantic knowledge, highlighting a trade‑off between specialization and generalization.

By Suguru Onda, Matthew Bailey, Ryan Farrell
arXiv Computer Vision
1d ago

Latent-Action-Guided Video-Language Feature Learning for Surgical Instrument-Tissue Interaction Recognition

The paper introduces Latent-Action-Guided Video-Language Feature Learning (LAG-VLFL) for recognizing surgical instrument–tissue interactions. By compressing frame-to-frame feature changes into latent actions and predicting next‑frame features, the method aligns video and textual action descriptions without requiring extra spatial or motion annotations. Experiments show that LAG-VLFL improves interaction grounding, temporal‑direction sensitivity, and achieves competitive recognition with faster inference and lower INT4 accuracy loss compared to V‑JEPA2/2.1.

By Jiajun Cheng, Sainan Liu, Subarna Tripathi, Xiaofan Yu, Shan Lin
arXiv AI
1d ago

Confidence Reasoning Graphs: Structured Confidence Estimation for LLM Agents

The paper introduces Confidence Reasoning Graphs (CRGs), an inference-time framework that estimates the probability an LLM agent has successfully completed its task using a single trajectory without requiring privileged model access or training data. CRGs decompose the agent’s overall claim of success into contextualized sub-claims, estimate confidence for each terminal claim based on trajectory evidence, and aggregate these into an overall confidence estimate. Experiments across multiple benchmarks, models, and agent frameworks show that CRGs produce better-calibrated confidence and more effective risk-aware decision making than existing verbalized, sampling-based, and white-box surrogate methods, while also providing transparent, auditable evidence for each estimate.

By Brendan King, Farima Fatahi Bayat, Jean-Flavien Bussotti, Pouya Pezeshkpour, Estevam Hruschka
arXiv AI
1d ago

Massive Activation Gating Channel in Large Language Models

The paper identifies a single input embedding channel, called the massive activation gating channel (MAGC), that controls the emergence of massive activations in large language models. When the MAGC value is sufficiently large or small, the spike feed‑forward network outputs exhibit exceptionally large magnitudes. The authors verify MAGC across six models and provide a theoretical explanation linking the channel to a quadratic form that mixes specific columns of the down‑projection matrix, which produce massive activations.

By Minjia Mao, Shi Chen, Bowen Yin, Xiao Fang
arXiv AI
1d ago

OTel: Open Telco AI Datasets, Benchmarks, and Models

OTel is an open telecom AI resource that provides derived datasets for retrieval, reranking, instruction tuning, and safety/abstention, along with 30 full‑parameter post‑trained baselines covering 10 embedding models, 3 rerankers, and 17 language models. The project has seen significant community engagement, with over 16 million model downloads and more than 157 media mentions by May 2026. Post‑training on OTel data improves performance across all model families, achieving 93.1% NDCG@10 for embeddings, 0.947 MRR@10 for rerankers, and 87.8% correctness for language models.

By Farbod Tavakkoli, Gregory Diamos, Kenneth Church, David Kanter, Mark Austin, Imtiaz Karim, Mirza Masfiqur Rahman, Merouane Abdelkader Debbah, Zeinab Nezami, Ali Maatouk, Leandros Tassiulas, Rex Ying, Nick Sorros, Louis Powell, Nikolaos Vasiloglou, Ashish Vaswani, Somanshu Singla, Adarsh Chaluvaraju
arXiv AI
1d ago

Thin Evidence, Thick Priors: How Language Models Substitute Identity for Missing Financial Facts

The study investigates how large language models (LLMs) compensate for missing financial information by substituting user identity cues. Using 96,600 prompts to Llama‑3.1‑8B‑Instruct, the authors varied the amount of financial facts provided while keeping the underlying finances constant, and measured changes in recommended equity allocations across 100 financial profiles, 138 personas, and seven disclosure conditions. Results show that as financial facts are removed, the influence of identity on advice grows dramatically—from 5 % of variation with full disclosure to 96 % with none—while household size becomes the most reliable predictor when evidence is scarce, and gender effects persist even after controlling for standard errors. "whyItMatters":"The findings highlight that LLM‑based advisory systems can produce biased financial recommendations when users provide incomplete information, underscoring the need for audits that reflect real‑world disclosure levels and consider the full spectrum of user identities."

By Saanvi Khetan, Sankar Balasubramanian
arXiv AI
1d ago

Personal-Agent Mediated Recommendation with Cross-Platform User History

The paper introduces Personal-Agent Mediated Recommendation, a new paradigm where a personal LLM agent uses cross‑platform user history to adjust a platform’s recommendation ranking. It presents MediateRec, a benchmark for evaluating this mediation, and proposes Personal Attribution Mediation Optimization (PAMO) to balance beneficial rescues against harmful overrides. Experiments show that PAMO improves over outcome‑only reinforcement learning, achieving a better rescue‑harm trade‑off on both synthetic and real cross‑platform tests.

By Yu Xia, Jiangfan Zhang, Jun Xiao, Julian McAuley, Xiangjun Fan
arXiv AI
1d ago

ShanLiangRen: A Nutrition Agent for Personalized Daily Meal Planning

ShanLiangRen is a nutrition agent designed to generate personalized daily meal plans that satisfy both user constraints and multidimensional nutritional goals. It transforms dietary specifications, nutrient data, user attributes, and natural language requirements into a constrained planning instance, then uses a retrieval‑augmented generation approach to narrow the candidate set from a large ingredient and recipe space. Finally, it refines plans via Pareto‑guided iterative revisions with an LLM, producing fully quantified meal plans with explicit ingredients, portion sizes, and compliance reports.

By Miao Xie, Xiao Zhang, Yuan Wang, Ruixin Zhu, Chunli Lv
arXiv AI
1d ago

Small Language Models for Smart Data Model Classification at the Edge: A Cost-Aware Hybrid Approach

The paper introduces a cost‑aware hybrid approach that uses small, open‑source language models to classify smart data models (SDMs) in Internet of Things (IoT) environments. It benchmarks general‑purpose, reasoning‑specialized, and code‑specialized models across domain‑specific datasets, highlighting their suitability for edge devices with limited resources. The study also compares these lightweight models to large language models and near‑zero‑cost baselines such as TF‑IDF and a lightweight sentence encoder to demonstrate practical performance gains.

By Cristian Martella, Angelo Martella, Antonella Longo, Motaz Saad