Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,035 stories · RSS feed

arXiv Machine Learning
4d ago

Conformal Prediction for Time Series with Deep Sequence Models

The paper investigates how deep sequence models—such as recurrent neural networks and Transformers—can be integrated into conformal prediction for time series. It explores three methods: conditional quantile regression, conditional quantile function estimation, and localized conformal prediction, providing theoretical asymptotic conditional coverage guarantees for each. Experiments on real-world datasets demonstrate the practical effectiveness of these approaches.

By Junghwan Lee, Jonghyeok Lee, Yao Xie
arXiv Machine Learning
4d ago

Probe the Harness: Setup Checks for Stale-Data RL Comparisons in Language Models

The paper introduces PTH (Probe The Harness), a set of checks designed to expose hidden details in experimental setups that can alter the ranking of stale-data reinforcement learning methods for language models. By applying PTH to a comparison between SAN and truncated importance sampling (TIS), the authors demonstrate that subtle harness configurations—such as how PPO ratios are computed, data seeding, replay queue reuse, and loss normalisation—can reverse the observed performance order. The study provides a detailed signature of each influencing factor, reference results for TIS and uncorrected GRPO, and a checklist to ensure fair comparisons.

By Taiheng Pan
arXiv AI
4d ago

When Terminal-Agent Training Stalls: Demystifying Data Generation and Verification Challenge

The paper investigates why training terminal agents often stalls when using a meta‑agent like Claude Opus to generate tasks and verifiers. It identifies three failure modes—benchmark invalidity, harness brittleness, and reward misalignment—and shows that prompt redesign and context extension can improve solvability by 5.6×, yet a 9B model still tops out at 81.3% mean pass@2. Adding hard tasks drops performance to 20.6%, highlighting that solvability depends on the model used.

By Xi Qin, Isabel Kurth, Xin Cui, Elin Park, Alexander Schaefer, Yaad Oren
arXiv Computation and Language
4d ago

Understanding Clinical Cognitive Dialogues Using Large Language Models

The paper introduces a de‑identified corpus of 33 in‑person cognitive assessment conversations, comprising 8,250 utterances annotated for three speaker roles and 56 dialogue acts. The authors benchmark large language models on fine‑grained dialogue‑act classification and next‑patient‑utterance generation, finding that instruction tuning and reasoning‑aware fine‑tuning improve performance but that models still struggle with closely related dialogue acts. The corpus and benchmark are presented as tools to measure interaction structure in cognitive assessments and to support future research on conversational markers, clinician education, and validated simulated patients.

By Vishalakshi Arumugam, Dan Schumacher, Veronica Rammouz, Erfan Nourbakhsh, Enrique Gonzalez Guerrero, Jeremy Davis, Anthony Rios
arXiv Computation and Language
4d ago

Encoded but Not Routed: Explaining the Table-Chart Gap in Scientific Claim Verification

The study investigates why multimodal large language models (VLMs) perform better at verifying scientific claims when evidence is presented as a table rather than a chart, despite both formats containing the same data. Using layer‑wise linear probing and attention analysis on three open‑weight VLMs, the authors find that chart information is indeed encoded in intermediate representations but never reaches the prediction layer, a gap absent for tables. Attention patterns reveal that this disconnect manifests differently across model families, suggesting the issue lies in how encoded visual data is utilized at prediction time rather than in the encoding process itself.

By Sunisth Kumar, Xanh Ho, Tim Schopf, Andre Greiner-Petter, Florian Boudin, Akiko Aizawa
arXiv AI
4d ago

Corrupted but Correct: Why Vision-Language Models Lie to Themselves Internally

The paper demonstrates that a targeted adversarial perturbation can reduce a vision‑language model’s training loss to near zero for a fixed target caption, yet the same model, when generating freely, still produces the correct description. This phenomenon, termed the train/inference gap, is traced to a single autoregressive step where the target token’s rank is fixed across all images, and further analysis shows that the language decoder, rather than the visual encoder, determines whether the corrupted signal is amplified or suppressed. The study uses a controlled two‑stage PGD attack on Qwen2.5‑VL‑7B‑Instruct and evaluates the effect on 200 held‑out COCO images, revealing that adversarial robustness in autoregressive VLMs largely depends on the language decoder’s prior. whyItMatters":"The findings suggest that defenses and faithfulness evaluations for deployed vision‑language models should focus on the language decoder rather than the visual encoder, as the former is the key determinant of robustness to adversarial perturbations."

By Arun Josephraj Arokiaraj, Zekun Wu, Adriano Koshiyama
arXiv AI
4d ago

Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT

The paper introduces Follow the Winners (FTW), a critic‑free policy‑learning algorithm that adapts the cross‑entropy method for reinforcement fine‑tuning of agentic large language models. FTW replaces group rollouts with an ordinal filter on replay‑buffer samples, achieving polynomial concentration in the order statistic of returns and offering a bounded risk‑seeking offset that balances variance reduction. Experiments on Sokoban and Search‑R1 show that FTW matches the performance of GRPO and PPO while reducing reliance on value models or repeated rollouts.

By Joery Ari\"en de Vries, Neil David Lawrence, Zhenwen Dai
arXiv Computation and Language
4d ago

A Guideline-Augmented Multi-Agent Framework for Schema-as-Code Biomedical Named Entity Recognition

The paper introduces GAMA, a guideline-augmented multi-agent framework designed to improve biomedical named entity recognition (BioNER) using large language models (LLMs). GAMA constructs dataset-specific guideline memory by inducing and verifying annotation rules from training data, then employs a planning component to generate span-type hypotheses with rationales, a coding component to produce schema-constrained entity objects, and a verification module for structural compliance and dual-loop refinement. Experiments across five BioNER datasets demonstrate that GAMA consistently outperforms strong LLM-based baselines, with ablation studies confirming the effectiveness of each component.

By Songtao Li, Yijia Zhang, Shidi Zhang, Jianyuan Yuan, Fengyu Zhang, Hongfei Lin
arXiv Computer Vision
4d ago

A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds

The paper introduces an automated pipeline that reconstructs editable 3D procedural models of field‑grown maize directly from raw 3D point clouds, eliminating the need for manual tuning or species‑specific training data. It uses a vision‑language model to annotate leaf midlines in rendered views, then applies deterministic geometric algorithms and differentiable NURBS fitting to generate accurate plant descriptors and refine leaf surfaces. The method achieves a median Chamfer distance of 5.4 mm on 100 diverse maize plants and recovers 99.4% of reference leaves with high overlap, outperforming previous semi‑automated approaches.

By Mozhgan Hadadi, Talukder Z. Jubery, Adarsh Krishnamurthy, Baskar Ganapathysubramanian
arXiv Computer Vision
4d ago

DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

DEPICT is a new training‑free metric for evaluating text‑to‑image alignment. It replaces fixed reference answers with an agreement rule that compares image‑based and caption‑only responses, weighting questions by how decisively the caption determines them. By merging this agreement score with a holistic score, DEPICT improves negation accuracy dramatically and outperforms existing training‑free metrics while matching or exceeding fine‑tuned evaluators on several benchmarks.

By Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei, Pedro Henrique Martins
arXiv AI
4d ago

Weave Forcing: Compositional Memory Routing for Interactive Long Video Generation

Weave Forcing is a training‑free framework that enhances interactive long video generation by enabling compositional memory routing. It uses an LLM to split user prompts into character and background slots, then applies masked memory weaving with semantic masks to selectively retrieve relevant historical references. The method also introduces coverage‑adaptive RoPE to adjust temporal offsets based on reference coverage, reducing visual artifacts and improving cross‑shot consistency while preserving visual quality and text alignment.

By Ziyi Wang, Junchi Yao, Heqian Qiu, Wenbo Shi, Chengjiu Wang, Jinyang He, Binkai Hong, Hongliang Li
arXiv Computer Vision
4d ago

FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views

FRUC is a feedforward 3D Gaussian Splatting framework that reconstructs dynamic scenes from uncalibrated collaborative driving views. It uses a visual‑grounded geometric Transformer backbone for one‑shot, calibration‑free inference and introduces an ego‑centric causal occlusion field to model occlusion evolution across agents. The method performs cross‑agent integration as a deterministic residual denoising process, achieving state‑of‑the‑art rendering quality and efficiency on V2X‑Real and UrbanIng‑V2X datasets.

By Yihang Tao, Yu Guo, Zhengru Fang, Haonan An, Yuguang Fang
arXiv Computation and Language
4d ago

How Robust Is Multimodal Claim Verification to LLM Rewriting?

The paper investigates how stylistic changes introduced by large language models (LLMs) affect multimodal claim verification, a task that determines whether a textual claim is supported by given evidence. Two rewriting strategies are used: natural rewriting, mimicking typical academic polishing, and controlled injection, adding a single LLM-associated word. Across 11 open‑weight models (2B–38B parameters) from five VLM families, the study finds that most models remain robust to these modifications, showing no significant accuracy drop, though consistent probability shifts—especially under hedging conditions—are observed.

By Yun-Ang Wu, Xanh Ho, Andre Greiner-Petter, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, Akiko Aizawa
arXiv Computation and Language
4d ago

Automatic Evaluation of Mental Health Stigma in Online Communication

The paper presents a new benchmark for automatically evaluating mental health stigma in online text, featuring a fine‑grained taxonomy that covers stigma mode, domain, and specific components across multiple mental health conditions. The authors annotate naturally occurring news and social media posts and test large language models and classifiers for sentiment, toxicity, and hate speech, finding that these models poorly capture stigma and often overpredict it without explicit rules. The benchmark, annotations, exemplar cases, and code are publicly released on GitHub.

By Naomi Baes, Jemima Kang, Nick Haslam, Chris Groot, Alsa Wu, Luc Raszewski, Yulia Otmakhova
arXiv Computation and Language
4d ago

Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models

The paper introduces Adaptive Mutual Distillation (AMD), a post‑training framework that jointly trains two large language models using different task‑balancing strategies. AMD evaluates and selects distillation weight adjustments via short training probes and task‑wise validation, leading to models that outperform supervised fine‑tuning baselines on six benchmarks across three backbones. Merging the two AMD models further improves performance, surpassing multi‑task fine‑tuning by an average of 2.91 points.

By Baohang Li, Xiaocheng Feng, Yichong Huang, Chengpeng Fu, Wenshuai Huo, Zekun Zhou, Zekun Yuan, Tingjia Zhang, Bing Qin
arXiv Computation and Language
4d ago

Evaluating VQA in Vision Language Models using Cooperative Principles

The paper evaluates Vision Language Models (VLMs) on Visual Question Answering tasks where questions violate Grice's maxims. By generating question modifiers that add non-essential, ambiguous, or false information, the authors show that VLMs such as ChatGPT, Claude, Gemini, and Llava exhibit reduced performance. They also compare human pragmatic reasoning to VLM reasoning, noting differences in how each handles human‑induced versus AI‑generated violations, and find that humans spend less time resolving VLM‑induced violations while VLMs are less accurate in those cases.

By Monika Shah, Sudarshan Balaji, Somdeb Sarkhel, Sanorita Dey, Deepak Venugopal
arXiv Computation and Language
4d ago

The Geometry of Knowledge Accessibility in Large Language Models

The paper investigates how easily large language models (LLMs) can retrieve knowledge for a given query, introducing the concept of knowledge accessibility. It discovers that the accessibility of a query is reflected in its geometric position in the model’s representation space: queries closer to a central point are more accessible, while those farther away are less so. This geometric insight identifies a knowledge boundary, shows that accessibility ordering is consistent across datasets, and informs which interventions—such as query rewriting, chain-of-thought reasoning, or retrieval—are most effective depending on a query’s position relative to the center.

By Lihu Chen
arXiv Computation and Language
4d ago

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

OmniConfess is a training‑free method designed to reduce hallucinations in omni‑modal large language models (OmniLLMs) that handle text, images, audio, and video. The approach fixes a candidate response and re‑scores it at token resolution while selectively intervening on evidence from each modality, producing a token‑by‑channel confession that shows which evidence supports each part of the response. Using this confession, OmniConfess preserves grounded content and corrects commitments that rely on irrelevant or contradictory evidence. The authors evaluated the method on OmniHalluBench, a 3,540‑example benchmark drawn from six datasets across multiple modalities and tasks, and found that OmniConfess mitigates hallucinations across diverse settings.

By Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou, Kaiwen Xue, Guoxin Zhang, Yifan Zhu
arXiv Machine Learning
4d ago

Drive vs. Decay: On the Training Dynamics of Joint-Embedding Predictive Architectures

The paper introduces a stability theory for Joint-Embedding Predictive Architectures (JEPAs), showing that training dynamics involve a driving force and a decay effect that determine representation collapse. By linearising the gradient flow, the authors derive a per‑mode stability ratio that separates data‑side and predictor‑side contributions, predicting a phase boundary confirmed across 800 configurations. Using this insight, they propose ResidualPred, a transformer predictor that biases attention toward the identity at initialization, improving representation rank and downstream accuracy on tabular and image benchmarks.

By Jos\'e Lucas De Melo Costa, Seong Woo Ahn, Fabrice Popineau, Arpad Rimmel, Bich-Li\^en Doan
arXiv Machine Learning
4d ago

Does Every User Need a Private LoRA? Decoupling Personalization from Per-User Adaptation

The paper investigates how to reduce the per-user adaptation burden in personalized large language models by separating reusable personalization capacity from user-specific adjustments. Through empirical studies, it shows that shared low‑rank factors can capture much of the cross‑user structure, while a tiny user code suffices for individual correction. The proposed LINEUP framework achieves state‑of‑the‑art performance on six personalized tasks while using only eight scalars per user compared to millions of private‑LoRA parameters.

By Songyuan Sui, Srikanth Malla, Chiho Choi, Joon Hee Choi