Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,449 stories · RSS feed

arXiv AI
5d ago

Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs

The paper introduces SymCE, a dataset of 4,707 false undergraduate‑algebra and real‑analysis conjectures each paired with a deterministic Python verifier that can be executed to check truth. It studies counterexample generation by training a 4B‑parameter language model (Qwen3‑4B) first with supervised fine‑tuning and then with reinforcement learning (GRPO) using the verifier as a reward. The results show that counterexample‑only fine‑tuning can destroy true‑theorem recognition, while reinforcement learning with a sparse outcome‑only reward restores and surpasses baseline performance, achieving higher success rates than larger open‑weight models and competitive performance on other math benchmarks. "whyItMatters":"The study demonstrates that reinforcement learning guided by a per‑theorem verifier can effectively repair the shortcomings of supervised fine‑tuning in theorem proving, offering a practical approach to improve language models’ mathematical reasoning capabilities."

By Omar Farouk Zouak, Houssam Eddine Boukhalfa, Soumaya Lakehal, Shiv Katiyar, Samia Nefti-Meziani
arXiv Machine Learning
5d ago

Clinical Concept Centers in LLMs

The paper investigates whether clinical concepts are represented as distinct, causally influential centers within the latent space of large language models (LLMs). By evaluating eleven open-weight LLMs, the authors discover that each model contains dedicated clinical concept centers that are interpretable, activate only on relevant clinical narratives, and drive model behavior in both constrained and open-ended contexts. These centers can be leveraged for evaluation and performance improvement, as steering models along them enhances downstream clinical outcomes and aligns with clinician preferences.

By Aishik Nagar, Abhishek Vaidyanathan, Arun-Kumar Kaliya-Perumal, Elijah Tzen Hsuen Boey, Stefan Winkler
arXiv AI
5d ago

From Retrieval to Typed Decisions: Calibrated System One Models from Biomedical Sentence Encoders

The paper introduces SBERT2S1, a method that transforms Sentence-Transformers encoders into typed decision models for biomedical text, and presents BIODECIDE, a suite for evaluating such models, along with MEDLINE‑S1, a large training set of 243k decisions derived from NLM indexing. Experiments show that retrieval‑trained encoders improve zero‑shot matching of content‑bearing options, and that a prior‑fused residual (PFR) head benefits most from retrieval pre‑training, while a cross‑head (C) head generally outperforms PFR across all objectives. The authors also release code, data, and a model, and discuss calibration and reward‑normalisation effects in their RLCD recipe.

By Pritam Deka
arXiv AI
5d ago

APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory

APDMem is a hierarchical long‑term memory system for personalized LLM assistants that uses progressive disclosure to retrieve conversation history. It organizes memory into four layers—thematic summaries, personalized key facts, turn‑level evidence notes, and raw messages—and a controller reads from the top layer, drilling down only when necessary. This approach balances cost and fidelity, allowing simple queries to finish early while deeper inspection is triggered for complex or exact‑evidence requests, and a note synthesizer structures retrieved evidence before answer generation. Experiments on LongMemEval show APDMem performs strongly while accessing only 8% of the total conversations.

By Chin-Lun Fu, Anagha Kulkarni, Hong Ni, Behrouz Madahian
arXiv AI
5d ago

CuBEs: Culturally-Situated Behavioral Evaluations and the Limitations of Culture-Blind LLM Judges

CuBEs introduces culturally‑situated behavioral evaluations for large language models, adding cultural context to test scenarios and assessments. The authors built a human‑labeled dataset covering 12 cultures, revealing significant cross‑cultural differences that one‑size‑fits‑all judgments miss. Evaluating 13 LLMs shows that culturally situated tests uncover varied behaviors, such as Western political bias versus non‑Western religious or colonial biases, which standard evaluations overlook.

By Hoda Ayad, Tanu Mitra, Abhishek Mukherji
arXiv Machine Learning
5d ago

XGenAct: Geometry-Enhanced World Action Models through Cross-Task Generation

XGenAct is a world action model that encodes RGB observations, robot actions, metric depth, surface normals, and functional role segmentation as RGB videos using deterministic codecs. It trains a single video diffusion transformer with a unified objective, sampling different perception and action tasks to learn temporal prediction across multiple spatial modalities without separate heads. Experiments on RLBench show that structured perception training boosts closed‑loop success, with XGenAct achieving 52% success on five external tasks compared to 26% for the best baselines, and it outperforms pipelines that first generate RGB and then apply a frozen perception expert for depth and segmentation prediction.

By Tingting Du, Ziyao Wang, Guoheng Sun, Ang Li
arXiv Computer Vision
5d ago

A Fully Automatic Pipeline for 3D Dendrite Instance Segmentation in SBF-SEM

The paper introduces a fully automatic pipeline for 3D dendrite instance segmentation in serial block-face scanning electron microscopy (SBF-SEM). It combines YOLOv6-guided Segment Anything Model prompting, iterative 2D mask refinement, random forest 3D linking, and high‑resolution nnU‑Net refinement to produce coherent, well‑separated dendrite reconstructions without manual prompting. Applied to hippocampal CA1 data from a control and an epileptic rat, the method achieves high semantic accuracy (Dice 0.93/0.91) and strong instance performance, while highlighting dense‑region recognition as the main challenge in epileptic tissue.

By Zewen Zhuo, Ilya Belevich, Eija Jokitalo, Alejandra Sierra, Jussi Tohka
arXiv Computer Vision
5d ago

ProgressNet: Sketching and Prompting with a Frozen Text-to-Image Model

ProgressNet is a training‑free framework that enables a frozen text‑to‑image model to follow a sketching session in real time. It allows strokes to be added or erased, prompts to be revised, and produces an updated image in about a second per turn without adding new parameters. The method uses three inference‑time mechanisms—Previous‑Concept Memory, Layer‑Selective K/V Injection, and Banded Adaptive Control—to maintain fidelity and coherence across sketch domains, outperforming existing approaches especially when sketches are partially erased.

By Arkaprabha Basu, Chaitat Utintu, Yi-Zhe Song
arXiv Computer Vision
5d ago

HakushoBench: A Japanese Chart and Table VQA Benchmark from Governmental White Papers

HakushoBench is a Japanese chart and table visual question answering benchmark created from 33 governmental white papers, comprising 2,053 images across more than ten types. The dataset includes manually annotated and independently verified QA pairs that evaluate holistic understanding of charts and tables rather than just local visual cues. Experiments show that HakushoBench is significantly harder than existing Japanese benchmarks, with sub‑10B open‑weight models achieving at most 58.6% accuracy and even large models like Qwen3.5-397B-A17B lagging behind Gemini‑3‑Pro by 8.1 points.

By Issa Sugiura, Shuhei Kurita, Yusuke Oda, Naoaki Okazaki
arXiv AI
5d ago

Large language models exhibit unreliable updating of clinical judgment as patient evidence evolves

Large language models (LLMs) are increasingly used for clinical reasoning, yet their ability to revise judgments as patient evidence changes is uncertain. In this study, researchers evaluated longitudinal belief updating using real intensive‑care patient trajectories and found that conditioning on prior judgments often increased prediction error. Controlled experiments revealed two failure modes: models were more responsive to worsening than improving respiratory evidence, and they shifted estimates significantly when prior risk was altered, indicating a causal influence of prior beliefs. Prompting did not improve reliability, and an Evidence‑Validated Longitudinal Update (EVLU) approach produced fewer but more trustworthy revisions, highlighting a reliability‑coverage trade‑off.

By Min Zeng, Rui Zhang
arXiv AI
5d ago

Coherence-Driven Belief Formation and Population Dynamics of Contagion in LLM Agents

The paper studies how language‑model agents adopt beliefs, measuring the probability that an agent accepts a claim based on the number of peers endorsing it. The adoption curve is sigmoid, indicating complex contagion, with a threshold influenced by the claim’s plausibility, the source’s reliability, and the agent’s disposition—dimensions that can be collapsed into a single coherence metric relative to the agent’s prior beliefs. In networks of AI agents, belief spread is stronger on clustered than random networks, and the system shows a bifurcating cascade window and hysteretic consensus that makes consensus difficult to reverse once formed.

By Tathagata Banerjee, Nima Moghaddas
arXiv AI
5d ago

CVE2AP: Automated Generation of PDDL-Encoded Attack Paths via Large Language Models

CVE2AP is an LLM-based system that automatically converts natural language CVE descriptions into PDDL-encoded attack paths. It uses structured prompting and an error‑feedback loop that refines outputs based on planner‑reported syntactic and solvability errors. Empirical tests across various LLMs show high quality results, with up to 86.9% syntax correctness, 78.6% solvability, and 93.1% semantic correctness, and GPT‑5.5 providing the best quality‑cost balance.

By Lin Cui, Vincenzo Scotti, Raffaela Mirandola
arXiv AI
5d ago

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

The paper investigates how training large language models to use fewer tokens in chain-of-thought (CoT) reasoning impacts the faithfulness and monitorability of the generated explanations. Three efficiency methods—fixed generation budget, per-example length target, and group-relative length reward—were applied during fine‑tuning, and the resulting models were evaluated on how well their CoT reflects decision processes and whether it signals changes due to input interventions. Results show that while faithfulness generally decreases because models become less consistent, monitorability remains relatively robust, with models still indicating the influence of input changes even when CoT is shortened.

By Samuel Lewis-Lim, Xingwei Tan, Mario Sanger, Zhixue Zhao, Nikolaos Aletras
arXiv AI
5d ago

Reasoning Models Are Accurate but Unsound on Identification

The paper introduces CERTID, a formal identification pipeline that uses the sound and complete causal identification algorithm ID to certify whether a causal effect is identifiable from a given graph and query. CERTID also verifies returned formulas against structural causal models with known interventional distributions, mitigating structural leakage and providing grading guarantees. The authors evaluate three frontier reasoning models on 1,200 certified instances, finding that accuracy is a poor proxy for soundness, with false-claim rates on non-identifiable queries varying by seventeen-fold across models.

By Arman Behnam, Binghui Wang
arXiv AI
5d ago

Causal discovery identifies pathways linking physical activity to dementia risk in the UK BioBank

The study used causal discovery and mediation analysis on 42,293 UK Biobank participants to map how moderate‑to‑vigorous physical activity (MVPA) reduces dementia risk. It found interconnected pathways involving depression, functional capacity, smoking, hypertension, cardiovascular disease, chronic kidney disease, and brain injury, with depression accounting for 15.1% of the effect. Sex‑specific analyses showed females mainly linked to metabolic and functional pathways, while males were more influenced by behavioral and cardiometabolic cascades.

By Wasif Khan, Panayiotis V. Benos, Joshua K. Wong, Ruogu Fang
arXiv AI
5d ago

Hop-Decayed Influence: New Vulnerabilities of Structural Auxiliary Indexing in GraphRAG Pipelines with LLM

The paper introduces the Hop-Decayed Influence (HDI) attack, which targets auxiliary schema-level structures—semantic summaries, hierarchical edges, and pre-computed scores—used in GraphRAG pipelines for retrieval prioritisation. By propagating query-aware influence, HDI identifies high-impact targets and corrupts a minuscule fraction (as low as 0.016%) of these structures, achieving an 88–94% success rate across HotpotQA and 2WikiMultiHopQA benchmarks on Microsoft GraphRAG and HippoRAG2 architectures. The modifications affect up to six queries each, yielding a 1:N amplification that instance-level attacks cannot achieve, and they evade perplexity and paraphrase defenses with over 99% evasion, exposing a structural blind spot in current GraphRAG defenses.

By Jisung Park, John Le, Heath Cooper
arXiv Machine Learning
5d ago

Inner Momentum for Differentially Private Muon

The paper introduces Inner Momentum (IM) for differentially private Muon training. It addresses distortion caused by per-example gradient clipping by averaging each example’s Muon gradient over the current and recent models before clipping, thereby bounding clipping-induced distortion. Experiments on private GPT‑2 fine‑tuning show that DP‑Muon‑IM consistently improves BLEU and ROUGE‑L scores and reduces polar error compared to standard DP‑Muon.

By Bishnu Bhusal, Minh Vu, Ben Southworth, Geigh Zollicoffer, Rohit Chadha, Manish Bhattarai
arXiv AI
5d ago

EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation

EVOL is a framework that uses a knowledge‑tracing simulator to generate per‑learner expert demonstrations via evolutionary search, and then trains a deployment‑free policy that distills these demonstrations into a feed‑forward learner. The method employs an asymmetric actor‑critic architecture, where the actor plans blindly as in real deployment while the critic has access to privileged simulator state during training. Across three educational datasets and varying path lengths, EVOL outperforms eight baselines and shows that the quality of evolutionary experts, rather than the imitation objective, drives final performance.

By Geonwoo Bang, Dongho Kim, Moohong Min
arXiv Machine Learning
5d ago

Exact Memory-Time Optimization for Prefix-Cached Language Model Serving

The paper introduces Prefix‑Certificate Retention (PCR), an exact optimization framework for deciding which prefix states of a language model to cache in order to balance recomputation and storage time. PCR models usable‑prefix rewards as nodes with prerequisites tied to timeout thresholds and preceding hit certificates, reducing the problem to a single minimum‑cut graph whose size scales linearly with block lookups and timeout choices. The authors provide a breakpoint theorem that extends the construction to all nonnegative timeouts without discretization error, a linear‑time dynamic program for ordered timeouts, and empirical validation on 39,632 public Mooncake requests, showing that ordered timeouts achieve the unrestricted optimum in most trace‑grouping cases and that heterogeneous retention can improve memory‑time tradeoffs. "whyItMatters":"The study offers a tractable, auditable optimization model for retention policies that directly measures usable prefix blocks and storage time, providing a concrete benchmark for evaluating memory‑time tradeoffs in language‑model serving."

By Shivam Gupta
arXiv AI
5d ago

Predicting Steering Vectors and Adapter Weights for Few-Shot Author-Style Transfer

The paper investigates how to adapt large language models to mimic an individual author's style using only a few example abstracts, a task made harder by the formal nature of scientific writing. It proposes three style‑conditioning methods—contrastive activation steering, a network predicting steering vectors, and a hypernetwork predicting LoRA adapters—and finds that while fine‑tuning captures the strongest style signal, it harms fluency; the hypernetwork offers the best balance between style imitation and output quality for both seen and unseen authors. The authors also show that steering can be performed at the author level by contrasting author abstracts against style‑neutral generations for the same content, eliminating the need for a predefined style inventory and outperforming inventory‑based approaches. whyItMatters":"The study provides practical techniques for author‑style transfer in scientific writing, revealing a trade‑off between style fidelity and fluency and demonstrating that hypernetworks can effectively balance these aspects."

By Leonard Popp, Danni Liu, Supriti Sinhamahapatra, Jan Niehues