Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,035 stories · RSS feed

arXiv AI
4d ago

SoftGene: Protein Language Model-Enhanced Soft Prompting for Interpretable Gene Set Annotation

SoftGene is a new framework that enhances gene set annotation by combining protein language model representations with a hybrid prompting scheme. It uses a hierarchical attention-based encoder built on ESM to encode gene sets from protein amino acid sequences, then merges soft prompts derived from these embeddings with hard prompts generated by a large language model. The approach is evaluated on Gene Ontology and MSigDB datasets, showing that integrating protein-sequence information with textual context improves overall annotation performance, though the benefit varies across biological domains.

By Drew Ross, Arya Hadizadeh Moghaddam, Dongjie Wang, Xiaoyu Zhang, Zijun Yao
arXiv Machine Learning
4d ago

Fisher-Guided Submodular Data Selection for Continual Pre-Training of Large Language Models

The paper introduces a Fisher-guided submodular data selection method for continual pre‑training of large language models, addressing the forgetting problem that arises when a target‑domain corpus overwrites pretrained knowledge. By decomposing candidate gradients into anchor and frontier components and optimizing a log‑determinant submodular objective, the method selects data that both improves target‑domain performance and limits forgetting. Experiments on TinyLlama‑1.1B and Llama‑3.1‑8B show that the selector achieves superior adaptation and forgetting control while using ten times fewer tokens than traditional replay strategies.

By Zhenghao Zhao, Gaowen Liu, Zhiling Lan, Yan Yan
arXiv AI
4d ago

The AI Risk Observatory: What Can We Learn from AI Disclosures in Annual Reports About Societal Resilience?

The study examines whether annual reports can reveal how companies disclose AI-related risks and responses. Using a two-stage classification pipeline on 9,821 reports from 1,362 UK listed firms (2020‑2026), the authors find that mentions of AI risk rose from 2.8% to 41.2% and AI adoption disclosures from 13.8% to 45.2%, with most risk mentions clustering around major vendors like Microsoft. Disclosure varies by sector and market segment, with Critical National Infrastructure and AIM reports lagging, and substantive risk disclosures remain rare—only 4.3% in 2025. "Why It Matters": The findings show that while AI risk is increasingly referenced in corporate reports, substantive disclosures are scarce, highlighting a gap in transparency that could affect societal resilience.

By Bart Jaworski
arXiv AI
4d ago

Toward SLM-based agentic task-tool intent matching

The paper proposes using Small Language Models (SLMs) as a task‑tool relevance classifier to verify each tool call made by AI agents. By evaluating every selected tool against the assigned task, the SLM provides a relevance signal that can be used for downstream enforcement. The authors introduce a novel dataset of multi‑tool tasks across distinct Model Context Protocol servers and explore prompt‑optimization, supervised fine‑tuning, and reinforcement learning (GRPO) to optimize and specialize the SLMs.

By Chiara Troiani, Arash Salarian, Majed El Helou, Benjamin Ryder, Jean Diaconu, Herv\'e Muyal, Marcelo Yannuzzi
arXiv Computer Vision
4d ago

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

The paper introduces Spatial Memory Intelligence (SMI), a framework that enhances long‑video world models by systematically managing spatial memory using an understanding model. SMI employs four coordinated operations—spatial clustering, within‑cluster sparsification, action‑aware retrieval, and reliability‑aware filtering—to handle increasingly complex and lengthy memory sequences. Experiments across various baselines and benchmarks show that SMI improves memory sparsity, generation stability, and spatial consistency, demonstrating its effectiveness and generalizability.

By Ying Yang, Guiyu Zhang, Lianghua Huang, Chang Nie, Chenyang Si, Haofan Wang, Shaoshuai Shi, Li Jiang
arXiv Machine Learning
4d ago

Context-Tower Conversion Preserves Generation While Freezing Retains Knowledge: Low-Budget AR-to-Diffusion Conversion of MoE LLMs

The paper compares two low‑budget methods for converting a 30B Mixture‑of‑Experts autoregressive language model into a diffusion language model. One method updates a subset of the model’s weights in‑place, while the other freezes the context tower and conditions on a frozen causal copy via cross‑attention. With only 1B training tokens, the frozen‑tower approach achieves a HumanEval pass@10 score of 71.60 versus 6.19 for the in‑place method, and retains 95% of the parent’s GSM8K and 99% of its MMLU‑Pro performance.

By Wentao Lu, Jesse Clark, Tianyu Zhu
arXiv AI
4d ago

LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios

LPS-Bench is a benchmark designed to evaluate the safety awareness of computer‑use agents (CUAs) in long‑horizon planning tasks that involve tool workflows. It uses a template‑guided multi‑agent pipeline to generate user instructions, simulated toolkits, and case‑specific safety criteria, followed by human review, allowing scalable expansion without building separate application environments. The benchmark includes 570 cases from 65 scenarios across seven task domains and nine planning‑risk types, and an LLM‑based evaluator assesses tool choices, arguments, and responses throughout execution. Evaluations of 13 LLM agents show persistent safety failures in both benign and adversarial settings, with prompt‑based interventions providing only model‑dependent improvements.

By Tianyu Chen, Chujia Hu, Dongrui Liu, Xia Hu, Wenjie Wang
arXiv AI
4d ago

SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models

SyntaxBench is a diagnostic benchmark and statistical evaluation framework for character‑level reasoning in large language models, comprising five core tasks—character counting, letter containment, palindrome detection, edit distance, and longest‑string selection—and a harder substring‑extraction stress test called index_to_span. The benchmark uses paired English and random‑string inputs, zero‑, one‑, and four‑shot prompts, and evaluates models from 2B to 32B parameters across multiple reasoning modes. It reports a wide range of metrics, including exact‑match and relaxed accuracy, Cohen’s kappa, McNemar tests, bootstrap confidence intervals, Kendall’s tau, class‑conditional metrics, tokenization analysis, and multiple‑comparison‑corrected tests.

By Mohsen Larni (Department of Computer Science, University of Nevada, Las Vegas), Sobhan Ebrahimi Azar (Department of Computer Science, University of Nevada, Las Vegas), Pouyan Nahed (Department of Computer Science, University of Nevada, Las Vegas), Kazem Taghva (Department of Computer Science, University of Nevada, Las Vegas)
arXiv AI
4d ago

AMBER: Multi-View Adaptive Budget Allocation for Listwise Vision-Language Reranking

AMBER is a new online, budgeted multi‑view reranking framework for vision‑language models that dynamically allocates computation to maximize information gain. It treats fragmented listwise VLM outputs as local tournaments and uses continuous Elo updates to maintain a lightweight global ranking state. Experiments on CIRR, CIRCO, and PhotoBench show that AMBER outperforms other multi‑call VLM reranking methods under comparable budgets, and remains effective even with lower budgets.

By Wenteng Chen, Jiachen Zhu, Rong Shan, Tianyi Xu, Yuxiang Chen, Congmin Zheng, Teng Wang, Junjie Wu, Weiwen Liu, Changwang Zhang, Weinan Zhang, Jun Wang, Jianghao Lin
arXiv AI
4d ago

LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling

The paper proposes a scalable traffic modeling approach that uses a single representative large language model (LLM) agent for each homogeneous traveler group, rather than one LLM per traveler. The representative agent maintains a mixed strategy over routes, updates it daily based on positive reinforcement signals, and uses a tunable step size to adjust its strategy. This design improves scalability, stabilizes learning, and produces interpretable dynamics that reproduce realistic behavioral patterns such as the decoy effect and income‑based willingness‑to‑pay differences.

By Hanlin Sun, Jiayang Li
arXiv AI
4d ago

Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation

Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation proposes a method to formally certify that disabling a specific circuit in a neural network removes a targeted skill while preserving another, for every input within a continuous region. The approach extends beyond testing by providing sound guarantees over continuous embedding-space regions, demonstrated on toy ReLU networks and a standard transformer. It also shows that no finite deterministic black-box test can guarantee removal, highlighting the necessity of formal verification.

By Md Sazid Uddin, Md. Khairul Alam Mazumder, M. F. Mridha
arXiv AI
4d ago

Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models

The paper introduces COMETH, a framework that learns interpretable moral contexts from human judgments using probabilistic clustering and large language models. It builds a dataset of 300 scenarios involving six core actions and three moral rules, and employs an LLM-based preprocessing pipeline to cluster actions and scenarios. COMETH achieves roughly double the alignment with human judgments compared to end‑to‑end LLM prompting, while providing transparent explanations of the contextual features driving its predictions.

By Geoffroy Morlat, Marceau Nahon, Augustin Chartouny, Raja Chatila, Ismael T. Freire, Mehdi Khamassi
arXiv AI
4d ago

Exploring Subnetwork Interactions in Heterogeneous Brain Network via Prior-Informed Graph Learning

The paper introduces KD-Brain, a Prior‑Informed Graph Learning framework that incorporates semantic and clinical priors to model interactions among functional subnetworks in brain networks. It employs a Semantic‑Conditioned Interaction mechanism to guide attention queries by subnetwork identities and a Pathology‑Consistent Constraint to align learned interactions with clinical priors. KD‑Brain achieves state‑of‑the‑art performance on disorder diagnosis tasks and identifies biomarkers that align with psychiatric pathophysiology.

By Siyu Liu, Guangqi Wen, Peng Cao, Jinzhu Yang, Xiaoli Liu, Fei Wang, Osmar R. Zaiane
arXiv AI
4d ago

Many Preferences, Few Policies: Compact Portfolios for Multi-Objective LLM Alignment

The paper introduces PALM, an algorithm that constructs a compact portfolio of aligned large language models (LLMs) capable of near‑optimal performance across a wide range of reward weightings for objectives like helpfulness, harmlessness, and conciseness. PALM uses a structured grid of weight vectors, lazy search, and pruning to efficiently identify a small set of policies that provably cover all target reward configurations within specified tolerances. Experiments demonstrate that PALM achieves smaller approximation gaps than portfolios built from uniformly spaced or randomly sampled weights and scales effectively to higher‑dimensional reward spaces.

By Cheol Woo Kim, Jai Moondra, Roozbeh Nahavandi, Andrew Perrault, Milind Tambe, Swati Gupta
arXiv AI
4d ago

DNAlign: Dynamic Null-Space Safe Alignment for LLMs

DNAlign is a lightweight alignment framework that uses control‑theoretic optimization and null‑space projection to steer large language models toward safe behavior while preserving core knowledge and response quality. By treating the LLM as a dynamic system, it introduces controllable perturbations that are restricted to a harmful‑related subspace derived from neutral hidden states, and a value function trained on human preference data adaptively optimizes these control signals. Extensive evaluations across multiple LLM backbones show that DNAlign consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility, outperforming prior alignment baselines without sacrificing generation diversity.

By Jisheng Dang, Yushuo Zhao, Dewei Liu, Junfeng Fang, Bimei Wang, Tiantian Rao, Hong Peng, Bin Hu, Tat-Seng Chua
arXiv AI
4d ago

Harness-Aware Distillation for Small Language Model Agents

The paper introduces Harness-Aware Distillation (HAD), a method for training smaller language model agents that preserves the surrounding harness—software managing context, tools, and feedback—while focusing distillation on the teacher’s contributions beyond the harness. HAD combines an action preference that contrasts teacher actions with and without harness information, and a validity check that filters out contradictory preference pairs. Experiments on long-horizon agent benchmarks show that HAD outperforms standard on‑policy distillation, reducing unproductive loops and improving error recovery without requiring task rewards or future information.

By Moonseok Choi, Taehong Moon, Giung Nam, Juho Lee
arXiv Computation and Language
4d ago

Enhancing Biomedical Named Entity Recognition via Multiple Programming Languages Instruction Tuning and Ensemble Method

The paper introduces MITE, a method that transforms biomedical named entity recognition (BioNER) into a structure‑to‑structure generation task by encoding instructions and outputs in multiple programming languages (Python, C++, Java). This approach provides structurally diverse supervision without extra biomedical knowledge, and during inference it aggregates predictions via entity‑level voting to reduce language‑specific variance. Experiments on six BioNER datasets show that MITE outperforms BERT‑based and LLM‑based baselines and generalizes well across datasets.

By Songtao Li, Yijia Zhang, Jianyuan Yuan, Shidi Zhang, Fengyu Zhang, Hongfei Lin
arXiv AI
4d ago

Assessing Rule Adherence of LLM Adjudicators in Call of Cthulhu TRPG

The paper introduces CoC‑Seduce, a multi‑agent adversarial benchmark for evaluating how well large language models (LLMs) adhere to the rules of the tabletop role‑playing game Call of Cthulhu when acting as adjudicators. Using three LLMs to generate 5,376 samples across diverse settings and skill categories, the study benchmarks 22 target adjudicators and finds that newer releases and explicit reasoning do not guarantee robustness, with pseudo‑logic framing emerging as the most effective manipulation. The benchmark highlights the vulnerability of LLM adjudicators to rhetorical injection attacks that exploit narrative framing to bypass rule enforcement.

By Weiying Chen, Junlong Shen, Zhanyuan Guo, Xiaoou Zhou
arXiv Computation and Language
4d ago

Silent Dissent: LLM Agents That Yield to the Majority Still Represent Their Original Premise

The paper investigates whether large language model (LLM) agents that yield to a unanimous majority in multi‑agent debate truly change their underlying premise or merely adjust their statement. Using two‑hop factual questions with an unstated intermediate entity, the authors show that agents that concede still encode the original bridge in their internal representations, as revealed by a Jacobian‑lens analysis, even when the majority answer is wrong. Experiments across several open‑weight models demonstrate that hiding or removing the agent’s earlier answer increases conformity and that the original premise can be recovered from the question alone.

By Ziang Ni, Peng Zou