Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,035 stories · RSS feed

arXiv Machine Learning
2d ago

Extending Music Annotation Schemas: Zero-Shot Prediction or Few-Shot Adaptation?

The paper investigates how to extend music annotation schemas when new attributes need to be added to commercial music catalogs. It compares zero‑shot prediction using audio‑language models, learning from pretrained representations, and supervised adaptation on existing annotations, using a benchmark built on the MGPHot dataset. The findings show that supervised adaptation outperforms zero‑shot prediction even with limited annotation budgets, while reusing frozen representations remains the best option for very modest budgets.

By Christos Plachouras, Emmanouil Benetos, Johan Pauwels
arXiv Machine Learning
2d ago

SkillFormer: Skill-Decomposed Adaptation for Audio Language Models

SkillFormer is a method for audio language models that decomposes audio understanding into skill‑specific low‑rank adapters and uses a learned router to activate the appropriate adapters at inference time. The router selects which adapters to engage based on the question, allowing different parameters to be used for tasks such as pitch comparison versus genre classification. An alternating training schedule updates each adapter on its own skill cluster before jointly calibrating the router, reducing gradient conflicts and adding fewer than 4% of the base model’s parameters. The approach improves average accuracy by 2.5 to 4.1 points across three distinct models on MMSU, MMAU‑Pro, and MMAR, achieving balanced gains across perception, reasoning, and semantic subcategories.

By Lee Seung-woo, Bowen Qi, Kim Min-jun, Jang Won-young
arXiv AI
2d ago

Principles that Guide, Actions that Inform: Agent Evolution via Knowledge Abstraction

The paper introduces SAGA, a framework that enables large language model agents to evolve by abstracting experiences into reusable principles, procedures, and episodic descriptions. SAGA transforms interaction trajectories into hierarchical knowledge with explicit applicability conditions, linking them back to execution evidence. Experiments on ScienceWorld and ALFWorld show that this execution–abstraction feedback loop improves task performance, and ablation studies confirm the importance of contextual instantiation and action regulation.

By Bowen Ye, Yongchao Xu, Junkai Ma, Xiang Yin, Wenzhao Li
arXiv AI
2d ago

Small Language Models for Smart Data Model Classification at the Edge: A Cost-Aware Hybrid Approach

The paper introduces a cost‑aware hybrid approach that uses small, open‑source language models to classify smart data models (SDMs) in Internet of Things (IoT) environments. It benchmarks general‑purpose, reasoning‑specialized, and code‑specialized models across domain‑specific datasets, highlighting their suitability for edge devices with limited resources. The study also compares these lightweight models to large language models and near‑zero‑cost baselines such as TF‑IDF and a lightweight sentence encoder to demonstrate practical performance gains.

By Cristian Martella, Angelo Martella, Antonella Longo, Motaz Saad
arXiv AI
2d ago

Sim-to-Real Transfer of Vision-Language Navigation in Continuous Environments Using an Ackermann-Steered Mobile Robot

The paper presents a vision‑language navigation system that transfers from simulation to a real Ackermann‑steered mobile robot without relying on navigation graphs or panoramic views. It uses a Cross‑Modal Attention architecture trained on simulated data and fine‑tuned with limited real‑world episodes, leveraging linear photometric adjustments and a camera‑LiDAR sensor suite. Evaluation with SPL and nDTW metrics shows robust, adaptable navigation in continuous environments.

By Chalindu Abeywansa, Sahan Gunasekara, Devindi De Silva, Seniru Dissanayake, Ranga Rodrigo, Peshala Jayasekara
arXiv AI
2d ago

MemMux: Runtime Verification and Honest Resource Attribution for Fleets of Parallel Coding Agents

MemMux is a local runtime designed to provide runtime verification and honest resource attribution for fleets of parallel coding agents. It emits observable signals that track per‑agent memory usage, ensure complete reclamation of terminated agents, detect escaped child processes, and keep the system from exceeding a bounded memory footprint. In benchmarks against tmux and a raw‑process baseline, MemMux keeps a fleet under a 7.5 GiB budget with zero swap, while ungoverned tools exceed the budget and spill into swap, and it achieves 100 % attribution with low overhead.

By Sumanyu Muku
arXiv AI
2d ago

Understanding and Mitigating Inference-Time Overreliance Using Agentic Memory

The paper investigates how large language model agents can over-rely on agentic memory, a phenomenon where retrieved memories distort inference even when they are correctly stored and retrieved. It shows that memory is helpful when past experience fully transfers to the current task but becomes misleading under partial query-memory overlap, a pattern confirmed by controlled experiments. To address this, the authors propose MEMTRIM, a plug‑and‑play framework that indexes memory evidence at write time and limits its reuse at read time, thereby reducing over-reliance without retraining and preserving useful memory benefits across models and memory architectures.

By Luoxi Tang, Yuqiao Meng, Nilesh Auradkar, Muchao Ye, Dazheng Zhang, Zhaohan Xi
arXiv AI
2d ago

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

Rationale-Guided Policy Optimization (RGPO) is a reinforcement‑learning framework that adaptively uses ground‑truth rationale information to scaffold a language model’s reasoning process. Instead of treating reference solutions as fixed imitation targets, RGPO temporarily incorporates rationales to help the model generate better responses, then reverts to unguided learning with higher‑reward, model‑generated solutions. Experiments in both language‑only and vision‑language tasks show that RGPO consistently outperforms RLVR baselines, with ablation studies confirming that adaptive rationale guidance is a key factor in its success.

By Hoang Phan, Minh Pham, Chau Pham, Chinmay Hegde, Trung Le, Qi Lei
arXiv AI
2d ago

Evaluating Escalation Signals for LLM Routing: Targets, Controls, and Five Ways to Fool Yourself

The paper investigates whether semantic entropy—a measure of disagreement among a language model’s sampled answers—can serve as a cheap signal for deciding when to route a query from a small to a larger language model. Experiments on GSM8K and other benchmarks show that semantic entropy can distinguish small‑model mistakes and improve routed accuracy, but the authors also reveal that a simple question‑difficulty rule can mimic its performance and that other factors (definition of success, benchmark design, live sampling cost) can undermine its effectiveness. They propose a checklist of checks to validate escalation signals and demonstrate how to predict when a cached‑outcome approach will fail. "whyItMatters":"The study highlights the importance of rigorous evaluation of escalation signals, showing that seemingly promising metrics can be misleading without proper controls and that practical routing decisions must account for cost and benchmark design."

By Ramin Pishehvar, Andrea Morandi, Mahesh Viswanathan
arXiv AI
2d ago

Defense-in-Depth for LLMs: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy

The paper presents a defense‑in‑depth framework for large language models that separates internal activation steering from external memory handling to combat sycophancy. It evaluates four open‑weight models on a new MemSyco‑Bench dataset, testing five memory‑defense configurations—including a Router Gate that selectively rewrites, keeps, or drops memories—and measures sycophancy and accuracy across 1,550 items. Results show that selective Router Gate filtering preserves more accuracy than complete memory removal, while inverse steering slightly reduces sycophancy but is not statistically significant.

By Ritvij Sharma, Russell Dlugosz, Ryan Zhou, Maheep Chaudhary
arXiv AI
2d ago

PsyCIDRA: A Dual-Agent Framework for Psychiatric Interviewing and Diagnostic Reasoning

PsyCIDRA is a dual‑agent framework that couples a free‑form psychiatric interviewer with a diagnostic reasoning agent to support expert review. The interviewer agent uses tools to keep working notes, load expert skills, and pull ICD‑11 references, while the diagnostic agent receives the interview transcript and generates hypotheses with supporting, conflicting, and missing evidence, withholding a final hypothesis if insufficient support exists. In simulations and a blinded human study, PsyCIDRA achieved higher diagnostic agreement and rank‑1 accuracy than direct prompting, indicating its promise for assisting psychiatric assessment through interactive dialogue.

By Milad Mohammadi, Fatemeh Akrami Shamsabadi, Zahra Mohseni, Amirhossein Safdarian, Malekfarhad Malek, Hadi Moradi, Hesham Faili
arXiv AI
2d ago

LOGIC: An LLM Benchmark for Intent-Grounded Change Impact in Aerospace Electrical Systems

LOGIC is a benchmark and evaluation framework that tests how language models can ground engineering requests in a deterministic inventory of candidate changes before propagating selected changes through an electrical traceability graph. The benchmark includes 168 scenarios—144 for selection and 24 for abstention—and evaluates three 7–8B models against intent‑agnostic, lexical, and structured‑evidence methods. Results show that structured evidence can achieve perfect candidate F1 on anchored cases, while large language models perform better on relational‑paraphrase cases; however, grounding accuracy drops as candidate inventories grow, and strict evidence gating reduces false positives but may also remove correct selections.

By Muhammad Faraz Shoaib, Muhammad Qasim, Raisulhaq Mohammed Rizwan, Rahmatullah Safdar, Muzammil Adnan Shaik, Abdul Aleem Mohammed
arXiv AI
2d ago

Personal-Agent Mediated Recommendation with Cross-Platform User History

The paper introduces Personal-Agent Mediated Recommendation, a new paradigm where a personal LLM agent uses cross‑platform user history to adjust a platform’s recommendation ranking. It presents MediateRec, a benchmark for evaluating this mediation, and proposes Personal Attribution Mediation Optimization (PAMO) to balance beneficial rescues against harmful overrides. Experiments show that PAMO improves over outcome‑only reinforcement learning, achieving a better rescue‑harm trade‑off on both synthetic and real cross‑platform tests.

By Yu Xia, Jiangfan Zhang, Jun Xiao, Julian McAuley, Xiangjun Fan
arXiv AI
2d ago

LSC-DPO: Learning-Signal-Controlled Direct Preference Optimization

The paper introduces LSC-DPO, a variant of Direct Preference Optimization that dynamically controls the learning signal to maintain sensitivity during training. By analyzing the logistic DPO loss geometrically, the authors identify the sigmoid factor as a key learning signal and propose a log‑space framework for stable target‑regime tracking. Experiments on AlpacaEval 2, MT‑Bench, and Anthropic‑HH demonstrate that LSC‑DPO outperforms standard DPO and other preference‑optimization baselines, and a signal‑budget compensation rule further reduces variability across different coefficient initializations.

By Yang Qu, Yusheng Han, Chengjia Feng, Handan Liu
arXiv AI
2d ago

Where Rules End and Judges Begin: Measuring the Judgment Boundary in Multi-Agent Systems Security

The paper introduces DEFER1, a deterministic-first enforcement system for large‑language‑model based multi‑agent systems that uses 28 checks to block most attacks and refers only a small fraction to human judges. In tests across four domains, DEFER1 reduces attack success from about 30% to roughly 3%, with 78% of attacks blocked deterministically and only a quarter reaching the judges. The study highlights that rules effectively handle clear policy violations while judges address ambiguous intent, but also reveals weaknesses such as a risk‑score gate that misclassifies many proposals.

By Shaswata Mitra, Raj Patel, Subash Neupane, Sudip Mittal, Md Rayhanur Rahman, Shahram Rahimi
arXiv AI
2d ago

Massive Activation Gating Channel in Large Language Models

The paper identifies a single input embedding channel, called the massive activation gating channel (MAGC), that controls the emergence of massive activations in large language models. When the MAGC value is sufficiently large or small, the spike feed‑forward network outputs exhibit exceptionally large magnitudes. The authors verify MAGC across six models and provide a theoretical explanation linking the channel to a quadratic form that mixes specific columns of the down‑projection matrix, which produce massive activations.

By Minjia Mao, Shi Chen, Bowen Yin, Xiao Fang
arXiv AI
2d ago

ST-Bench: A Spatial-Temporal Benchmark for Multi-Agent System Generation on Scientific Research Tasks

ST-Bench is a new benchmark that tests whether multi‑agent systems (MAS) outperform single agents on complex scientific data analysis tasks. It includes 100 Earth‑science data‑science tasks expanded into 2,067 queries, validated by domain experts. Evaluations show that most MAS configurations beat the cheapest single‑agent baseline, with the best reaching nearly three times its score, though at higher inference cost.

By Qi Cheng, Rongchao Dong, Shengyu Chen, Licheng Liu, Dan Lu, Zhengzhang Chen, Wei Cheng, Yiqun Xie, Haifeng Chen, Xiaowei Jia, Haoyu Wang
arXiv AI
2d ago

OTel: Open Telco AI Datasets, Benchmarks, and Models

OTel is an open telecom AI resource that provides derived datasets for retrieval, reranking, instruction tuning, and safety/abstention, along with 30 full‑parameter post‑trained baselines covering 10 embedding models, 3 rerankers, and 17 language models. The project has seen significant community engagement, with over 16 million model downloads and more than 157 media mentions by May 2026. Post‑training on OTel data improves performance across all model families, achieving 93.1% NDCG@10 for embeddings, 0.947 MRR@10 for rerankers, and 87.8% correctness for language models.

By Farbod Tavakkoli, Gregory Diamos, Kenneth Church, David Kanter, Mark Austin, Imtiaz Karim, Mirza Masfiqur Rahman, Merouane Abdelkader Debbah, Zeinab Nezami, Ali Maatouk, Leandros Tassiulas, Rex Ying, Nick Sorros, Louis Powell, Nikolaos Vasiloglou, Ashish Vaswani, Somanshu Singla, Adarsh Chaluvaraju