Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

26,449 stories · RSS feed

arXiv Computation and Language
5d ago

Understanding Why Language Models Hallucinate: Testing Reasoning Against Priors

The paper investigates why large language models hallucinate by proposing that failures often stem from inference misalignment rather than missing knowledge. It introduces a latent key-task model that shows pretraining-frequency imbalance can cause shortcut inference paths to dominate, leading to hallucinations. The authors create TrapQA, a diagnostic testbed with ScientistQA and Real-Life Constrained QA, to demonstrate two failure modes—task-retrieval bias and key-selection bias—where biased latent inference produces hallucinated answers.

By Yangfan Hu, Xuhan Tong, Haoyue Bai, Xi Ding, Shashank Muralidhar Bharadwaj, Siyang Cao, Robert Nowak, Jiawei Zhang
arXiv AI
5d ago

Learning to Assess Heartbeat Observability for mmWave Heart-Rate Sensing

The paper presents HEAR, a dual‑task Transformer that learns to assess heartbeat observability from mmWave radar phase spectra and simultaneously estimates heart rate. Using a controllable FMCW simulator, the authors generate labeled data where observability is defined by the agreement between the dominant heartbeat‑band peak and the known heart rate. Trained only on simulated data, HEAR transfers zero‑shot to real‑world datasets at 60 and 120 GHz, enabling selective heart‑rate estimation that dramatically reduces error and achieves sub‑50 ms latency on edge hardware.

By Yuxuan Hu, Shilin Shan, Jianfei Yang, Feng Xu
arXiv Computation and Language
5d ago

Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality

Memristor-based analog compute-in-memory architectures promise efficient deployment of Large Language Models, yet their intrinsic non-idealities degrade reasoning performance across benchmarks. The study evaluates how these non-idealities affect LLM reasoning and tests three training-free mitigation strategies: thinking mode, in-context learning, and module redundancy. Findings show shallow layer redundancy boosts robustness, thinking mode helps only at low noise, and in-context learning shortens output with a modest performance cost.

By Taiqiang Wu, Yuxin Cheng, Chenchen Ding, Runming Yang, Xincheng Feng, Wenyong Zhou, Zhengwu Liu, Ngai Wong
arXiv Machine Learning
5d ago

What Is Lost in Post-Training? Default Collapse and the Loss of In-Context Steerability Across Diverse Perspectives

The paper investigates how post‑training fine‑tuning of large language models can reduce their ability to adapt to in‑context information, particularly when the models are fine‑tuned toward one side of cultural‑value disagreements. Experiments show that as a model is trained to favor a specific perspective, its capacity to recognize and enact opposing viewpoints diminishes over time. The authors propose an alternative objective that balances reward maximization with a prescribed distribution over expressed perspectives, offering a practical stance‑distribution matching implementation.

By Jessica Dierking, Itai Shapira, Niclas Boehmer
arXiv Computation and Language
5d ago

HARPO: Hallucination-Aware Reinforcement Learning for Faithful and Creative Language Generation

HARPO is a reinforcement learning framework that jointly optimizes faithfulness and creativity in language generation. It uses a Hallucination-Aware Generative Reward Model (HA‑GRM) to evaluate both faithfulness and writing quality, and a Selective Activation Mechanism (SAM) that applies writing rewards only to hallucination‑free outputs. Experiments on Qwen models show that HARPO improves faithfulness scores and reduces hallucination rates while boosting creative‑writing performance.

By Tiezheng Yu, Yuxin Jiang, Jinpeng Li, Shuning Sun, Fei Mi, Haoli Bai, Lifeng Shang
arXiv Machine Learning
5d ago

AIGS: Adaptive Incremental Gating System for Online Representation Learning in Non-Stationary Data Streams

The paper introduces AIGS, a lightweight Adaptive Incremental Gating System designed for online representation learning in non‑stationary data streams. AIGS uses a Shock Ratio feedback signal to drive a Continuous Plasticity Controller, enabling smooth adaptation between learning plasticity and memory retention while keeping per‑step complexity linear in the feature dimension. Experiments on smart‑city traffic, meteorological, and industrial datasets show that AIGS improves early‑warning lead times, accelerates recovery after abrupt changes, and enhances anomaly recall in noisy environments.

By SiRui He, Kai Liang Lew, Chui Zi Ong, Chean Khim Toa
arXiv Machine Learning
5d ago

Effects of interpulse-interval variation on deep-learning classification of bat vocalizations

The study examined how variation in the interpulse interval (IPI) of bat vocalizations affects deep‑learning classification. Two datasets—one preserving natural IPI timing and another normalizing call spacing to 50 ms—were used to fine‑tune EfficientNet‑B0 and PaSST models. Results showed that IPI normalization had mixed effects: PaSST performance remained stable while EfficientNet improved, yet models trained on natural IPI data performed better on natural test sets, indicating limited but non‑negligible influence of natural IPI variation.

By Welmoed R. Eversteijn, Burooj Ghani, A. Leonie Baier, Dan Stowell
arXiv AI
5d ago

Reliable Self-Evolution with Imperfect Proxy Rewards

The paper introduces Conformal Interval-Driven Self-Evolution (CISE), a method that uses conditional conformal inference and online density-ratio estimation to create candidate‑specific reward intervals for self‑evolving search in materials science. CISE applies conservative interval‑based rewards, ensuring that only candidates whose required property intervals lie entirely within feasible regions are returned. Experiments on three self‑evolving search tasks show that all candidates returned by CISE are true positives under high‑fidelity evaluation, whereas baseline methods produce more candidates but include false positives.

By Kangjun Noh, Soyu Kim, Kyungwoo Song
arXiv AI
5d ago

PLCWorld: Benchmarking LLM-Generated PLC Programs in Closed-Loop Plant Simulation

PLCWorld is a closed‑loop execution environment and benchmark that evaluates large language model (LLM)–generated programmable logic controller (PLC) programs by coupling Structured Text (ST) execution with simulated plant responses and sensor feedback. It includes 100 synthetic tasks and 473 task‑condition pairs across Motion Control and Material Handling, with difficulty levels based on control‑dependency scope. The benchmark reports Task Success and Safety Violation separately and validates results through practitioner review, reference programs, counterexamples, and comparisons with independent ST runtimes, revealing performance differences among six LLMs and four generation‑and‑verification workflows.

By Yunji Kim, Yunseok Lee, Hyunwoo Seo, Jaerim Choi, Woojin Lee
arXiv Computation and Language
5d ago

SEDIMA: Cross-Run Hierarchical Insight Memory for Evolutionary Search Agents

SEDIMA is a persistent hierarchical insight memory designed for evolutionary search agents that use large language models. It transforms raw search traces into natural‑language insights, clusters them by semantic similarity, and retrieves relevant guidance to inform future mutations, thereby accumulating transferable knowledge across runs and problems. When added as a drop‑in module, SEDIMA improves average final performance by 5.5% on AlgoTune and 6.6% on ALE‑Bench LITE, and reduces the number of iterations needed to reach baseline‑best performance by 32.3% on OpenEvolve.

By Amirhossein Abaskohi, Mahdi Mostajabdaveh, Zirui Zhou
arXiv Computation and Language
5d ago

CUEing User Simulators: Calibrated User Embeddings for Multi-Turn Benchmarking

The paper introduces Calibrated User Embeddings (CUE), a framework that encodes real user sessions into continuous representations and decodes them into persona commands to steer large language models (LLMs) as user simulators without additional training. CUE enables user-conditioned replay of past interactions and generates novel personas that better align with real-user success rates and failure patterns. Evaluations on the τ²‑Bench dataset show that CUE‑driven simulators commit fewer simulator‑attributed errors, more accurately reproduce real‑user failure modes, and maintain competitive user fidelity across various tasks and LLMs.

By Anjali Kantharuban, Jonas Mueller
arXiv Computation and Language
5d ago

Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text

The paper investigates why supervised AI‑text detectors, specifically a RoBERTa‑based model, can be fooled by subtle changes in language. By applying semantic, structural, and tokenizer‑level perturbations to a large dataset and controlled Mistral‑7B‑Instruct outputs, the authors show that increasing verb diversity makes machine text easier to detect and that detection scores correlate with statistical complexity, leading to a high false‑positive rate on formal human writing. They also evaluate an event‑based latent space detector, finding that paraphrasing and homoglyphs significantly alter extracted event sequences and verbs, yet the detector’s performance remains modest (AUC 0.577).

By Claudiu Creanga, Liviu Dinu
arXiv Machine Learning
5d ago

Predicting and Repairing Merge Collapse in Large Language Models

The paper introduces a method for predicting and preventing merge collapse when combining large language models fine‑tuned from a shared base. By measuring the variance of specialists’ task vectors—termed interference—a pre‑merge score identifies destructive merges. The authors present PRISM, an operator that averages task vectors and then soft‑thresholds each layer based on interference, which successfully preserves model performance without additional data or tuning.

By Jungseob Lee, Seungyoon Lee, Sugyeong Eo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim
arXiv AI
5d ago

GISTBench: Evaluating LLM User Understanding via Evidence-Based Interest Verification

GISTBench is a benchmark designed to assess how well Large Language Models understand users by extracting and verifying user interests from their interaction histories in recommendation systems. It introduces two new metric families—Interest Groundedness (IG) and Interest Specificity (IS)—to measure the accuracy and distinctiveness of LLM-generated user profiles. The benchmark includes a synthetic dataset built from real user interactions on a global short‑form video platform, validated through user surveys, and evaluates a range of open‑weight and proprietary LLMs, uncovering limitations in their ability to count and attribute engagement signals.

By Iordanis Fostiropoulos, Muhammad Rafay Azhar, Abdalaziz Sawwan, Boyu Fang, Yuchen Liu, Jiayi Liu, Hanchao Yu, Qi Guo, Jianyu Wang, Fei Liu, Xiangjun Fan
arXiv AI
5d ago

What Should World Models Forget? Stratified Retention for Continual Adaptation

The paper argues that continual learning for world models must differentiate between knowledge that should never be revised—such as physics and object permanence—and knowledge that should be updated when the environment changes. It critiques existing forgetting metrics and benchmarks for failing to capture this distinction, and proposes a differential retention approach that tracks invariant regression testing and revision latency throughout the adaptation process.

By Nishit Anand, Ramani Duraiswami, Dinesh Manocha
arXiv AI
5d ago

MaskCoFT: Masked Co-Adaptive Fine-Tuning for Memory-Efficient MoE Inference

MaskCoFT introduces a masked co‑adaptive fine‑tuning approach for mixture‑of‑experts language models, training both routers and experts jointly with a cross‑entropy loss. A learnable binary mask limits each layer’s Top‑K routing to a subset of experts during fine‑tuning, allowing experts to adapt to the tokens they receive. In simulated GPU cache scenarios, MaskCoFT reduces expert fetches per token by 23.7% for Mixtral‑8x7B and 10.1% for DeepSeek‑V2‑Lite, and lowers inference time per output token by up to 16.4% and 5.5% respectively, while maintaining or improving accuracy across nine benchmarks.

By Junfeng Wu, Zehao Fan, Hadjer Benmeziane, Kaoutar El Maghraoui, Liu Liu, Yinan Wang
arXiv Machine Learning
5d ago

To Explore The Strange New World Beyond Data Distribution: System Behavior, Causality Tax, and Non-causal Base Model

The paper argues that the causality of language models may be unnecessary or suboptimal when system behavior—extra dominant factors beyond data distribution—is treated as a first‑principle Bayesian feature. It introduces the SBD framework, incorporating system behavior into the evidence lower bound, and demonstrates a counter‑intuitive Causality Tax where ignoring these factors leads to structural error. Using a non‑causal variational family called Green Shell, the authors show through theoretical bounds, implicit measurements, and Neural Tangent Kernel analysis that this approach yields tighter error bounds and improved generalization compared to causal models.

By Xianzhi Zeng, Jiangneng Li, Gao Cong
arXiv AI
5d ago

Automating the Application of HCI Principles: Skills for On-Demand UI Construction, the Human-AI Space to Think, and the Future of HCI

The article discusses how large language models can now generate user interfaces from natural‑language descriptions, and proposes a framework that turns HCI design knowledge into machine‑readable ‘skills’ that guide on‑demand UI construction. By treating the dialogue between user and AI as a shared cognitive workspace, these skills encode established HCI principles (e.g., Nielsen’s heuristics, WCAG criteria) as editable, version‑controlled artifacts. The authors outline a research agenda that envisions HCI moving from heuristic checklists to a machine‑executable, open‑source craft.

By Nathan Conklin, Miranda Capra, Chris North
arXiv AI
5d ago

VERSE: Verified Self-Evolving Optimizer for Agent Harnesses

The paper introduces VERSE, a Verified Self‑Evolving optimizer that enhances LLM agent harnesses by allowing the optimizer to test edits, replay failures, and perturb steps while tracking fixes and regressions. VERSE builds its own tools for failure analysis, verification, training audits, and workflow control, and uses this feedback to revise the harness’s prompts, skills, tools, hooks, and notes without changing model weights. In experiments across five executors and multiple languages, VERSE improves all evaluated harness optimizers, achieving higher accuracy on held‑out and out‑of‑distribution tasks compared to the strongest baselines.

By Zekai Wang, Yingqiang Ge, Zekun Wang, Hai Wang, Yuhui Xu, Joshua Frandsen, Shancong Fu, Ashia C. Wilson, Chandan K. Reddy
arXiv Machine Learning
5d ago

Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies

The paper examines how transformer policies, which process agents as ordered token sequences, perform in multi‑agent robot learning where agent teams are unordered. It finds that low permutation error can mask action collapse, where all agents choose the same action, and proposes additional diagnostics such as action diversity and same‑action fraction. Experiments show that while a PPO‑ID baseline avoids collapse, it remains order‑sensitive, and that strong equivariance regularization can still cause homogeneous behavior; a weak penalty improves robustness and preserves diversity for three‑agent teams, but four‑agent teams need much smaller regularization weights.

By Amit Thakur, Mukesh Singhal