The paper introduces threat‑preserving representation sensitivity (TPRS) to assess how changes in agent‑visible wording affect attack success rates (ASR) while keeping the underlying task and evaluation criteria constant. Experiments on Agent Security Bench, MCPTox, and AgentDojo show that renaming threat‑related tools can significantly alter ASR—up to +13.21 percentage points on some models—yet may also degrade benign utility. The findings suggest that security scores based on a single representation may not generalize, urging robustness claims to be validated across multiple threat‑preserving representations.
By Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, Dinghao Wu
OLMo-Detect is a new benchmark for membership inference attacks (MIAs) on large language models, covering pre‑training, mid‑training, and post‑training stages. It aligns member and non‑member samples on key axes and rigorously filters non‑members using infini‑gram, also offering a shifted variant to test distribution robustness. Evaluation of 15 unsupervised and 3 supervised MIAs on the OLMo 2 family shows limited overall performance (best AUC 0.68), peak detection at mid‑training driven by data type, and sensitivity to distribution shifts, with findings generalizing to OLMo 3 and other models.
By Tao Shi, Chaoyi Xiang, Qiongkai Xu, Jey Han Lau
Gaze Attention is a new mechanism for multimodal large language models that selectively focuses on relevant visual regions during each generation step, rather than attending to all visual tokens. By grouping tokens into spatial regions and using learnable context tokens to retain global information, it reduces attention computation and visual key‑value entries by up to 90%. Experiments on 13 image and 6 video benchmarks show that Gaze Attention matches or outperforms dense‑attention baselines while using fewer visual resources.
By Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun
The paper investigates how well large language models (LLMs) can serve as judges in summarization evaluation by applying psychometric techniques. Using Many‑Facet Rasch Models, the authors decompose human and LLM ratings into latent summary quality, rater severity, dimension severity, and rating‑scale thresholds, and introduce a residual hardness metric to capture judging difficulty. Their analysis of 17 open‑weight LLM judges on the SummEval dataset reveals that moderate alignment in latent quality does not translate to alignment in residual hardness; humans and LLMs differ in which summary–dimension units are hard, with LLMs tending to find consistency hard and humans tending to find coherence hard, and some hard cases can be predicted from observable source–summary properties.
By Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne
The paper investigates whether adding a short classification instruction after a user’s prompt improves the ability of activation probes to detect malicious inputs in large language models. Across 13 safety benchmarks and three open‑weight model families, a classification suffix consistently boosts out‑of‑distribution detection (up to ~4 AUC points) compared to no suffix, and the benefit transfers to multi‑position pooling probes used in production. The improvement stems from the classification format itself rather than the specific content of the instruction, though the optimal suffix varies with the model and readout type.
By Elad David, Max Fomin
NeutronGym is the first executable environment that lets language‑model agents design neutron instruments, using tools that validate their builds, McStas ray‑tracing for simulation, and a level‑resolved grading ladder that evaluates syntax, runtime, structure, and science without an LLM judge. The platform provides procedural families of instrument layouts with held‑out parameter regimes and a curated set of 16 tasks from published instruments, called McStasBench, which includes memorization probes and a sandbox. Experiments show that while seven models can reproduce at most seven of the 16 tasks and none reaches a reference design, reinforcement learning can dramatically improve performance—e.g., Qwen3‑8B’s success rate jumps from 11% to 77% on held‑out instances—highlighting the ladder’s importance for partial credit and the potential of RL to match classical optimizers under realistic simulation budgets.
By Lijie Ding, Changwoo Do
RT‑SFT is a method for text style transfer that uses roundtrip translation through a pivot language to strip stylistic information from monolingual corpora, creating pseudo‑parallel data. This data is then used to LoRA‑finetune an instruction‑tuned large language model as a stylizer, allowing the model to rewrite sentences in a target style while preserving meaning. Experiments across four style domains show that RT‑SFT surpasses state‑of‑the‑art approaches, including few‑shot in‑context learning, and offers effective retrieval augmentation for expert style domains with strict terminology.
By Ruoxi Liu, Philipp Koehn
ReFract is a new benchmark that tests whether language model agents can act appropriately based on a user’s role, a capability called Perspective Awareness. It contains 150 expert‑validated entries derived from anonymized domain support queries, and uses Text World Models to simulate the agents’ operating environments and generate perspective‑aware action trajectories. Current state‑of‑the‑art LLMs solve at most 69% of the tasks, with over half of their trajectories attempting perspective‑violating actions.
By Hainiu Xu, V\'{i}tor N. Louren\c{c}o, Mohnish Dubey, Yunfei Bai, Yulan He, Caroline Catmur, Aline Paes, Marco Caserta, Akash Chandrayan, Luca D'Angelo
The paper investigates why certain synthetic relational datasets lead to better relational foundation models. By training four Relational Transformer checkpoints on data from four different generators, the authors trace a measurable property of the data—specifically the predictive necessity of cross‑table information—to the emergence of relational computation in the models. They find that the RelDiff generator yields the largest predictive gain from foreign‑key‑linked parents, and its model uniquely responds to foreign‑key interventions, a dependence that persists across random initializations and grows with corrupted links. Disrupting this mechanism during downstream inference eliminates RelDiff’s advantage on relational tasks while leaving structure‑insensitive models largely unchanged, thereby linking synthetic data properties to learned mechanisms and downstream behavior.
By Shivam Dubey, Mohamed Bouadi, Nassim Bouarour, Varun Kulkarni, Aditya Tanna, Vinay Kumar Sankarapu
The paper introduces THPL, a vision-to-language framework for precision feeding of rainbow trout in recirculating aquaculture systems. It uses Fishsort to extract feeding trajectories, a Hierarchical Behavior Encoder to model individual and collective dynamics, and integrates these with environmental data and expert rules to fine‑tune a large language model via LoRA and counterfactual multimodal Direct Preference Optimization. Experimental results show strong correlation between the Activity Coefficient and expert feeding intensity, and significant improvements in decision accuracy and language metrics when using dual‑evidence tokens and counterfactual optimization.
By Meng Liang, Guanbo Feng, Haozhuang Chi, Shilong Zhao, Zhixin Xiong, Yuhang He, Wenfeng Han, Tianhao Zhao, Zhihong Ma, Ying Liu
The paper introduces an information‑theoretic framework for studying intent‑hiding jailbreaks against large language models. It defines prior‑posterior matching, where auxiliary tasks are chosen so that the average probability of harmful intent in a bundle matches the overall prior, thereby concealing a harmful target. The authors analyze both query‑independent and query‑dependent settings, proving computational hardness for exact matching, deriving a water‑filling solution for fractional weights, and evaluating the trade‑off between bundle size and target preservation across several open‑source models.
By Fengwei Tian, Ravi Tandon
The study investigates how large language models encode rhetorical questions by applying linear probes to two social‑media datasets. It finds that rhetorical signals appear early in the model’s representations, are most stable in last‑token embeddings, and can be distinguished from information‑seeking questions with AUROC 0.7–0.8 even across datasets. However, probes trained on different datasets rank target instances differently, revealing that multiple, distinct linear directions capture various rhetorical cues rather than a single shared representation.
By Louie Hong Yao, Vishesh Anand, Yuan Zhuang, Tianyu Jiang
MLCommons Jailbreak Benchmark v1.0 is a new methodology for assessing how well large language models resist single‑turn, text‑based jailbreak attacks. It evaluates eight open‑weight systems with 264 seed prompts across eleven hazard categories, using the AILuminate Assessment Standard v1.4 to measure the Resilience Gap—the difference in safety performance between baseline and adversarial conditions. The benchmark found that unsafe‑response rates rose from 11.08% to 18.65% under jailbreak conditions, yielding an average Resilience Gap of 7.57% and highlighting variability in attack effectiveness and evaluator reliability.
By Carsten Maple (Victor), Cagatay Yucel (Victor), Isaac Holeman (Victor), Chris Knotz (Victor), Peter Mattson (Victor), James Goel (Victor), Jonathan Petit (Victor), Sean McGregor (Victor), James Ezick (Victor), Abhishek Kumar (Victor), Alicia Parrish (Victor), Murali Emani (Victor), Kashyap Iyer (Victor), Faiza Khan Khattak (Victor), Washington Mbonu (Victor), Daniel Machlab (Victor), Eileen Long (Victor), Shaona Ghosh (Victor), Jibin Varghese (Victor), Roman Lutz (Victor), Andrew Gruen (Victor), Bennett Hillenbrand (Victor), Prabal Gupta (Victor), Mohammed Serrhini (Victor), Dhivya Nagasubramanian (Victor), Aakash Gupta (Victor), Jun (Victor), Lu, Kurt Bollacker, Chang Liu, Jonathan Petit, Cong Chen, Jean-Philippe Monteuuis, Brent Miller, Apurv Verma, Roman Eng, Armstrong Foundjem, Mohammed Serrhini
VTBench is a new benchmark that evaluates visual tokenizers (VTs) used in autoregressive image generation. It assesses VTs on image reconstruction, detail preservation, and text preservation across diverse scenarios, revealing that continuous VAEs outperform discrete VTs in maintaining spatial structure and semantic detail. The study also explores GPT‑4o’s potential autoregressive behavior and releases the benchmark publicly to encourage development of robust, open‑source VTs.
By Huawei Lin, Tony Geng, Zhaozhuo Xu, Weijie Zhao
The paper investigates why large language models hallucinate by proposing that failures often stem from inference misalignment rather than missing knowledge. It introduces a latent key-task model that shows pretraining-frequency imbalance can cause shortcut inference paths to dominate, leading to hallucinations. The authors create TrapQA, a diagnostic testbed with ScientistQA and Real-Life Constrained QA, to demonstrate two failure modes—task-retrieval bias and key-selection bias—where biased latent inference produces hallucinated answers.
By Yangfan Hu, Xuhan Tong, Haoyue Bai, Xi Ding, Shashank Muralidhar Bharadwaj, Siyang Cao, Robert Nowak, Jiawei Zhang
The paper presents HEAR, a dual‑task Transformer that learns to assess heartbeat observability from mmWave radar phase spectra and simultaneously estimates heart rate. Using a controllable FMCW simulator, the authors generate labeled data where observability is defined by the agreement between the dominant heartbeat‑band peak and the known heart rate. Trained only on simulated data, HEAR transfers zero‑shot to real‑world datasets at 60 and 120 GHz, enabling selective heart‑rate estimation that dramatically reduces error and achieves sub‑50 ms latency on edge hardware.
By Yuxuan Hu, Shilin Shan, Jianfei Yang, Feng Xu
Memristor-based analog compute-in-memory architectures promise efficient deployment of Large Language Models, yet their intrinsic non-idealities degrade reasoning performance across benchmarks. The study evaluates how these non-idealities affect LLM reasoning and tests three training-free mitigation strategies: thinking mode, in-context learning, and module redundancy. Findings show shallow layer redundancy boosts robustness, thinking mode helps only at low noise, and in-context learning shortens output with a modest performance cost.
By Taiqiang Wu, Yuxin Cheng, Chenchen Ding, Runming Yang, Xincheng Feng, Wenyong Zhou, Zhengwu Liu, Ngai Wong
The paper investigates how post‑training fine‑tuning of large language models can reduce their ability to adapt to in‑context information, particularly when the models are fine‑tuned toward one side of cultural‑value disagreements. Experiments show that as a model is trained to favor a specific perspective, its capacity to recognize and enact opposing viewpoints diminishes over time. The authors propose an alternative objective that balances reward maximization with a prescribed distribution over expressed perspectives, offering a practical stance‑distribution matching implementation.
By Jessica Dierking, Itai Shapira, Niclas Boehmer
HARPO is a reinforcement learning framework that jointly optimizes faithfulness and creativity in language generation. It uses a Hallucination-Aware Generative Reward Model (HA‑GRM) to evaluate both faithfulness and writing quality, and a Selective Activation Mechanism (SAM) that applies writing rewards only to hallucination‑free outputs. Experiments on Qwen models show that HARPO improves faithfulness scores and reduces hallucination rates while boosting creative‑writing performance.
By Tiezheng Yu, Yuxin Jiang, Jinpeng Li, Shuning Sun, Fei Mi, Haoli Bai, Lifeng Shang
The paper introduces AIGS, a lightweight Adaptive Incremental Gating System designed for online representation learning in non‑stationary data streams. AIGS uses a Shock Ratio feedback signal to drive a Continuous Plasticity Controller, enabling smooth adaptation between learning plasticity and memory retention while keeping per‑step complexity linear in the feature dimension. Experiments on smart‑city traffic, meteorological, and industrial datasets show that AIGS improves early‑warning lead times, accelerates recovery after abrupt changes, and enhances anomaly recall in noisy environments.
By SiRui He, Kai Liang Lew, Chui Zi Ong, Chean Khim Toa