Pivot‑SD is an offline self‑distillation framework for masked diffusion language models that focuses training on high‑impact commitments, called pivots, identified by an information‑gain metric. By supervising only these pivots—using cross‑entropy for successful trajectories and targeted unlikelihood for failed ones—Pivot‑SD improves LLaDA‑8B‑Instruct on math and code benchmarks with just 200 questions and four rollouts each.
By Seo Hyun Kim, Sunwoo Hong, Younwoo Choi, Chen-Hao Chao, Se-Young Yun, Rahul G. Krishnan
The paper introduces a new loss function called the horizon loss for training classifiers. It argues that the exact policy gradient used in reinforcement learning is myopic, whereas cross‑entropy is patient, and the horizon loss interpolates between the two by truncating the total error at the remaining learning. Experiments on MNIST and ImageNet with ResNet and ViT models show that horizon loss consistently improves top‑1 accuracy over cross‑entropy, especially when label noise is present.
By Ian Osband
arXiv:2610. 02741v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring provides a promising approach for detecting undesirable model behavior.
By Han Wang, Ishwar B Balappanawar, Huan Zhang
Dynamic LLM routers aim to reduce inference costs by directing each query to the cheapest capable model. In a study of six commercial routers across 14 settings and eight task categories, none surpassed a simple random router that selects between two well-chosen models at the same cost, with some underperforming by over 10 percentage points. The authors identify four common patterns—difficulty blindness, length reversal, semantic matching, and roster suboptimality—that explain this gap and propose a new evaluation method and a simple two-model router that mitigates these patterns, though its advantage over random routing remains modest.
By Sam Wang, Julia White, Sahibzada Allahyar, Dhruv Atreja, Urchade Zaratiana, Kelton Zhang
OptiSelect is a framework that incorporates the optimizer’s effect into online data selection for large language model pretraining. The study shows that optimizers like Lion and Muon, which use sign-based or polar-tangential preconditioners, suffer from a discriminability collapse that limits selection gains, whereas diagonal‑adaptive optimizers such as AdamW and Sophia can achieve higher gains. Experiments on 124M and 720M models confirm the theory and demonstrate that OptiSelect remains effective even when data is rephrased.
By Simin Fan, Alireza Abdollahpoorrostam, Martin Jaggi
The paper evaluates how robust large language models (LLMs) are when using retrieval‑augmented generation (RAG) in practical settings. It investigates whether RAG always outperforms non‑RAG approaches, whether adding more retrieved documents helps, and whether the order of documents matters, using a benchmark of 1,891 samples across five datasets and three task categories. Experiments with 11 LLMs show generally high retrieval robustness, but performance varies by task and prompting strategy, indicating that adopting RAG should be considered on a case‑by‑case basis.
By Shuyang Cao, Karthik Radhakrishnan, David Rosenberg, Steven Lu, Pengxiang Cheng, Lu Wang, Shiyue Zhang
The paper introduces threat‑preserving representation sensitivity (TPRS) to assess how changes in agent‑visible wording affect attack success rates (ASR) while keeping the underlying task and evaluation criteria constant. Experiments on Agent Security Bench, MCPTox, and AgentDojo show that renaming threat‑related tools can significantly alter ASR—up to +13.21 percentage points on some models—yet may also degrade benign utility. The findings suggest that security scores based on a single representation may not generalize, urging robustness claims to be validated across multiple threat‑preserving representations.
By Neeraj Karamchandani, Piyush Nagasubramaniam, Xinhong Xie, Sencun Zhu, Dinghao Wu
OLMo-Detect is a new benchmark for membership inference attacks (MIAs) on large language models, covering pre‑training, mid‑training, and post‑training stages. It aligns member and non‑member samples on key axes and rigorously filters non‑members using infini‑gram, also offering a shifted variant to test distribution robustness. Evaluation of 15 unsupervised and 3 supervised MIAs on the OLMo 2 family shows limited overall performance (best AUC 0.68), peak detection at mid‑training driven by data type, and sensitivity to distribution shifts, with findings generalizing to OLMo 3 and other models.
By Tao Shi, Chaoyi Xiang, Qiongkai Xu, Jey Han Lau
Gaze Attention is a new mechanism for multimodal large language models that selectively focuses on relevant visual regions during each generation step, rather than attending to all visual tokens. By grouping tokens into spatial regions and using learnable context tokens to retain global information, it reduces attention computation and visual key‑value entries by up to 90%. Experiments on 13 image and 6 video benchmarks show that Gaze Attention matches or outperforms dense‑attention baselines while using fewer visual resources.
By Junha Song, Byeongho Heo, Geonmo Gu, Jaegul Choo, Dongyoon Han, Sangdoo Yun
The paper investigates how well large language models (LLMs) can serve as judges in summarization evaluation by applying psychometric techniques. Using Many‑Facet Rasch Models, the authors decompose human and LLM ratings into latent summary quality, rater severity, dimension severity, and rating‑scale thresholds, and introduce a residual hardness metric to capture judging difficulty. Their analysis of 17 open‑weight LLM judges on the SummEval dataset reveals that moderate alignment in latent quality does not translate to alignment in residual hardness; humans and LLMs differ in which summary–dimension units are hard, with LLMs tending to find consistency hard and humans tending to find coherence hard, and some hard cases can be predicted from observable source–summary properties.
By Longwei Cong, Sonja Hahn, Sebastian Gombert, Leon Camus, Fabian Zehner, Hendrik Drachsler, Ulf Kroehne
The paper investigates whether adding a short classification instruction after a user’s prompt improves the ability of activation probes to detect malicious inputs in large language models. Across 13 safety benchmarks and three open‑weight model families, a classification suffix consistently boosts out‑of‑distribution detection (up to ~4 AUC points) compared to no suffix, and the benefit transfers to multi‑position pooling probes used in production. The improvement stems from the classification format itself rather than the specific content of the instruction, though the optimal suffix varies with the model and readout type.
By Elad David, Max Fomin
NeutronGym is the first executable environment that lets language‑model agents design neutron instruments, using tools that validate their builds, McStas ray‑tracing for simulation, and a level‑resolved grading ladder that evaluates syntax, runtime, structure, and science without an LLM judge. The platform provides procedural families of instrument layouts with held‑out parameter regimes and a curated set of 16 tasks from published instruments, called McStasBench, which includes memorization probes and a sandbox. Experiments show that while seven models can reproduce at most seven of the 16 tasks and none reaches a reference design, reinforcement learning can dramatically improve performance—e.g., Qwen3‑8B’s success rate jumps from 11% to 77% on held‑out instances—highlighting the ladder’s importance for partial credit and the potential of RL to match classical optimizers under realistic simulation budgets.
By Lijie Ding, Changwoo Do
RT‑SFT is a method for text style transfer that uses roundtrip translation through a pivot language to strip stylistic information from monolingual corpora, creating pseudo‑parallel data. This data is then used to LoRA‑finetune an instruction‑tuned large language model as a stylizer, allowing the model to rewrite sentences in a target style while preserving meaning. Experiments across four style domains show that RT‑SFT surpasses state‑of‑the‑art approaches, including few‑shot in‑context learning, and offers effective retrieval augmentation for expert style domains with strict terminology.
By Ruoxi Liu, Philipp Koehn
ReFract is a new benchmark that tests whether language model agents can act appropriately based on a user’s role, a capability called Perspective Awareness. It contains 150 expert‑validated entries derived from anonymized domain support queries, and uses Text World Models to simulate the agents’ operating environments and generate perspective‑aware action trajectories. Current state‑of‑the‑art LLMs solve at most 69% of the tasks, with over half of their trajectories attempting perspective‑violating actions.
By Hainiu Xu, V\'{i}tor N. Louren\c{c}o, Mohnish Dubey, Yunfei Bai, Yulan He, Caroline Catmur, Aline Paes, Marco Caserta, Akash Chandrayan, Luca D'Angelo
The paper investigates why certain synthetic relational datasets lead to better relational foundation models. By training four Relational Transformer checkpoints on data from four different generators, the authors trace a measurable property of the data—specifically the predictive necessity of cross‑table information—to the emergence of relational computation in the models. They find that the RelDiff generator yields the largest predictive gain from foreign‑key‑linked parents, and its model uniquely responds to foreign‑key interventions, a dependence that persists across random initializations and grows with corrupted links. Disrupting this mechanism during downstream inference eliminates RelDiff’s advantage on relational tasks while leaving structure‑insensitive models largely unchanged, thereby linking synthetic data properties to learned mechanisms and downstream behavior.
By Shivam Dubey, Mohamed Bouadi, Nassim Bouarour, Varun Kulkarni, Aditya Tanna, Vinay Kumar Sankarapu
The paper introduces THPL, a vision-to-language framework for precision feeding of rainbow trout in recirculating aquaculture systems. It uses Fishsort to extract feeding trajectories, a Hierarchical Behavior Encoder to model individual and collective dynamics, and integrates these with environmental data and expert rules to fine‑tune a large language model via LoRA and counterfactual multimodal Direct Preference Optimization. Experimental results show strong correlation between the Activity Coefficient and expert feeding intensity, and significant improvements in decision accuracy and language metrics when using dual‑evidence tokens and counterfactual optimization.
By Meng Liang, Guanbo Feng, Haozhuang Chi, Shilong Zhao, Zhixin Xiong, Yuhang He, Wenfeng Han, Tianhao Zhao, Zhihong Ma, Ying Liu
The paper introduces an information‑theoretic framework for studying intent‑hiding jailbreaks against large language models. It defines prior‑posterior matching, where auxiliary tasks are chosen so that the average probability of harmful intent in a bundle matches the overall prior, thereby concealing a harmful target. The authors analyze both query‑independent and query‑dependent settings, proving computational hardness for exact matching, deriving a water‑filling solution for fractional weights, and evaluating the trade‑off between bundle size and target preservation across several open‑source models.
By Fengwei Tian, Ravi Tandon
The study investigates how large language models encode rhetorical questions by applying linear probes to two social‑media datasets. It finds that rhetorical signals appear early in the model’s representations, are most stable in last‑token embeddings, and can be distinguished from information‑seeking questions with AUROC 0.7–0.8 even across datasets. However, probes trained on different datasets rank target instances differently, revealing that multiple, distinct linear directions capture various rhetorical cues rather than a single shared representation.
By Louie Hong Yao, Vishesh Anand, Yuan Zhuang, Tianyu Jiang
MLCommons Jailbreak Benchmark v1.0 is a new methodology for assessing how well large language models resist single‑turn, text‑based jailbreak attacks. It evaluates eight open‑weight systems with 264 seed prompts across eleven hazard categories, using the AILuminate Assessment Standard v1.4 to measure the Resilience Gap—the difference in safety performance between baseline and adversarial conditions. The benchmark found that unsafe‑response rates rose from 11.08% to 18.65% under jailbreak conditions, yielding an average Resilience Gap of 7.57% and highlighting variability in attack effectiveness and evaluator reliability.
By Carsten Maple (Victor), Cagatay Yucel (Victor), Isaac Holeman (Victor), Chris Knotz (Victor), Peter Mattson (Victor), James Goel (Victor), Jonathan Petit (Victor), Sean McGregor (Victor), James Ezick (Victor), Abhishek Kumar (Victor), Alicia Parrish (Victor), Murali Emani (Victor), Kashyap Iyer (Victor), Faiza Khan Khattak (Victor), Washington Mbonu (Victor), Daniel Machlab (Victor), Eileen Long (Victor), Shaona Ghosh (Victor), Jibin Varghese (Victor), Roman Lutz (Victor), Andrew Gruen (Victor), Bennett Hillenbrand (Victor), Prabal Gupta (Victor), Mohammed Serrhini (Victor), Dhivya Nagasubramanian (Victor), Aakash Gupta (Victor), Jun (Victor), Lu, Kurt Bollacker, Chang Liu, Jonathan Petit, Cong Chen, Jean-Philippe Monteuuis, Brent Miller, Apurv Verma, Roman Eng, Armstrong Foundjem, Mohammed Serrhini
VTBench is a new benchmark that evaluates visual tokenizers (VTs) used in autoregressive image generation. It assesses VTs on image reconstruction, detail preservation, and text preservation across diverse scenarios, revealing that continuous VAEs outperform discrete VTs in maintaining spatial structure and semantic detail. The study also explores GPT‑4o’s potential autoregressive behavior and releases the benchmark publicly to encourage development of robust, open‑source VTs.
By Huawei Lin, Tony Geng, Zhaozhuo Xu, Weijie Zhao