ReFract is a new benchmark that tests whether language model agents can act appropriately based on a user’s role, a capability called Perspective Awareness. It contains 150 expert‑validated entries derived from anonymized domain support queries, and uses Text World Models to simulate the agents’ operating environments and generate perspective‑aware action trajectories. Current state‑of‑the‑art LLMs solve at most 69% of the tasks, with over half of their trajectories attempting perspective‑violating actions.
By Hainiu Xu, V\'{i}tor N. Louren\c{c}o, Mohnish Dubey, Yunfei Bai, Yulan He, Caroline Catmur, Aline Paes, Marco Caserta, Akash Chandrayan, Luca D'Angelo
The paper investigates why certain synthetic relational datasets lead to better relational foundation models. By training four Relational Transformer checkpoints on data from four different generators, the authors trace a measurable property of the data—specifically the predictive necessity of cross‑table information—to the emergence of relational computation in the models. They find that the RelDiff generator yields the largest predictive gain from foreign‑key‑linked parents, and its model uniquely responds to foreign‑key interventions, a dependence that persists across random initializations and grows with corrupted links. Disrupting this mechanism during downstream inference eliminates RelDiff’s advantage on relational tasks while leaving structure‑insensitive models largely unchanged, thereby linking synthetic data properties to learned mechanisms and downstream behavior.
By Shivam Dubey, Mohamed Bouadi, Nassim Bouarour, Varun Kulkarni, Aditya Tanna, Vinay Kumar Sankarapu
The paper introduces THPL, a vision-to-language framework for precision feeding of rainbow trout in recirculating aquaculture systems. It uses Fishsort to extract feeding trajectories, a Hierarchical Behavior Encoder to model individual and collective dynamics, and integrates these with environmental data and expert rules to fine‑tune a large language model via LoRA and counterfactual multimodal Direct Preference Optimization. Experimental results show strong correlation between the Activity Coefficient and expert feeding intensity, and significant improvements in decision accuracy and language metrics when using dual‑evidence tokens and counterfactual optimization.
By Meng Liang, Guanbo Feng, Haozhuang Chi, Shilong Zhao, Zhixin Xiong, Yuhang He, Wenfeng Han, Tianhao Zhao, Zhihong Ma, Ying Liu
The paper introduces an information‑theoretic framework for studying intent‑hiding jailbreaks against large language models. It defines prior‑posterior matching, where auxiliary tasks are chosen so that the average probability of harmful intent in a bundle matches the overall prior, thereby concealing a harmful target. The authors analyze both query‑independent and query‑dependent settings, proving computational hardness for exact matching, deriving a water‑filling solution for fractional weights, and evaluating the trade‑off between bundle size and target preservation across several open‑source models.
By Fengwei Tian, Ravi Tandon
The study investigates how large language models encode rhetorical questions by applying linear probes to two social‑media datasets. It finds that rhetorical signals appear early in the model’s representations, are most stable in last‑token embeddings, and can be distinguished from information‑seeking questions with AUROC 0.7–0.8 even across datasets. However, probes trained on different datasets rank target instances differently, revealing that multiple, distinct linear directions capture various rhetorical cues rather than a single shared representation.
By Louie Hong Yao, Vishesh Anand, Yuan Zhuang, Tianyu Jiang
MLCommons Jailbreak Benchmark v1.0 is a new methodology for assessing how well large language models resist single‑turn, text‑based jailbreak attacks. It evaluates eight open‑weight systems with 264 seed prompts across eleven hazard categories, using the AILuminate Assessment Standard v1.4 to measure the Resilience Gap—the difference in safety performance between baseline and adversarial conditions. The benchmark found that unsafe‑response rates rose from 11.08% to 18.65% under jailbreak conditions, yielding an average Resilience Gap of 7.57% and highlighting variability in attack effectiveness and evaluator reliability.
By Carsten Maple (Victor), Cagatay Yucel (Victor), Isaac Holeman (Victor), Chris Knotz (Victor), Peter Mattson (Victor), James Goel (Victor), Jonathan Petit (Victor), Sean McGregor (Victor), James Ezick (Victor), Abhishek Kumar (Victor), Alicia Parrish (Victor), Murali Emani (Victor), Kashyap Iyer (Victor), Faiza Khan Khattak (Victor), Washington Mbonu (Victor), Daniel Machlab (Victor), Eileen Long (Victor), Shaona Ghosh (Victor), Jibin Varghese (Victor), Roman Lutz (Victor), Andrew Gruen (Victor), Bennett Hillenbrand (Victor), Prabal Gupta (Victor), Mohammed Serrhini (Victor), Dhivya Nagasubramanian (Victor), Aakash Gupta (Victor), Jun (Victor), Lu, Kurt Bollacker, Chang Liu, Jonathan Petit, Cong Chen, Jean-Philippe Monteuuis, Brent Miller, Apurv Verma, Roman Eng, Armstrong Foundjem, Mohammed Serrhini
VTBench is a new benchmark that evaluates visual tokenizers (VTs) used in autoregressive image generation. It assesses VTs on image reconstruction, detail preservation, and text preservation across diverse scenarios, revealing that continuous VAEs outperform discrete VTs in maintaining spatial structure and semantic detail. The study also explores GPT‑4o’s potential autoregressive behavior and releases the benchmark publicly to encourage development of robust, open‑source VTs.
By Huawei Lin, Tony Geng, Zhaozhuo Xu, Weijie Zhao
The paper investigates why large language models hallucinate by proposing that failures often stem from inference misalignment rather than missing knowledge. It introduces a latent key-task model that shows pretraining-frequency imbalance can cause shortcut inference paths to dominate, leading to hallucinations. The authors create TrapQA, a diagnostic testbed with ScientistQA and Real-Life Constrained QA, to demonstrate two failure modes—task-retrieval bias and key-selection bias—where biased latent inference produces hallucinated answers.
By Yangfan Hu, Xuhan Tong, Haoyue Bai, Xi Ding, Shashank Muralidhar Bharadwaj, Siyang Cao, Robert Nowak, Jiawei Zhang
The paper presents HEAR, a dual‑task Transformer that learns to assess heartbeat observability from mmWave radar phase spectra and simultaneously estimates heart rate. Using a controllable FMCW simulator, the authors generate labeled data where observability is defined by the agreement between the dominant heartbeat‑band peak and the known heart rate. Trained only on simulated data, HEAR transfers zero‑shot to real‑world datasets at 60 and 120 GHz, enabling selective heart‑rate estimation that dramatically reduces error and achieves sub‑50 ms latency on edge hardware.
By Yuxuan Hu, Shilin Shan, Jianfei Yang, Feng Xu
Memristor-based analog compute-in-memory architectures promise efficient deployment of Large Language Models, yet their intrinsic non-idealities degrade reasoning performance across benchmarks. The study evaluates how these non-idealities affect LLM reasoning and tests three training-free mitigation strategies: thinking mode, in-context learning, and module redundancy. Findings show shallow layer redundancy boosts robustness, thinking mode helps only at low noise, and in-context learning shortens output with a modest performance cost.
By Taiqiang Wu, Yuxin Cheng, Chenchen Ding, Runming Yang, Xincheng Feng, Wenyong Zhou, Zhengwu Liu, Ngai Wong
The paper investigates how post‑training fine‑tuning of large language models can reduce their ability to adapt to in‑context information, particularly when the models are fine‑tuned toward one side of cultural‑value disagreements. Experiments show that as a model is trained to favor a specific perspective, its capacity to recognize and enact opposing viewpoints diminishes over time. The authors propose an alternative objective that balances reward maximization with a prescribed distribution over expressed perspectives, offering a practical stance‑distribution matching implementation.
By Jessica Dierking, Itai Shapira, Niclas Boehmer
HARPO is a reinforcement learning framework that jointly optimizes faithfulness and creativity in language generation. It uses a Hallucination-Aware Generative Reward Model (HA‑GRM) to evaluate both faithfulness and writing quality, and a Selective Activation Mechanism (SAM) that applies writing rewards only to hallucination‑free outputs. Experiments on Qwen models show that HARPO improves faithfulness scores and reduces hallucination rates while boosting creative‑writing performance.
By Tiezheng Yu, Yuxin Jiang, Jinpeng Li, Shuning Sun, Fei Mi, Haoli Bai, Lifeng Shang
The paper introduces AIGS, a lightweight Adaptive Incremental Gating System designed for online representation learning in non‑stationary data streams. AIGS uses a Shock Ratio feedback signal to drive a Continuous Plasticity Controller, enabling smooth adaptation between learning plasticity and memory retention while keeping per‑step complexity linear in the feature dimension. Experiments on smart‑city traffic, meteorological, and industrial datasets show that AIGS improves early‑warning lead times, accelerates recovery after abrupt changes, and enhances anomaly recall in noisy environments.
By SiRui He, Kai Liang Lew, Chui Zi Ong, Chean Khim Toa
The study examined how variation in the interpulse interval (IPI) of bat vocalizations affects deep‑learning classification. Two datasets—one preserving natural IPI timing and another normalizing call spacing to 50 ms—were used to fine‑tune EfficientNet‑B0 and PaSST models. Results showed that IPI normalization had mixed effects: PaSST performance remained stable while EfficientNet improved, yet models trained on natural IPI data performed better on natural test sets, indicating limited but non‑negligible influence of natural IPI variation.
By Welmoed R. Eversteijn, Burooj Ghani, A. Leonie Baier, Dan Stowell
The paper introduces Conformal Interval-Driven Self-Evolution (CISE), a method that uses conditional conformal inference and online density-ratio estimation to create candidate‑specific reward intervals for self‑evolving search in materials science. CISE applies conservative interval‑based rewards, ensuring that only candidates whose required property intervals lie entirely within feasible regions are returned. Experiments on three self‑evolving search tasks show that all candidates returned by CISE are true positives under high‑fidelity evaluation, whereas baseline methods produce more candidates but include false positives.
By Kangjun Noh, Soyu Kim, Kyungwoo Song
PLCWorld is a closed‑loop execution environment and benchmark that evaluates large language model (LLM)–generated programmable logic controller (PLC) programs by coupling Structured Text (ST) execution with simulated plant responses and sensor feedback. It includes 100 synthetic tasks and 473 task‑condition pairs across Motion Control and Material Handling, with difficulty levels based on control‑dependency scope. The benchmark reports Task Success and Safety Violation separately and validates results through practitioner review, reference programs, counterexamples, and comparisons with independent ST runtimes, revealing performance differences among six LLMs and four generation‑and‑verification workflows.
By Yunji Kim, Yunseok Lee, Hyunwoo Seo, Jaerim Choi, Woojin Lee
SEDIMA is a persistent hierarchical insight memory designed for evolutionary search agents that use large language models. It transforms raw search traces into natural‑language insights, clusters them by semantic similarity, and retrieves relevant guidance to inform future mutations, thereby accumulating transferable knowledge across runs and problems. When added as a drop‑in module, SEDIMA improves average final performance by 5.5% on AlgoTune and 6.6% on ALE‑Bench LITE, and reduces the number of iterations needed to reach baseline‑best performance by 32.3% on OpenEvolve.
By Amirhossein Abaskohi, Mahdi Mostajabdaveh, Zirui Zhou
The paper introduces Calibrated User Embeddings (CUE), a framework that encodes real user sessions into continuous representations and decodes them into persona commands to steer large language models (LLMs) as user simulators without additional training. CUE enables user-conditioned replay of past interactions and generates novel personas that better align with real-user success rates and failure patterns. Evaluations on the τ²‑Bench dataset show that CUE‑driven simulators commit fewer simulator‑attributed errors, more accurately reproduce real‑user failure modes, and maintain competitive user fidelity across various tasks and LLMs.
By Anjali Kantharuban, Jonas Mueller
The paper investigates why supervised AI‑text detectors, specifically a RoBERTa‑based model, can be fooled by subtle changes in language. By applying semantic, structural, and tokenizer‑level perturbations to a large dataset and controlled Mistral‑7B‑Instruct outputs, the authors show that increasing verb diversity makes machine text easier to detect and that detection scores correlate with statistical complexity, leading to a high false‑positive rate on formal human writing. They also evaluate an event‑based latent space detector, finding that paraphrasing and homoglyphs significantly alter extracted event sequences and verbs, yet the detector’s performance remains modest (AUC 0.577).
By Claudiu Creanga, Liviu Dinu
The paper introduces a method for predicting and preventing merge collapse when combining large language models fine‑tuned from a shared base. By measuring the variance of specialists’ task vectors—termed interference—a pre‑merge score identifies destructive merges. The authors present PRISM, an operator that averages task vectors and then soft‑thresholds each layer based on interference, which successfully preserves model performance without additional data or tuning.
By Jungseob Lee, Seungyoon Lee, Sugyeong Eo, Hyeonseok Moon, Jaehyung Seo, Heuiseok Lim