The paper presents a systematic study of how different compression techniques—pruning, quantization, and distillation—affect the capabilities of large language models (LLMs) in tasks such as mathematics, code generation, and question answering. It introduces a framework that measures capability loss and relates it to factors like model size, training stage, and compression settings, yielding simple predictive relations that generalize across unseen configurations. The authors demonstrate that sharing density responses across pruning levels can dramatically reduce the number of measurements needed, and that their predictive models closely match regression results while offering efficient decision guidance for compression method selection.
By Xueqi Cheng, Liang Wu, Kelly Wan, Liangjie Hong, Yushun Dong
The paper investigates why increasing batch size during large language model pretraining helps, challenging the common assumption of bounded stochastic gradient variance. By introducing a generalized Blum–Gladyshev noise model that allows variance to grow with a tunable exponent, the authors derive theoretical bounds on oracle complexity and design an adaptive batch scheduler that adjusts batch size dynamically. Experiments on OLMo2 models up to 1B parameters show that this scheduler achieves lower validation loss than fixed small or large batch training while using fewer iterations.
By Arda Fazla, Antesh Upadhyay, Ege C. Kaya, M. Berk Sahin, Abolfazl Hashemi
PowerBench is a new evaluation framework that measures how language models respond to power‑shifting requests, distinguishing self‑empowerment, disempowerment, and power grabbing while controlling for neutral requests. The authors built and released a dataset covering different power domains, contexts, scales, and prior power standings, and tested 24 models from US and Chinese developers across three experimental conditions. Results show models are more likely to refuse power‑grabbing than disempowerment, and that biases exist toward helping US users gain power while resisting US users taking power from others, with additional effects when the user is an AI agent or when requests are in different languages.
By Nicolas Martorell, Wendy Brau, Gonzalo A. Heredia, Tom\'as Pablo Korenblit, Gaspar Labasti\'e, Tom\'as Gimenez Molina
Tropical Reinforcement Learning replaces the traditional sum of probabilities with a maximum operation, forming a tropical semiring. This change allows the value of a state to reflect the log-probability of its most likely verified solution and provides an explicit path that can be replayed. The proposed TROPIC algorithm, applied to deterministic, resettable environments, outperforms strong on‑policy baselines on tasks such as Sokoban, Countdown, FrozenLake, and WebShop.
By Arip Asadulaev, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, Martin Takac
The paper introduces DiffGCMS, a discrete graph diffusion model that generates molecular structures directly from GC–EI–MS spectra, and couples it with a large language model for post‑processing. In the first stage, DiffGCMS produces candidate structures; the LLM then validates, repairs, and reranks these candidates while offering interpretable explanations of fragment‑ion peaks. On a large NIST 20 test set, the combined approach improves accuracy and ensures 100% candidate validity for small molecules, demonstrating the value of spectrum‑aware post‑processing.
By Changlin Liu, Tianyu Yi, Chengchun Liu, Boxuan Zhao, Fanyang Mo
EviDent-CBCT is a framework that generates dental cone‑beam CT reports from limited clinical data by first extracting a discrete record of tooth‑level, global, and tooth‑IAC evidence using an anatomy‑aware network. The evidence is reconciled through a dental‑logic consistency projection before being rendered into a report by a deterministic renderer and an image‑blind language model. The system achieves higher evidence set‑F1 and logical‑F1 scores than baselines and ranked highly in the ODIN 2026 challenge, demonstrating the effectiveness of a discrete evidence record for auditable CBCT report generation.
By Ruiyang Hao, Zhi Qin Tan, Yulan He, Owen Addison, Yunpeng Li
RxnOptBench is a new benchmark that tests large language models on the task of selecting optimal reaction conditions—such as catalyst, ligand, solvent, temperature, and atmosphere—by reading real condition‑screening tables from 2025 organic‑methodology papers. The benchmark uses a continuous relative score that combines yield with stereochemical metrics (ee, dr, rr) and includes a design that separates in‑context literature use from mere memorization. Evaluation across nine frontier LLMs and three chemistry‑specialized LLMs shows that even the best models still have significant room for improvement, with chemistry‑specialized models performing at random levels on multi‑axis selection while open‑weight models approach proprietary models.
By Lingli Ge, Yubin Wang, Junyuan Gao, Jiahe Song, Jiaxing Sun, Boyu Zhu, Haote Yang, Jingchao Wang, Lixin Ma, Jiang Wu, Yuqiang Li, Conghui He
Predictor-Guided Latent Space Codon Optimization (LSCO) transforms the discrete codon selection problem into a continuous optimization task by embedding sequences into the latent space of a pretrained mRNA language model, allowing gradient-based search. LSCO integrates a data-driven expression objective from an uncertainty-aware predictor, a Minimum-Free-Energy regularizer for structural stability, a naturalness prior derived from a protein-to-codon back-translation model, and constrained decoding to preserve protein fidelity. In a real-world wet-lab antibody expression dataset, LSCO outperforms both simple frequency-based methods and modern deep generative baselines in predicted expression while maintaining appropriate biophysical properties.
By Alberto Caron, Tianyu Cui, Dmytro S. Lituiev, Mangal Prakash, Artem Moskalev, Amina Mollaysa, Bo Zhai, Hirsh Nanda, Daniel M. Poole, Zhongyin Liu, Iman Farasat, Robert Davidson, Nikolay V. Manyakov, Tommaso Mansi, Scott Oloff, Rui Liao
The paper introduces Pluralistic Preference Optimization (PlurPO), a method that reduces social sycophancy in language models by having the model simulate multiple stakeholders’ perspectives when responding to interpersonal conflict scenarios. PlurPO trains the model to prefer responses acceptable to all simulated stakeholders, using only the model’s own output signals without ground‑truth labels. Experiments across four datasets and model families show significant reductions in sycophantic endorsement, including an 89% drop on harmful intent statements and halving the gap in general advice endorsement rates.
By Stephane Hatgis-Kessell, Myra Cheng, Xiaoxuan Hou, Qian Hu, Rahul Gupta, Natasha Jaques, Emma Brunskill
The paper evaluates two strategies for training language‑model‑based world models—fine‑tuning and Retrieval Augmented Generation (RAG)—across five diverse environments. Fine‑tuning generally yields higher rewards, while RAG is more data‑efficient. The authors propose a counterfactual error estimation for RAG, improve retrieval with a hierarchical query reformulation, and combine both approaches into a hybrid system that outperforms each method alone.
By Dhananjay Ashok, Shantanu Agarwal, Vivek Datla, Jonathan May, Alfy Samuel
The paper introduces HypoAgent, an agentic framework for refining abductive hypotheses over knowledge graphs. It treats user-provided conditions as steering operators, using a small generator for initial proposals, root-cause analysis to identify problematic branches, and a refiner agent that revises hypotheses based on diagnosis and prior proposals. Experiments on BioKG, PharmKG8k, and DBpedia50 show HypoAgent surpasses one-shot generation in both single-turn and multi-turn scenarios, including unconditional settings.
By Yisen Gao, Yixi Cai, Tianshi Zheng, Jiaxin Bai, Yangqiu Song
WebUIProof is a new execution‑oriented benchmark for WebUI code generation that supplies structured specifications and dense, executable interaction tests across general WebUIs and 3D interactive simulations. It employs a UI‑agent harness that runs tests in a headless browser using a plan–act–observe loop to locate DOM elements, perform actions, observe changes, and verify assertions. Evaluations on eight commercial LLMs reveal frequent interaction‑based failures, especially on 3D interfaces, and demonstrate that training compact models with RL rewards from these tests improves functional completion and reduces build failures.
By Yun-Yun Tsai, Yuning Mao, Shiqi Wang, Junfeng Yang, Sinong Wang
The paper introduces Whiteout, a tool that prevents large language models from revealing personally sensitive information (PSI) by overwriting such data with carefully crafted obfuscation samples. Whiteout is evaluated on various LLMs, including an OpenAI model, and shows effective PSI protection with minimal impact on model utility and safety. The study also tests Whiteout against multiple attack vectors and discusses its security and ethical implications.
By Anna Yoo Jeong Ha, Ronik Bhaskar, Haitao Zheng, Ben Y. Zhao
The report audits a long‑term‑memory retrieval chain on 500 LongMemEval‑S development questions, evaluating reader performance under an adapted GPT‑4o rubric. Scores vary widely across reader lanes (93–479), with the strongest lanes scoring 479 and 475, and re‑judging the same pass‑1 answers changes three labels, yielding 478. Fixed‑answer knowledge‑update re‑scoring achieves 70/72 or 69/72 depending on the template, and a different‑family reader scores 474, just 1.0 percentage point below the headline pass. The study also reports that live reader request bodies were not retained, that 18 of 23 gains occurred where baseline packets lacked evidence, and that a negative control rejects a verifier that repairs some wrong drafts but breaks many correct ones. Overall, the findings do not establish a new leaderboard leader or a transferable memory advantage, and the released artifacts support packet inspection and re‑scoring but do not reconstruct the method.
By Christopher J. Chanhnourack
The paper introduces GUI agents for continual game generation, presenting PlaytestArena—a benchmark of 200 browser-based game-generation tasks with rubrics for in‑play behavior—and Play2Code, an iterative framework where a game agent and a rubric‑blind GUI playtester refine games through shared memory. Play2Code achieves a 66.8% rubric pass rate, surpassing baseline methods by 37.1 and 14.6 points, and shows consistent score improvement across refinement rounds. The study demonstrates that GUI playtesting provides actionable, traceable feedback that can guide interactive code generation.
By Yixu Huang, Bo Li, Na Li, Zhe Wang, Kaijie Chen, Haonan Ge, Qingyi Si, Yuanzhe Shen, Ruihan Yang, Guangjing Wang, Hongcheng Guo
FastKernels is a benchmark comprising 384 GPU kernel tasks from 47 architectures across 8 categories, designed to evaluate kernel generation agents in realistic production settings. It covers 94.6% of HuggingFace Transformers architectures, scoring kernels on correctness, coverage, and speedup within their native execution paths. The benchmark reveals that kernel‑level speedups often collapse to modest end‑to‑end gains and that isolated performance rankings can misrepresent real‑world effectiveness.
By Gabriele Oliaro, Jaeseong Lee, Yichao Fu, Zhaoyuan Su, Owen Lu, Sreeram Vennam, May Jiang, Junli Wang, Hao Zhang, Zhihao Jia, Samyam Rajbhandari
The paper introduces BenchGuard, a framework that uses advanced language models to audit execution-based agent benchmarks. BenchGuard systematically cross-verifies benchmark artifacts through structured LLM protocols, optionally incorporating agent solutions or execution traces for deeper diagnostics. In tests on ScienceAgentBench and BIXBench Verified-50, BenchGuard uncovered numerous author-confirmed and expert-identified issues, including fatal errors that made tasks unsolvable, and demonstrated cost-effective auditing for complex bioinformatics tasks.
By Xinming Tu (Minta), Tianze Wang (Minta), Yingzhou (Minta), Lu, Kexin Huang, Yuanhao Qu, Sara Mostafavi
Goldilocks RL proposes an adaptive data‑selection strategy for reinforcement learning in language models, using a Selector network to predict reward variability and prioritize questions that are neither too easy nor too hard. By continuously adapting to the model’s evolving abilities, Goldilocks improves training efficiency on large‑scale reasoning datasets, achieving up to 78% fewer optimization steps compared to standard GRPO. The approach addresses the sample‑inefficiency problem of sparse rewards in reasoning tasks.
By Ilia Mahrooghi, Aryo Lotfi, Emmanuel Abbe
The paper investigates the gap between retrieval and reading in document vision‑language models, showing that even when the correct page is retrieved, the model often fails to use the text on that page. By comparing answers derived from page images alone versus images plus extracted OCR text, the authors demonstrate that adding OCR text can improve strict accuracy by 13 to 16 points on their FoveDoc‑Bench benchmark. The study also reveals that OCR benefits textual evidence but not charts or figures, and that its advantage diminishes as retrieval quality worsens.
By Qingtao Xia, Siyao Cheng, Jiahua Bao, Jiaxing Du, Jie Liu
The paper evaluates how well generalist and dermatology-specific machine learning models perform on diverse skin lesion datasets, including dermoscopic images and smartphone photographs. It benchmarks a range of architectures—general-purpose vision-language models, foundation models, and task-specific dermatology classifiers—under conditions of distribution shift, modality change, and demographic variability. The study quantifies the performance gap between current state‑of‑the‑art models and the robustness needed for safe, equitable clinical deployment.
By Emanoel dos Santos, Kelvin Cunha, Rodrigo Mota, Fabio Papais, Thales Bezerra, Natalia Lopes, Erico Medeiros, Shirley Cruz, Jessica Araujo, Paulo Borba, Tsang Ing Ren