Transduced language models (TLMs) combine a pretrained source language model with a finite‑state transducer to produce a language model over target strings. The paper introduces an unbiased stochastic estimator that resamples source prefixes without replacement and reweights them, allowing accurate estimation of target prefix probabilities while reducing computation compared to threshold‑pruned beam summing. Experiments on encyclopedic text, DNA, and DNA‑to‑amino‑acid transduction show improved compute–variance trade‑offs and significant runtime reductions, and the method also lowers estimated corpus surprisal in a reading‑time analysis without altering its conclusions.
By V\'esteinn Sn{\ae}bjarnarson, Samuel Kiegeland, Manuel de Prada Corral, Ryan Cotterell, Tim Vieira
arXiv:2607. 24192v1 Announce Type: cross Abstract: We study the problem of lossless compression of source code, motivated by the storage demands of large-scale software archives, such as Software Heritage (https://www.
By Angelo Nardone, Paolo Ferragina
The paper introduces a new complexity measure, based on a restricted class of "layered zipline programs," which is weaker than the previously defined Arithmetic Repetition Complexity (ARC). This measure allows for a prediction algorithm that operates in quasilinear time and polylogarithmic space on highly-structured sequences, offering a more efficient solution at the cost of reduced expressivity compared to ARC. The work highlights a tradeoff between algorithmic efficiency and the breadth of sequences that can be effectively predicted.
By Vanessa Kosoy
arXiv:2608. 01320v1 Announce Type: cross Abstract: Language generation in the limit is a theoretical framework for studying how a generator can learn to produce new valid strings from a stream of positive examples.
By Ziyi Cai, Shuangping Li, Yiheng Shen, Kangning Wang, Peng Zhang
SimLens is a training‑free decoder that improves early‑layer predictions in large language models by keeping only the start token and a candidate answer token and performing a lightweight continuation through the remaining layers. It outperforms direct linear readouts, yielding higher accuracy on tasks such as ARC, BoolQ, and HeadQA with LLaMA‑7B and Vicuna‑7B. The method is extended to Linear SimLens for confidence estimation and combined into SimExit, a hybrid early‑exit mechanism that achieves significant speedups while maintaining accuracy.
By Ming Ma, Bowen Zheng, Zhongqiao Lin, Tianming Yang
arXiv:2606. 29203v1 Announce Type: new Abstract: We study the Bayesian fixed-budget best-arm identification problem in which a learner can abstain from making a terminal recommendation.
By Yuqi Huang, Yunlong Hou, Vincent Y. F. Tan