arXiv Machine Learning

Exact ReLU realization of binary affine refinement iterates via reflection folding and cone switching

The paper investigates vector‑valued binary affine refinement operators with finite matrix masks and compactly supported continuous piecewise‑linear data. It demonstrates that every finite refinement iterate can be exactly realized by a ReLU network of fixed width and depth linear in the number of iterations, using a universal reflection‑doubling mechanism that replaces two binary transition matrices with a single fixed block matrix and a swap involution. The construction allows exact branch selection via a continuous piecewise‑linear cone switch, propagates full vectorized profiles without decomposing inputs, and handles stage‑dependent forcing while reducing the doubled cascade to a single parity sector through genuine reflection equivariance.

arXiv Machine Learning
Sep 10

Compressed Recurrent Feedback in Tsetlin Machines: A Reproducible Boolean-FSM Study

The paper proposes a fixed‑width recurrent feedback scheme for Recurrent Tsetlin Machines (RTMs) by folding clause activations with XOR, retaining the folded bits at two time scales, and thresholding them back to binary. This compression reduces 480 clause activations to 96 recurrent bits while achieving comparable accuracy (61.47 ± 6.74% and 62.94 ± 9.92%) on a reproducible Boolean finite‑state‑machine benchmark across 144 runs. The study shows that raw clause feedback offers only marginal accuracy gains but increases recurrent width and execution time significantly, and highlights the importance of no‑feedback controls in sequence model benchmarking.

By Ankit Kumar, Utkarsh Raj, Rishad Shafik, Sudip Roy
arXiv Machine Learning
Aug 27

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

The paper introduces Gated Recurrent Transformers, a depth‑sharing architecture that brackets a single shared core with fixed prelude and coda blocks and uses a lightweight projection and element‑wise update gate to modulate recurrent updates. This design allows functional specialization across recurrences while reducing memory footprint. Experiments show that, under equal FLOPs or parameter budgets, the recurrent model matches or surpasses deeper GPT‑2 Small baselines, achieving similar or better accuracy with fewer parameters and lower peak decoding memory.

By Amr Hegazy, Amr Alanwar, Mostafa Elhoushi
arXiv AI
Sep 23

SPECTRA: Adaptive Execution of Speculative Decoding on a Runtime-Reconfigurable Tiled Architecture

The paper introduces SPECTRA, a runtime‑reconfigurable tiled architecture designed to accelerate speculative decoding for large language models on edge devices. SPECTRA adapts its compute engine within each tile between systolic GEMM execution and vector‑lane GEMV execution, while dynamically adjusting tile count, kernel partitioning, and communication patterns across tiles. Experiments on a 20‑tile FPGA prototype demonstrate up to a 2.09× speedup from tile‑level reconfiguration and an additional 1.25× improvement from system‑level adaptability compared to fixed designs.

By Gabriele Tombesi, William Baisi, Je Yang, Elisavet Lydia Alvanaki, Kevin Lee, Michael Lippe, Biruk Seyoum, Luca P. Carloni
arXiv Machine Learning
1d ago

Decoding Looped Transformers Better for (Almost) Free

The paper introduces LoopCD, a training‑free contrastive decoding framework that improves token selection in Loop‑Transformer models by comparing the final prediction with earlier recurrent passes. LoopCD operates either in logit space (LoopCD‑Logits) with a single extra output pass or in hidden‑state space (LoopCD‑Hidden) with no output overhead. Across multiple looped Transformer families, LoopCD yields significant performance gains—raising pass@1 scores on tasks such as AIME 2024 and HumanEval—while enabling a reduction in the number of recurrent loops and a corresponding decrease in inference FLOPs.

By Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
arXiv Machine Learning
Sep 14

Dual-guided Hierarchical Edge Localization for Large-scale Optimal Transport Across Dimensions

The paper introduces HELLO, a hierarchical solver for large‑scale discrete optimal transport that reduces the problem to edge localization guided by dual potentials. HELLO uses a coarse‑to‑fine initialization across a recursive subsampling hierarchy and a refinement step that inserts the largest dual violators until a KKT residual tolerance is met, achieving linear memory usage. Experiments show that HELLO outperforms strong baselines by an order of magnitude in runtime while attaining lower transport objectives, and it scales to over a million samples in high‑dimensional settings, supporting various OT variants.

By Wenzhou Xia, Qiaoqiao Ding, Jingwei Liang, Xiaoqun Zhang