arXiv Machine Learning

Permutation Learning with Only N Parameters: From SoftSort to Self-Organizing Gaussians

arXiv:2503. 13051v3 Announce Type: replace Abstract: Sorting and permutation learning are key concepts in optimization and machine learning, especially when organizing high-dimensional data into meaningful spatial layouts.

arXiv Machine Learning
Aug 27

Deep greedy unfolding: Sorting out argsorting in greedy sparse recovery algorithms

The paper introduces Soft-OMP and Soft-IHT, permutation‑based variants of Orthogonal Matching Pursuit and Iterative Hard Thresholding that replace the non‑differentiable argsort with continuous soft‑sort operators. These differentiable algorithms enable the construction of fully trainable neural network architectures—OMP‑Net and IHT‑Net—while preserving the core greedy sparse recovery logic. The authors show both theoretically and numerically that the soft variants approximate their hard counterparts with controllable accuracy and can be extended to structured sparse recovery by learning structure‑aware weights.

By Sina Mohammad-Taheri, Matthew J. Colbrook, Simone Brugiapaglia
arXiv AI
2d ago

GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning

GPart introduces a new parameter‑efficient fine‑tuning technique that directly maps a low‑dimensional trainable vector into the full weight space using a sparse, isometric partition matrix. Unlike LoRA, GPart eliminates the bilinear reconstruction step, preserving exact end‑to‑end isometry and reducing the checkpoint to just the vector and a random seed. Experiments across NLP, vision, and reasoning tasks show that GPart matches or surpasses existing PEFT methods while using far fewer parameters and offering a simpler, more tractable parameterization.

By Paolo Mandica, Micha{\l} Brzozowski, Zuzanna Dubanowska, Neo Christopher Chung
arXiv Machine Learning
Sep 7

Inducing Permutation Invariant Priors in Bayesian Optimization for Carbon Capture and Storage Applications

The paper introduces a new Gaussian Process kernel, GP‑Perm, that incorporates permutation invariance for Bayesian Optimization tasks involving well placement in Carbon Capture and Storage (CCS) projects. It compares sets via a stable divergence between their empirical representations and can be combined with standard kernels for additional inputs. The authors also explore a Deep Kernel Learning model using a Deep Sets architecture as a learned invariant baseline, evaluating both approaches on eight use cases, including seven synthetic benchmarks and a realistic CCS case study in the Johansen formation.

By Sofianos Panagiotis Fotias, Vassilis Gaganis
arXiv Machine Learning
Sep 18

A Table-Free Index for Tapered Memoization Grids: Compact Out-of-Core Evaluation of Functions of Sorted Arguments

The paper presents a table‑free index for tapered memoization grids, enabling compact out‑of‑core evaluation of functions that depend on sorted arguments. By showing that the grid’s key set corresponds to multiset combinations, the authors derive a closed‑form O(d) ranking and unranking scheme that removes the need for large preprocessing tables and allows order‑free parallel construction. The resulting values‑only flat array uses significantly less memory than hash‑map memoization, offers faster query times once cache limits are exceeded, and remains operable with memory‑mapped storage beyond RAM.

By Tamal Maharaj
arXiv AI
4d ago

ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs

ThinQuant introduces efficient rotation learning for low‑bit weight and activation quantization of large language models by reducing calibration data through a geometric selection of activations and solving a lower‑dimensional optimization problem via an ADMM algorithm. The method achieves comparable or better quantization performance with dramatically fewer calibration points, completing rotation calibration for Llama‑3‑70B in under 12 minutes and for Llama‑3.1‑405B in just over 2 hours on a single GPU. ThinQuant outperforms existing gradient‑free approaches such as DartQuant and gradient‑based SpinQuant in both speed and perplexity metrics on WikiText‑2.

By Mehdi Makni, Ryan Lucas, Rahul Mazumder