arXiv Machine Learning By Joshua Nunley

TANGO: Treating Tokens as Operators

Read the original on arXiv Machine Learning →

The paper introduces TANGO, a Token‑Aggregated Nonlinear Gating Operator that blends cross‑token mixing and token‑wise transformation in transformer architectures. By computing a nonlinear gate per source token and averaging these gates for each destination, TANGO forms a source‑conditioned linear operator that improves predictive performance. Experiments on web text, formal mathematics, and code show that full‑prefix TANGO achieves the lowest test negative log‑likelihood across 16 settings, while a narrower variant offers substantial throughput gains with only a modest increase in loss.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Aug 25

TANGO: Token-Aggregated Nonlinear Gating Operators for Natural and Formal Language Modeling

The paper introduces TANGO, a new Transformer variant that replaces the standard self‑attention and feed‑forward sublayers with a single cross‑token gated residual update. Each token generates a SwiGLU gate, and query‑key similarities weight these gates to rescale destination features, yielding a quadratic‑time model. A windowed variant, WANGO, limits the gated interactions to a recent window and uses prefix statistics for older tokens, achieving linear‑time complexity while maintaining competitive performance.

By Joshua Nunley
arXiv Machine Learning
1d ago

Fast Polynomial Transcendentals for LLMs

The paper investigates using short polynomial approximations to accelerate special‑function operations in large language models on NVIDIA Blackwell GPUs. By replacing native sigmoid, tanh, and SiLU with degree‑3 or degree‑4 bfloat16 programs, the authors achieve up to 2.19× speed‑ups in isolated FP16 benchmarks and modest training‑step throughput gains (2.7–8.0%) across four integration tasks. The study also evaluates model behavior, finding negligible training‑loss differences within 100 billion tokens.

By Robert Hu
arXiv AI
Aug 11

Full-bandwidth transformer

arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.

By Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford