arXiv AI By Jiajun Bao, Zihao Qi, Toni J. B. Liu, Gurbir Arora, Rapha\"el Sarfati, Nicolas Boull\'e, Christopher J. Earls

A Graph Signal Processing Perspective on Numerical Sequence Representations in LLM In-Context Learning

Read the original on arXiv AI →

arXiv:2608. 03015v1 Announce Type: cross Abstract: Pretrained large language models (LLMs) have demonstrated in-context learning (ICL) capabilities for numerical inference over sequences serialized as text.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
4d ago

Lost in Tokenization: Fundamental Trade-offs in Graph Tokenization for Transformers

The paper investigates how the choice of graph tokenization affects transformer expressivity. It analyzes three tokenization families—spectral, random‑walk, and adjacency—showing that each induces different depth requirements and that some tokenizations are inherently lossy or ill‑conditioned for certain tasks. The authors prove lower bounds and impossibility results for converting between tokenizations and validate these findings with experiments on synthetic and real‑world data.

By Maya Bechler-Speicher, Gilad Yehudai, Gil Harari, Clayton Sanford, Amir Globerson, Joan Bruna