arXiv Machine Learning By Weilun Wang, Wantong Li

Gram-Space: Structure-Preserving Codebook Compression for Memory-Efficient Neuro-Symbolic AI

Read the original on arXiv Machine Learning →

arXiv:2608. 01528v1 Announce Type: new Abstract: Vector symbolic architectures (VSA) are widely used for reasoning in neuro-symbolic (NeSy) AI, yet high-dimensional codebooks often create severe memory bottlenecks that limit scalability and deployment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

Neurosymbolics for Data Engineering: Achieving Long Context Token Reduction Without Finetuning

Neurosymbolics for Data Engineering introduces a neurosymbolic layer that can be added to existing LLM backbones to improve logical reasoning and reduce long‑context token usage. The layer boosts accuracy by an average of 85% on benchmarks such as BIRD‑CRITIC and LiveSQLBench without any task‑specific finetuning or RLHF. It also cuts effective token usage by over 50% and lowers time complexity from O(n²) to roughly O(n) for long‑context tasks.

By Vishvesh Bhat
arXiv Machine Learning
5d ago

Benchmarking Attention for Tabular Foundation Models

The paper introduces a reproducible benchmark for evaluating attention mechanisms in tabular foundation models, focusing on the distinct row and column attention patterns that differ from language model attention. It compares several backends—Torch SDPA, FlashAttention variants, vLLM, and SageAttention—across realistic tabular shapes on A100, H100, and B200 GPUs, revealing that optimal backend choice varies by attention type, hardware, and model specifics. The study finds FlashAttention generally performs best, but CuDNN can outperform it for column attention on longer sequences, while SageAttention excels for large row sequences beyond 16k rows.

By Maximilian Schambach, Clemens Biehl, Sam Thelin
arXiv AI
Sep 1

Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.

By Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki