arXiv Computation and Language
Aug 28

TabuLM: Morphology-Aware Tabular Pre-training for Low-Resource Languages

TabuLM is a new language model pre‑trained on Kinyarwanda tabular data, extending KinyaBERT‑large with row, column, and cell‑type embeddings and a table‑structure attention bias. It introduces two pre‑training objectives—Masked Cell Recovery and Column Type Prediction—and is trained on 172 Rwandan government tables. On the TabQA‑kin benchmark, TabuLM achieves 62.0% exact match, outperforming KinyaBERT‑large and multilingual baselines by significant margins.

By Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre
arXiv Computation and Language
Aug 28

KinyaEmbed: Contrastive Sentence Embeddings for Kinyarwanda via Multi-Stage Curriculum Training

KinyaEmbed is the first sentence‑embedding model specifically designed for Kinyarwanda, built on KinyaBERT‑large and trained through a four‑stage curriculum that incorporates paraphrase pairs, translated MNLI triplets, OPUS‑100 translation pairs, and high‑quality KinyaCOMET pairs. It outperforms existing multilingual embeddings on the SemRel2024‑rw benchmark, achieving a Spearman ψ of 0.7298, and introduces the Wiki‑RW‑STS benchmark of 300 contamination‑free Kinyarwanda sentence pairs. All model checkpoints, filtered pairs, and the new benchmark are publicly released.

By Ireddi Rakshitha, Devavarapu Yashwanth, Ntakirutimana Pierre
arXiv Machine Learning
5d ago

Benchmarking Attention for Tabular Foundation Models

The paper introduces a reproducible benchmark for evaluating attention mechanisms in tabular foundation models, focusing on the distinct row and column attention patterns that differ from language model attention. It compares several backends—Torch SDPA, FlashAttention variants, vLLM, and SageAttention—across realistic tabular shapes on A100, H100, and B200 GPUs, revealing that optimal backend choice varies by attention type, hardware, and model specifics. The study finds FlashAttention generally performs best, but CuDNN can outperform it for column attention on longer sequences, while SageAttention excels for large row sequences beyond 16k rows.

By Maximilian Schambach, Clemens Biehl, Sam Thelin
arXiv AI
Aug 18

Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting

arXiv:2605. 20254v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have shown promising results on NLP tasks, however, their performance on tabular data still needs research attention, because Table Question-Answering (TQA) requires precise cell retrieval and multi-step structured reasoning.

By Amritansh Maurya, Navjot Singh, Mohammed Javed, Omar Moured