arXiv AI

When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors

arXiv:2606. 32029v1 Announce Type: cross Abstract: While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.

arXiv AI
Sep 24

Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation

The paper proposes using large language models (LLMs) to identify disagreements among models as a way to focus expert effort on revising codebooks for large‑scale text annotation. Three expert feedback methods are evaluated: editing LLM‑generated revisions (Codebook Verifying), answering questions about disagreements (Question Answering), and labeling disagreement cases with rationales (Rationale Labeling). Experiments on tutoring‑session transcripts show that Rationale Labeling achieves the highest LLM‑labeling accuracy (64.9%) compared to the expert‑revised codebook (57.8%), with Question Answering also outperforming the baseline (60.5%).

By Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang
arXiv AI
Sep 25

TabSieve: Explicit In-Table Evidence Selection for Tabular Prediction

TabSieve is a select‑then‑predict framework that explicitly chooses a small set of informative rows from a table as evidence before predicting a missing target. The authors build a large synthetic dataset, TabSieve‑SFT‑40K, and introduce a reinforcement learning method, TAB‑GRPO, to jointly optimize evidence selection and prediction. Experiments on 75 classification and 52 regression tables show consistent performance gains, with TabSieve improving classification by 2.92% and regression by 4.45% over the best baseline while enhancing robustness to noisy context.

By Yongyao Wang, Ziqi Miao, Lu Yang, Haonan Jia, Wenting Yan, Chen Qian, Lijun Li
arXiv AI
Aug 18

Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting

arXiv:2605. 20254v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have shown promising results on NLP tasks, however, their performance on tabular data still needs research attention, because Table Question-Answering (TQA) requires precise cell retrieval and multi-step structured reasoning.

By Amritansh Maurya, Navjot Singh, Mohammed Javed, Omar Moured
arXiv AI
Aug 26

PARTAB: Partition-Aware Reasoning with Structured Evidence for Scalable Table Understanding

PARTAB is a framework that improves large language model reasoning on tables by constructing a structured evidence interface. It represents query‑relevant evidence as semantically coherent, row‑linked table regions and performs hierarchical selection over column groups and row‑level partitions before composing the evidence for answer generation. Evaluations on multiple table reasoning benchmarks show that PARTAB consistently outperforms full‑table prompting and recent methods, achieving strong performance on WikiTableQuestions and TabFact while remaining competitive on numerical reasoning tasks.

By Md Mahadi Hasan Nahid, Davood Rafiei
arXiv AI
Jun 3

Fixing FOLIO and MALLS: Verified Annotations and an LLM-assisted Framework to Focus Human Relabeling

arXiv:2606. 02837v1 Announce Type: cross Abstract: Accurate translation from Natural Language to First-Order Logic (NL-to-FOL) underpins neurosymbolic AI systems and Natural Language Inference (NLI), making the quality of NL-to-FOL benchmarks essential -- yet these datasets have never been rigorously audited.

By Andrea Brunello, Cristian Curaba, Luca Geatti, Michele Mignani, Angelo Montanari, Nicola Saccomanno
arXiv Computation and Language
Sep 3

User Feedback Provides a Unique Signal that LLMs Can not Detect

The paper argues that user feedback from real interactions is a valuable learning signal for Large Language Models (LLMs), contrary to recent claims that it is too noisy to use. By creating synthetic data with a clear ground truth and testing on naturalistic data, the authors show that revisions guided by user feedback fix targeted issues more often than baseline revisions. They further reveal that current evaluation methods bias against feedback‑driven improvements, as judges tend to overlook genuinely corrected responses and favor inferior baselines.

By Shachar Don-Yehiya, Leshem Choshen, Omri Abend
arXiv Machine Learning
Sep 21

Semantic Calibration Prevails Where Token Confidence Fails: Benchmarking Long-Form Scientific QA

The paper presents the first large‑scale benchmark for uncertainty quantification (UQ) calibration in long‑form scientific question answering, evaluating four UQ methods on 685,000 responses from up to 20 large language models across seven datasets. It shows that instruction tuning leads to token‑level probability polarization, undermining token‑level uncertainty signals, while reasoning model families differ in how they handle this effect. Only semantic consistency—consistency of the final answer—provides well‑calibrated outputs, demonstrating that semantic calibration remains robust in multi‑step, dependency‑rich reasoning.

By Philip M\"uller, Nicholas Popovi\v{c}, Michael F\"arber, Peter Steinbach