The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.
By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
arXiv:2609.01383v1 Announce Type: new
Abstract: Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic...
By Maeve Hutchinson, Syed Mahbubul Huq, Mohammad Albinhassan, Radu Jianu, Aidan Slingsby, Pranava Madhyastha
arXiv:2609.00454v1 Announce Type: new
Abstract: Pretrained transformer-based language models achieve strong performance across a wide range of NLP tasks but remain limited in encoding geo-locational...
By Gokul Srinivasagan, Munir Georges
arXiv:2508.11857v3 Announce Type: replace-cross
Abstract: Tokenization remains a persistent bottleneck in language modeling, especially when vocabulary learning is limited by whitespace boundaries. W...
By Andrei-Valentin T\u{a}nase, Elena Pelican
arXiv:2609.00513v1 Announce Type: new
Abstract: Retrieval-Augmented Generation (RAG) mitigates large language models (LLMs) hallucinations, yet conventional dense retrieval struggles with the complex...
By Siyuan Zhang, Hanchen Wang, Dong Wen, Ying Zhang, Wenjie Zhang
DiscoTrace is a method that identifies rhetorical strategies used by answerers to information‑seeking questions by representing answers as sequences of question‑related discourse acts paired with interpretations of the original question, annotated on top of rhetorical structure theory parses. When applied to answers from nine different communities, DiscoTrace reveals that these communities exhibit diverse preferences for answer construction, whereas large language models (LLMs) lack such rhetorical diversity even when prompted to follow specific community guidelines. Additionally, LLMs tend to adopt a breadth‑oriented approach, addressing interpretations of questions that human answerers often ignore, highlighting a systematic difference in how LLMs and humans respond to information needs.
By Neha Srikanth, Jordan Boyd-Graber, Rachel Rudinger
arXiv:2609.01604v1 Announce Type: cross
Abstract: LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the int...
By Himil Vasava, Ming Jiang
The study investigates whether attention weights in Vision‑Language Models (VLMs) accurately reflect model reasoning for visual inputs. Using causal perturbation analysis, it identifies three distinct processing modes—Faithful‑Sufficient, Faithful‑Distributed, and Non‑Focal—indicating heterogeneous visual attention faithfulness. The research also shows that human‑annotated ground‑truth regions align with model attention in only about 60% of cases, highlighting a systematic divergence between model visual reliance and human intuition across VQA, document, and chart tasks.
By Xurui Song, Weishi Wang, Zhongqi Yue, Kuluhan Binici, Tao Bai, Hongxin Shao, Daniel Dahlmeier, Jun Luo
The paper introduces WHALE, a method that alternates between updating a language model’s weights and searching for a better harness (the code that manages context and control flow). By iteratively fine‑tuning the model under the current harness and then optimizing the harness under the updated model, WHALE improves performance across search QA, math reasoning, and chess puzzles, outperforming weight‑only, harness‑only, and Fast‑Slow Training by 4.15–24.38 percentage points in mean@8 accuracy. The approach uses either fixed phase lengths or an adaptive patience rule to decide when to switch phases, and the authors provide code on GitHub.
By Haechan Kim, Yoonho Lee, Gisang Lee, Chelsea Finn, Kangwook Lee
LLM-based evaluators of natural language generation (NLG) quality are widely deployed as scoring tools and as automated training signals, yet the internal procedure by which they assign a rating remai...
Large language models (LLMs) struggle to classify text into taxonomies with many semantically similar labels, as the distinctions are domain-specific and not captured by pre-training. To handle large...
Compressed context is usually carried as human-readable text or as rendered images that must be decoded, even when its consumer is a language model. We introduce LatentPress, which writes conversation...
Grounded question answering systems should answer only when the supplied evidence supports the answer. In multi-hop QA, this requirement is difficult because partial evidence can make an unsupported a...
Vision Language Models have demonstrated remarkable proficiency in interpreting static visual artifacts, but modern data analysis is inherently dynamic, requiring the active interrogation of interacti...
Reliable evaluation of open-ended question answering remains a bottleneck for measuring answer correctness of modern LLMs. Unlike multiple-choice tasks, free-form answers may be correct in many surfac...
Tables are ubiquitous across diverse domains, yet reasoning over them remains a significant challenge for modern large language models (LLMs). Current approaches typically linearize tables into sequen...
Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is cri...
Mixed-format medical visual question answering (VQA) requires stable option selection and machine-readable free-text output. The two formats fail differently: multiple-choice predictions can change wi...
Pixel-space diffusion has recently emerged as a promising direction for high-fidelity image generation by modeling images directly in the original pixel domain. However, pixel-space diffusion is compu...
arXiv:2601.02933v4 Announce Type: replace
Abstract: Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is n...
By Vil\'em Zouhar, Tom Kocmi