WildSEEK is a new dataset of 3,000 real user information‑seeking queries, manually annotated for risk‑sensitive domains and whether the query is factoid or analytical. The accompanying evaluation framework tests LLM responses against four failure criteria—sycophantic behavior, overreliance, a default US‑centric perspective, and poor handling of vulnerable populations—finding higher failure rates for analytical queries. The authors also train classifiers on WildSEEK to analyze over 1.8 million realistic queries, revealing that more than a third are high‑risk and often analytical.
By Tanise Ceron, Joachim Baumann, Elisa Bassignana, Berat Cabuk, Dirk Hovy, Debora Nozza
arXiv:2609.38256v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to answer questions about politically contentious issues, yet evaluations typically treat a model's...
By Olivia Macmillan-Scott, Michael Jacobs, Nils Metternich, Mirco Musolesi
PolERo presents a new dataset of 3,574 Romanian question‑answer pairs from presidential transcripts, annotated for political evasion using a two‑level taxonomy of response clarity and fine‑grained evasion strategies. The study evaluates various classification methods—including TF‑IDF baselines, fine‑tuned encoders, a sliding‑window encoder, and zero/few‑shot LLM prompting—under matched conditions. Cross‑lingual transfer experiments via joint bilingual training and machine‑translation augmentation reveal that fine‑tuned encoders perform competitively, transfer is asymmetric, and ambivalent evasion categories with pragmatic cues remain the most challenging across all models.
By Gabriel Stefan, Sergiu Nisioi
arXiv:2608.29198v1 Announce Type: new
Abstract: As Large Language Models (LLMs) increasingly encourage users to disclose personal profiles for tailored assistance, measuring their political alignment...
By Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen, Hen-Hsen Huang, I-Chen Wu
The paper investigates covert dialect bias in large language models (LLMs) by analyzing how internal probability distributions associate different English varieties—Standard American English, African American Vernacular English, Nigerian Standard English, and Nigerian Pidgin—with housing-related adjectives. Using 260 meaning‑matched sentence quadruples and log‑probability scoring across ten open‑weight LLMs, the study finds that AAVE and NP are consistently linked to more negative adjectives than SAE, with NP experiencing the greatest penalty. The bias varies by context and stereotype cluster, and Nigerian Standard English shows a context‑dependent shift, being favored in formal tenant screening but penalized in more socially proximate scenarios.
By Chowdhury Mohammad Abdullah, Rita Orji
arXiv:2609.00319v1 Announce Type: cross
Abstract: Online health information seeking is shifting from keyword search, where users consider a ranked list of links, to conversational systems that compos...
By Phuong Anh Nguyen, Jill Noorily, Matthew Flathers, Haruka Notsu, Laura Ospina-Pinillos, Tommy Nguyen, Samantha Clark, Aoife Keane, Grace Thompson, John Torous
arXiv:2610.00606v1 Announce Type: cross
Abstract: Large Language Models increasingly serve as interfaces for knowledge-intensive information seeking tasks across languages by synthesizing multilingua...
By Dayeon Ki, Ruochen Zhang, Silviu Cucerzan, Ryen W. White, Ning Gao
The paper examines how large language model (LLM) services route queries to models of varying size based on a cheap complexity estimate. It finds that this routing is not register neutral: queries written in non‑standard English registers (e.g., African American English or second‑language English) are systematically assigned to lower‑capacity models because they appear shorter due to omitted function words. Experiments on 37,704 learner sentence pairs and a controlled corpus show that this bias leads to significantly lower accuracy across all model tiers, including the highest‑capacity cloud models, while the routing decision itself adds little marginal cost.
By Simran Koul
arXiv:2603. 13891v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring.
By Petter T\"ornberg
arXiv:2607. 14888v1 Announce Type: cross Abstract: Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains.
By Robert Graham, Edward Stevinson, Yariv Barsheshat
The paper introduces the concept of Language Specific Knowledge (LSK), showing that multilingual language models can answer certain queries better when prompted in a language other than English, sometimes even in low‑resource languages. It defines a language‑selection problem and presents several baseline methods, including the authors’ LSKExtractor, to empirically demonstrate that choosing the optimal language can improve question‑answering performance across datasets covering cultural and social norms. Experiments reveal non‑intuitive mappings, such as Gemma models excelling on Chinese and Middle Eastern topics in Spanish and Qwen models performing best on authority and responsibility queries in Arabic and Chinese.
By Ishika Agarwal, Nimet Beyza Bozdag, Dilek Hakkani-T\"ur
The study investigates why language models exhibit systematic performance gaps across English dialects, a phenomenon termed the "dialect tax." Using parallel dialect corpora that preserve meaning while altering surface form, the authors confirm that models treat Standard American English and dialectal texts as semantically equivalent, yet find representational disparities that persist through tokenization, pre‑training, post‑training, and inference. Even a character‑level tokenizer does not eliminate input/output asymmetries or accuracy gaps, and dialect pairs produce more divergent gradient updates than unrelated Standard texts, indicating that dialectal content is harder for models to learn.
By Elle