arXiv Machine Learning By Han-yu Wang

Persistent Priors, Preserved Targets: A Stroop-Style Paradigm for Lexical Override

Read the original on arXiv Machine Learning →

arXiv:2606. 07555v5 Announce Type: replace-cross Abstract: Local definitions can assign a familiar word a temporary meaning while its usual associations remain useful elsewhere.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computation and Language
Sep 24

MWE-ECL: Recoverable Long-Range Context Does Not Always Override Local Lexical Priors

The paper introduces MWE‑ECL, a bilingual diagnostic framework that tests whether distant discourse anchors can override local lexical priors in multi‑word expression interpretation. It evaluates models on a 0‑128K context grid, finding that while retrieval of anchors is near perfect, the ability to change locally preferred readings varies, especially when the model’s default conflicts with the anchor. The study shows that explicit recoverability does not always translate into behavioral influence, with gaps differing across models and languages.

By Wei He, Aline Villavicencio, Rodrigo Wilkens, Zhenyun Deng
arXiv Computation and Language
Sep 18

WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution

The paper investigates why Word-in-Context (WiC) remains difficult for language models, suggesting that the lack of an explicit sense inventory contributes to the challenge. By evaluating open LLMs on both WiC and traditional Word Sense Disambiguation (WSD) tasks, the authors find that providing candidate senses—akin to WSD—consistently improves WiC performance. Human evaluation indicates that many WiC errors stem from label ambiguity or mismatched sense boundaries, with models often over‑discriminating senses and making overly fine‑grained distinctions.

By Yi Zhou, Kiamehr Rezaee, Danushka Bollegala, Mohammad Taher Pilehvar, Jose Camacho-Collados
arXiv AI
Sep 15

Domain-Specific Jargon in Large Language Models: A Comparative Analysis between General-Purpose and Specialist Models

The paper investigates how large language models handle domain-specific jargon, comparing a general-purpose Llama‑3.1 with a version fine‑tuned on medical data. Two new medical jargon benchmarks reveal that the general model actually outperforms the fine‑tuned variant, and interpretability tools show the fine‑tuned model over‑emphasizes a few components linked to jargon predictions. Reweighting these components narrows the performance gap, and some jargon‑sensitive components also aid materials‑science tasks, indicating a partially domain‑agnostic representation of specialized terminology.

By Darin Keng, Zhewei Sun