The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.
By Fabio De Ponte
The study examines Urdu light verbs, which add schematic event meaning while staying lexically linked to their main verbs. Using contextual embeddings from UrduBERT, DunbaaBERT, and multilingual BERT on 1,126 sentences, the authors find significant representational separation between main and light uses across all verb–model pairs, yet main and light uses of the same lemma remain closer than mismatched pairs. In a seven‑way prediction task limited to light uses, UrduBERT achieves 0.866 accuracy and 0.852 macro‑F1, and maintains 0.782 accuracy when tested on unseen preceding forms, demonstrating generalization beyond local verb combinations.
By Farah Adeeba, Miriam Butt
The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.
By Salima Lamsiyah, Ruslan Mitkov
arXiv:2607. 03981v1 Announce Type: cross Abstract: Memes have become influential communication tools on social media, combining viral visuals with concise messaging to convey impactful ideas.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md. Mahfuzur Rahman, Pronay Debnath, Asif Iftekher Fahim, Faisal Muhammad Shah
arXiv:2606. 07522v1 Announce Type: cross Abstract: We propose an unsupervised method of resolving slang, unique entities, and folklore from online communities by isolating words in the lexicon that have the highest magnitude of semantic shift.
By Julia Kruk, Sanchita Porwal, Amitrajit Bhattacharjee, Mansi Phute
arXiv:2603. 18482v2 Announce Type: replace-cross Abstract: Standard decoding strategies for text generation, including top-$k$, nucleus sampling, and contrastive search, select tokens based on likelihood, restricting outputs to high-probability regions.
By Esteban Garces Arias, Nurzhan Sapargali, Christian Heumann, Matthias A{\ss}enmacher
arXiv:2608. 01935v1 Announce Type: cross Abstract: Prior work in Ancient Greek NLP relies on corpora that do not disambiguate the phonemic vowel length of alpha, iota, and ypsilon, together known as the dichrona.
By Albin Th\"orn Cleland, Eric Cullhed
The paper introduces TSS (Triple-Stream Stress probe), a diagnostic framework that splits text into lexical, morpho-syntactic, and psycholinguistic style channels to analyze mental health NLP classifiers. Across four English datasets, TSS uncovers a lexical interference effect where adding lexical features harms performance on human-labeled data but not on auto-labeled data, and proposes the Degree of Divergence (DoD) statistic to audit label-source bias. The study demonstrates that style features largely remain effective even after masking content words, emphasizing that shortcut learning is label-source specific rather than clinically relevant.
By Moustafa Yehia Hassan
arXiv:2607. 23675v1 Announce Type: cross Abstract: The embedding of word sequences into continuous vector spaces has been one of the most important developments in Natural Language Processing in the recent past.
By Salomon Kabongo
arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.
By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv:2606. 20255v1 Announce Type: cross Abstract: We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent.
By Celestine Achi
Statistical regularities are routinely offered as evidence that undeciphered sign systems encode language; the Indus script debate is the canonical example. Any such inference rests on specificity: the reported outcome must be unusual among plausible structured non-languages.