arXiv Computation and Language

PunGraph: Retrieval-Enhanced Phonetic-Semantic Graph Reasoning for Pun Understanding

PunGraph is a retrieval‑enhanced knowledge‑graph framework designed to improve pun understanding by combining phonetic and semantic information. It builds a phonetic‑semantic lexical graph using the Unisyn phonetic dictionary, IPA and G2P representations, and WordNet definitions, and uses this graph to retrieve candidate words or senses that constrain large language model reasoning. The authors also introduce WebPun, a new dataset of 5,730 annotated heterographic and homographic puns, and demonstrate that PunGraph consistently boosts the performance of small‑scale LLMs on SemEval‑2017 and WebPun, achieving results competitive with strong proprietary models.

arXiv AI
Jun 3

ReaLM: Residual Quantization Bridging Knowledge Graph Embeddings and Large Language Models

arXiv:2510. 09711v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have recently emerged as a powerful paradigm for Knowledge Graph Completion (KGC), offering strong reasoning and generalization capabilities beyond traditional embedding-based approaches.

By Wenbin Guo, Xin Wang, Jiaoyan Chen, Lingbing Guo, Zhao Li, Zirui Chen
arXiv AI
Jun 15

Knowledge Graph Enhanced Memory-Augmented Retrieval for Long Context Modeling

arXiv:2606. 14047v1 Announce Type: cross Abstract: Long-context language modeling requires not only extending context windows but maintaining coherent understanding of entity states and relationships across thousands of tokens -- a challenge that semantic similarity alone cannot address.

By Ghadir Alselwi, Basem Suleiman, Hao Xue, Shoaib Jameel, Hakim Hacid, Flora D. Salim, Imran Razzak
arXiv AI
Jun 17

findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding

arXiv:2603. 26292v2 Announce Type: replace-cross Abstract: Syllable-level units offer compact and linguistically meaningful representations for spoken language modeling and unsupervised word discovery, but research on syllabification remains fragmented across disparate implementations, datasets, and evaluation protocols.

By H\'ector Javier V\'azquez Mart\'inez
arXiv Computation and Language
Aug 24

Jokes Aside: Measuring the Semantic Distance of Double Meanings

The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.

By Fabio De Ponte
arXiv AI
Sep 2

Do General NLP Embeddings Capture Ontological Reasoning?

The paper introduces AVA, a framework that tests whether general NLP embeddings can differentiate logic-sensitive relational semantics in ontologies and knowledge graphs. AVA uses 171,007 contrastive triplets from 163 ontologies, each containing an ontology statement, a paraphrase, and a hard negative with contradictory meaning. Evaluation of over 25 embedding models shows significant limitations, with the best model achieving only 0.739 triplet accuracy and 0.135 for hard negatives; fine‑tuning helps but does not transfer well to downstream Semantic Web tasks.

By Hamed Babaei Giglou, Jennifer D'Souza, S\"oren Auer