The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.
By Fabio De Ponte
arXiv:2607. 07047v1 Announce Type: cross Abstract: Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety.
By Szczepan Konior, Alexandre Quemy, Przemys{\l}aw Klocek, Gr\'egoire Cattan, Bart{\l}omiej Sobieski
arXiv:2606. 18717v1 Announce Type: cross Abstract: Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language models split words by corpus statistics, fragmenting semantically loaded suffixes and -- in the case of WordPiece and rule-based analyzers -- failing to decode their output back to the original text.
By Tolga \c{S}akar
arXiv:2609.08372v2 Announce Type: replace
Abstract: A dense retriever encodes a question as one vector, but the question arrives one word at a time. We read 2,144 held-out headlines from Thu Vien Pha...
By Tran Minh Quan
The article presents a technical manual for an open toolkit designed to measure how transformer language models individuate word meanings across different contexts. It introduces the concept of a "bridge form"—a single word that appears unchanged in multiple domains but with distinct senses—and outlines a full pipeline from specifying these forms to extracting layer-wise representations, computing silhouette-based separation metrics, and visualizing results. The manual details each design choice and its intended methodological safeguards, emphasizing that it serves as a methodological reference rather than reporting empirical findings.
By Jos\'e Luciano Ver\c{c}osa Marques, Frederico Jorge Heitmann, Daniel Omar Perez, Marcelo Vinicius de Paula, T\'arcio Andr\'e dos Santos Barros
The study examines how the geometric placement of synthetic data affects discourse‑pragmatic function classification. Using 410 annotated instances of the word "look" from the British National Corpus, synthetic examples were generated with Llama 3.1 and grouped by cosine distance from real data in RoBERTa space. Six training conditions were compared, showing that examples close to real data (NEAR) yield the largest macro‑F gain, while a distance‑balanced mix gives the highest accuracy, yet none improve AUC.
By Sara Sorahi, Kevin Tang, Reza Kazemian
The paper investigates why supervised AI‑text detectors, specifically a RoBERTa‑based model, can be fooled by subtle changes in language. By applying semantic, structural, and tokenizer‑level perturbations to a large dataset and controlled Mistral‑7B‑Instruct outputs, the authors show that increasing verb diversity makes machine text easier to detect and that detection scores correlate with statistical complexity, leading to a high false‑positive rate on formal human writing. They also evaluate an event‑based latent space detector, finding that paraphrasing and homoglyphs significantly alter extracted event sequences and verbs, yet the detector’s performance remains modest (AUC 0.577).
By Claudiu Creanga, Liviu Dinu
arXiv:2608.21975v1 Announce Type: new
Abstract: This study examines the performance of the state-of-the-art MARBERT model in identifying the lexical/pragmatic category associated with emoji use on X...
By Mohammed Q. Shormani, Yehia A. AlSohbani, Mohammed Q. Shormani
The study investigates whether pretrained transformer models encode functional words—such as pronouns and adverbs—in a way that mirrors human usage. By comparing embeddings of nouns with those of their functional counterparts in both isolated and parallel sentences, the authors find that functional words occupy a central yet distinct position in embedding space and that parallel lexicalized and functional sentences reside in different subspaces. Experiments show that only a mixed training set of functional and lexicalized sentences reveals shared syntactic and semantic structure, whereas training on either type alone fails to capture this parallelism.
By Giuseppe Samo, Vivi Nastase, Paola Merlo
The study investigates how textual neural models degrade when inputs contain noise such as typos, OCR errors, or dropped words. It finds that model performance decline is largely consistent across architectures under word‑level noise but diverges under character‑level noise, a difference attributed to tokenization rather than architecture. By applying a short contrastive training recipe, diverse encoders converge to a common robustness curve, enabling prediction of a model’s noise resilience and the ability to enhance robustness at specific noise scales through targeted training.
By Yefan Tao, Gerald Friedland, Luyang Kong
arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.
By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
The study examines how the geometric placement of synthetic data affects discourse-pragmatic function classification. Using 410 annotated instances of the word "look" and synthetic examples generated by Llama 3.1, the authors partitioned the synthetic data by cosine distance in RoBERTa space and compared six training conditions. While all augmentation conditions improved macro‑F and accuracy over a real‑only baseline, the nearest synthetic examples (NEAR) yielded the largest macro‑F gain (0.113) and a distance‑balanced mix achieved the highest accuracy (0.748), though none improved AUC.