arXiv AI

An empirical investigation into the properties of standard word embeddings

arXiv:2607. 23675v1 Announce Type: cross Abstract: The embedding of word sequences into continuous vector spaces has been one of the most important developments in Natural Language Processing in the recent past.

arXiv Computation and Language
Aug 24

Jokes Aside: Measuring the Semantic Distance of Double Meanings

The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.

By Fabio De Ponte
arXiv AI
Jun 18

Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish

arXiv:2606. 18717v1 Announce Type: cross Abstract: Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language models split words by corpus statistics, fragmenting semantically loaded suffixes and -- in the case of WordPiece and rule-based analyzers -- failing to decode their output back to the original text.

By Tolga \c{S}akar
arXiv AI
Sep 7

Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models

The article presents a technical manual for an open toolkit designed to measure how transformer language models individuate word meanings across different contexts. It introduces the concept of a "bridge form"—a single word that appears unchanged in multiple domains but with distinct senses—and outlines a full pipeline from specifying these forms to extracting layer-wise representations, computing silhouette-based separation metrics, and visualizing results. The manual details each design choice and its intended methodological safeguards, emphasizing that it serves as a methodological reference rather than reporting empirical findings.

By Jos\'e Luciano Ver\c{c}osa Marques, Frederico Jorge Heitmann, Daniel Omar Perez, Marcelo Vinicius de Paula, T\'arcio Andr\'e dos Santos Barros
arXiv Computation and Language
Sep 4

The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification

The study examines how the geometric placement of synthetic data affects discourse‑pragmatic function classification. Using 410 annotated instances of the word "look" from the British National Corpus, synthetic examples were generated with Llama 3.1 and grouped by cosine distance from real data in RoBERTa space. Six training conditions were compared, showing that examples close to real data (NEAR) yield the largest macro‑F gain, while a distance‑balanced mix gives the highest accuracy, yet none improve AUC.

By Sara Sorahi, Kevin Tang, Reza Kazemian
arXiv Computation and Language
4d ago

Ontological Instability and Statistical Amplification: The Paradox of "Humanizing" LLM-Generated Text

The paper investigates why supervised AI‑text detectors, specifically a RoBERTa‑based model, can be fooled by subtle changes in language. By applying semantic, structural, and tokenizer‑level perturbations to a large dataset and controlled Mistral‑7B‑Instruct outputs, the authors show that increasing verb diversity makes machine text easier to detect and that detection scores correlate with statistical complexity, leading to a high false‑positive rate on formal human writing. They also evaluate an event‑based latent space detector, finding that paraphrasing and homoglyphs significantly alter extracted event sequences and verbs, yet the detector’s performance remains modest (AUC 0.577).

By Claudiu Creanga, Liviu Dinu
arXiv Computation and Language
Sep 18

Generalization through Lexical Abstraction in Transformer Models: The Case of Functional Words

The study investigates whether pretrained transformer models encode functional words—such as pronouns and adverbs—in a way that mirrors human usage. By comparing embeddings of nouns with those of their functional counterparts in both isolated and parallel sentences, the authors find that functional words occupy a central yet distinct position in embedding space and that parallel lexicalized and functional sentences reside in different subspaces. Experiments show that only a mixed training set of functional and lexicalized sentences reveals shared syntactic and semantic structure, whereas training on either type alone fails to capture this parallelism.

By Giuseppe Samo, Vivi Nastase, Paola Merlo
arXiv Machine Learning
Aug 28

When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models

The study investigates how textual neural models degrade when inputs contain noise such as typos, OCR errors, or dropped words. It finds that model performance decline is largely consistent across architectures under word‑level noise but diverges under character‑level noise, a difference attributed to tokenization rather than architecture. By applying a short contrastive training recipe, diverse encoders converge to a common robustness curve, enabling prediction of a model’s noise resilience and the ability to enhance robustness at specific noise scales through targeted training.

By Yefan Tao, Gerald Friedland, Luyang Kong
arXiv Machine Learning
Jul 28

BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis

arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.

By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
Hugging Face Trending Papers
Sep 3

The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification

The study examines how the geometric placement of synthetic data affects discourse-pragmatic function classification. Using 410 annotated instances of the word "look" and synthetic examples generated by Llama 3.1, the authors partitioned the synthetic data by cosine distance in RoBERTa space and compared six training conditions. While all augmentation conditions improved macro‑F and accuracy over a real‑only baseline, the nearest synthetic examples (NEAR) yielded the largest macro‑F gain (0.113) and a distance‑balanced mix achieved the highest accuracy (0.748), though none improved AUC.