arXiv Computation and Language

Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure

The paper audits a multilingual affective generation benchmark that uses emoji summaries for Bangla, English, and Hindi sentences. It finds that the benchmark’s conclusions are largely artifacts of the measurement instrument, with no system significantly outperforming another when annotators are treated as random factors. The study shows that annotator identity and output length drive most variance, and proposes a new stable metric called emoji‑affect decodability.

arXiv AI
Sep 3

EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision

The paper introduces EmoStance, a method for controlling the affective orientation of empathetic responses by leveraging weak supervision from emoji distributions. It builds the EmojiDialogue dataset, extending EmpatheticDialogues with emoji votes and confidence scores, and uses a frozen instruction‑tuned LLM steered by continuous prefix embeddings to generate responses that align with the listener’s stance. In blind pairwise evaluations, EmoStance achieves a 62.2% decisive win rate, notably improving contextual specificity and perceived responsiveness compared to baselines.

By Ziyuan Jin, Yuxuan Ge, Zheng Tian
Hugging Face Trending Papers
Sep 2

EmoStance: Response-Side Affective-Orientation Control for Empathetic Response Generation via Emoji Weak Supervision

The paper introduces EmoStance, a method for controlling the affective orientation of empathetic responses in dialogue systems. It leverages weak supervision from multi‑annotator emoji distributions to create a latent control space that approximates listener stance, and uses a frozen instruction‑tuned LLM steered by continuous prefix embeddings. Evaluation shows a 62.2% decisive win rate over baselines, especially in contextual specificity and perceived responsiveness.

arXiv Computation and Language
Aug 24

Jokes Aside: Measuring the Semantic Distance of Double Meanings

The paper investigates how semantic distance and ambiguity contribute to joke humor by revisiting and extending metrics from prior work. It introduces a new symmetry metric—measuring how close the ambiguous element Z is to both X and Y—and evaluates it using two embedding models on three joke datasets, including expanded versions with paired ambiguous sentences. Although models based on these metrics performed poorly in predicting humor ratings, the symmetry metric consistently correlated with higher-rated jokes, hinting it captures a key, though not sole, property of humor.

By Fabio De Ponte
arXiv AI
Aug 5

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

arXiv:2608. 03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable.

By Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov
arXiv Computer Vision
Aug 27

Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces

The study audits six vision‑language models (VLMs) to assess whether they consistently encode affective qualities of 3D shapes, using Kansei adjective pairs as affective axes. Across ten ShapeNet categories, models show moderate agreement (mean rank correlation 0.36) that is lower than geometric controls but higher than unrelated adjective pairs, with convergence varying widely by category and axis. The authors demonstrate how this audit informs a UI prototype that selectively exposes Kansei descriptors for generative design interfaces.

By Luca Bux, Thiago Rios, Ingo Scholtes, Stefan Menzel
arXiv AI
Sep 17

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

The paper investigates why large vision‑language models sometimes misclassify harmful memes, attributing failures to either missing internal evidence or poor routing of evidence to the output. Using sparse autoencoders, role‑conditioned probes, and causal interventions on Gemma‑3 and Qwen3.5, the authors show that sparse readouts consistently outperform native predictions across six harmful content benchmarks, revealing a readout gap that is largely due to routing rather than representation. The study also demonstrates that calibration‑only routing recovers most of the performance gap and that the issue persists across languages and is not solely driven by OCR signals.

By Girish A. Koushik, Diptesh Kanojia, Helen Treharne
arXiv AI
Aug 20

Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation

This study investigates whether large language models (LLMs) exhibit different safety vulnerabilities when prompted with emojis instead of plain text. By testing 50 emoji‑augmented prompts on four open‑source LLMs—Mistral 7B, Qwen 2 7B, Gemma 2 9B, and Llama 3 8B—the authors found varying success rates of unsafe responses: Gemma 2 9B and Mistral 7B each had a 10% success rate, Llama 3 8B had 6%, while Qwen 2 7B was fully resistant. A chi‑square test confirmed significant differences in the outcome distributions, suggesting that input representation can markedly affect model robustness.

By M P V S Gopinadh
arXiv Computation and Language
Sep 7

A Systematic Comparison of Multilingual Interpretability Methods Reveals Anisotropy-Driven Failures

The paper evaluates four metrics—CKA, ANC, GMM dominance per token, and ILO—used to measure cross‑lingual representation sharing in multilingual language models. Across 21 models ranging from 125 M to 14 B parameters, the metrics disagree, and the authors attribute this to anisotropy, where representations cluster in a narrow embedding cone. Only ILO shows a strong, robust correlation with cross‑lingual transfer performance (Spearman’s ρ = 0.90) after controlling for model size, family, and task variation, leading the authors to recommend ILO as the primary metric alongside anisotropy diagnostics.

By Oskar Holmstr\"om, Marcel Bollmann, Marco Kuhlmann
arXiv Machine Learning
Aug 28

Vowel Signs Are Not Letters: A Pre-tokenization Ceiling on Multilingual Tokenizer Fertility

The paper shows that HuggingFace’s ByteLevel pre‑tokenizer, which treats a word as a sequence of Unicode letters, splits abugida scripts at every vowel sign, creating a training‑free lower bound on tokenizer fertility. Across 26 languages, all 17 abugidas exhibit increased token counts (up to 9×), while Latin, Cyrillic, Hangul, and Han remain unchanged. The authors demonstrate that correcting the character class reduces Nepali token counts, improves model performance, and that this issue is widespread in popular HuggingFace models.

By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun