Hugging Face Trending Papers

Learning Moral Diversity: Modelling Individual Perspectives in Moral Classification of Texts

Understanding moral values in social media text offers insight into moral judgement formation, and supervised NLP models trained on crowdsourced data have achieved strong classification performance. However, most approaches simplify the problem by aggregating multiple annotators' labels into a single "ground truth", overlooking the inherent subjectivity of the task.

arXiv AI
Sep 2

ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation

ValueGraph is a graph pre‑training framework that incorporates automatically inferred moral‑value signals as soft constraints to learn contextualized user representations. By constructing post‑reply graphs, it captures both semantic and structural information and aligns users through contrastive and clustering objectives based on relative value similarity. Experiments on stance detection and Twitter bot detection demonstrate consistent improvements over strong text‑based, graph‑based, and LLM baselines, underscoring the benefit of value‑signal guidance for socially informed user modeling.

By Yitong Han, Wei Gao, Yi Zhao, Prasanta Bhattacharya, Fengzhu Zeng, Mohammad Amanlou
arXiv Computation and Language
Sep 4

Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour

The paper introduces CHARM, a lightweight fine‑tuned language model framework for detecting moral foundations in text. CHARM combines MAC cross‑attention, rationale alignment, and hate‑speech modulation to operationalize distinct psychological constructs, achieving up to 15.3% higher AUC in‑domain and outperforming supervised baselines on all out‑of‑domain datasets. The authors demonstrate CHARM’s scalability by applying it to large‑scale COVID‑19 Twitter data, revealing a strong link between moral value alignment and online endorsement behavior.

By Huixiang Fu, Marian-Andrei Rizoiu
arXiv AI
Sep 4

Representational alignment yields generalizable safety in language models

The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.

By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu
arXiv AI
Sep 21

Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation

The paper introduces a multilingual story moral generation task to evaluate cultural alignment in large language models. Using a dataset of human-written story morals from 14 language‑culture pairs, the authors compare model outputs to human interpretations through semantic similarity, a preference survey, and value categorization. They find that advanced models like GPT‑4o and Gemini produce morally similar and preferred responses but show less cross‑linguistic variation, focusing on a narrower set of shared values, indicating a limitation in capturing the diversity of human narrative understanding.

By Sophie Wu, Andrew Piper
arXiv Machine Learning
Aug 24

MIL-BERT: Classification of Arbitrarily Large Text with Performance and Explanatory Guarantees

MIL-BERT is a neural network algorithm that classifies large texts by selecting relevant excerpts, inspired by multiple instance learning. It scales to samples with nearly 1 million tokens and has been evaluated on seven datasets, achieving state‑of‑the‑art results on three long‑text tasks such as political bias detection, trigger warning identification, and author demographic inference. The model also generalizes from weakly‑labeled text bags to accurately classify smaller instances.

By John Cadigan, Dayne Freitag, Eric Yeh
Hugging Face Trending Papers
Jun 10

A Resource for Enthymeme Detection in Controversial Political Discourse

Enthymemes, arguments with unstated premises or conclusions, are pervasive in persuasive discourse, yet their annotation remains notoriously subjective. We present a resource of 1,482 tweets from politically controversial discourse, annotated by five annotators for the presence of enthymemes and their argument structure, designed to study label variation.