Hugging Face Trending Papers

Learning Moral Diversity: Modelling Individual Perspectives in Moral Classification of Texts

Read the original on Hugging Face Trending Papers →

Understanding moral values in social media text offers insight into moral judgement formation, and supervised NLP models trained on crowdsourced data have achieved strong classification performance. However, most approaches simplify the problem by aggregating multiple annotators' labels into a single "ground truth", overlooking the inherent subjectivity of the task.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 2

ValueGraph: Value-Signal Guided Graph Pre-training for Contextualized User Representation

ValueGraph is a graph pre‑training framework that incorporates automatically inferred moral‑value signals as soft constraints to learn contextualized user representations. By constructing post‑reply graphs, it captures both semantic and structural information and aligns users through contrastive and clustering objectives based on relative value similarity. Experiments on stance detection and Twitter bot detection demonstrate consistent improvements over strong text‑based, graph‑based, and LLM baselines, underscoring the benefit of value‑signal guidance for socially informed user modeling.

By Yitong Han, Wei Gao, Yi Zhao, Prasanta Bhattacharya, Fengzhu Zeng, Mohammad Amanlou
arXiv Computation and Language
Sep 4

Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour

The paper introduces CHARM, a lightweight fine‑tuned language model framework for detecting moral foundations in text. CHARM combines MAC cross‑attention, rationale alignment, and hate‑speech modulation to operationalize distinct psychological constructs, achieving up to 15.3% higher AUC in‑domain and outperforming supervised baselines on all out‑of‑domain datasets. The authors demonstrate CHARM’s scalability by applying it to large‑scale COVID‑19 Twitter data, revealing a strong link between moral value alignment and online endorsement behavior.

By Huixiang Fu, Marian-Andrei Rizoiu
arXiv AI
Sep 4

Representational alignment yields generalizable safety in language models

The paper argues that aligning large language models (LLMs) at the level of latent representations—specifically by matching their internal categorization of moral concepts to human prototype-based judgments—improves safety. Current alignment methods that focus on observable responses fail to preserve fine-grained moral categorization, leaving models vulnerable to adversarial rephrasings. By optimizing representational similarity, the authors demonstrate that LLMs can maintain more robust moral categorization and exhibit better adversarial robustness across multiple benchmarks and model sizes.

By Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu