arXiv AI

Linguistic Monoculture in LLM-Assisted Language Use

arXiv:2607. 27134v1 Announce Type: new Abstract: Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text.

arXiv Computation and Language
Sep 2

Evaluating Style-Personalized Text Generation: Challenges and Directions

The paper "Evaluating Style-Personalized Text Generation: Challenges and Directions" examines the difficulties of assessing text that is tailored to individual users’ styles. It critiques common metrics such as BLEU, embeddings, and LLM-as-judges, and introduces a style discrimination benchmark covering domain discrimination, authorship attribution, and LLM-generated personalized versus non-personalized discrimination across eight writing tasks. The study finds that ensembles of diverse evaluation metrics outperform single-evaluator approaches and offers guidance for reliable assessment of style-personalized generation.

By Anubhav Jangra, Bahareh Sarrafzadeh, Silviu Cucerzan, Adrian de Wynter, Sujay Kumar Jauhar
arXiv AI
Jun 12

Authorship Attribution in Multilingual Machine-Generated Texts

arXiv:2508. 01656v2 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) have reached human-like fluency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly difficult.

By Lucio La Cava, Dominik Macko, R\'obert M\'oro, Ivan Srba, Andrea Tagarelli
arXiv Computation and Language
6d ago

FAVoR: Measuring and Mitigating Author-Style Homogenization in Federated Personalized Generation

The paper introduces FAVoR, a method for federated personalized generation that mitigates author‑style homogenization caused by standard aggregation in parameter‑efficient fine‑tuning. Using the BlogText benchmark and ASCE diagnostics, the authors show that common federated PEFT baselines preserve semantic utility but blur author‑specific style. FAVoR employs a shared‑private adapter design, where clients upload shared updates while keeping author‑specific residual corrections locally, leading to improved style retention with minimal utility loss.

By Lu Han, Jingyao Zhang, Katy Ilonka Gero, Nguyen H. Tran
arXiv AI
4d ago

Population Fidelity: Evaluating Population Representativeness in LLMs

The paper introduces Population Fidelity, an evaluation framework for assessing how well large language models (LLMs) represent human population attitudes. It focuses on three dimensions: group-level accuracy, between-group variation, and the structure of that variation. Using the framework, the authors replicate a prior study on machine bias and test cultural fine-tuning, finding that while fine-tuning improves overall alignment, it does not enhance representation of within-population differences.

By Neemias B. da Silva, Martin Lukk, Ali Sutani, Abhishek Moturu, Harris Yang, Daniel Silver, Matt Ratto, Thiago H. Silva