arXiv:2608.30754v1 Announce Type: new
Abstract: Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine h...
By Mohammad Reza Modarres, Armin Tourajmehr, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar
The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.
By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine how reliably LLMs evaluate short literary text in...
The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.
By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
arXiv:2609.16006v1 Announce Type: cross
Abstract: Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test wha...
By Enes Altinisik, Hamdy Mubarak, Masoomali Fatehkia, Husrev_Taha_Sencar Husrev Taha Sencar
A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen «This is not a course. It is a journey of transformation».
The paper introduces a training‑free method for uncovering prompt‑conditional stylistic axes in large language models (LLMs). By repeatedly sampling completions of a single prompt at high temperature and applying Principal Component Analysis (PCA) to the pooled hidden activations, the authors automatically label the resulting axes using the extreme (pole) generations. Validation against 245 human‑elicited stylistic annotations shows that, for the Qwen‑3.5‑4B‑Instruct model, the top two axes align with human dimensions with 72.8% precision and 43.6% macro‑recall, and 75.6% of validity ratings confirm the axes’ polar generations, while other models exhibit varying degrees of discoverability.
By Ajit Mallavarapu, Ziwei Gu
arXiv:2607. 15847v1 Announce Type: cross Abstract: Metaphor in Arabic is a culturally grounded mechanism for constructing meaning, encoding cultural knowledge that shapes interpretation.
By Suzan Awinat, Alfonso Ortega del Puente
arXiv:2604. 26269v2 Announce Type: replace-cross Abstract: In the era of large language models, creative writing quality lacks a computable theoretical anchor.
By Bo Zou, Chao Xu
The paper "Evaluating Style-Personalized Text Generation: Challenges and Directions" examines the difficulties of assessing text that is tailored to individual users’ styles. It critiques common metrics such as BLEU, embeddings, and LLM-as-judges, and introduces a style discrimination benchmark covering domain discrimination, authorship attribution, and LLM-generated personalized versus non-personalized discrimination across eight writing tasks. The study finds that ensembles of diverse evaluation metrics outperform single-evaluator approaches and offers guidance for reliable assessment of style-personalized generation.
By Anubhav Jangra, Bahareh Sarrafzadeh, Silviu Cucerzan, Adrian de Wynter, Sujay Kumar Jauhar
arXiv:2607. 21498v1 Announce Type: cross Abstract: A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen {\guillemotleft}This is not a course.
By Federico Boggia
arXiv:2606. 17350v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet the question of whether these models are capable of generating diverse outputs remains contested.
By Thennal DK, Hans Ole Hatzel