arXiv:2510.08831v2 Announce Type: replace
Abstract: As AI writing tools become widespread, we need to understand how both humans and machines evaluate literary style, a domain where objective standar...
By Wouter Haverals, Meredith Martin
arXiv:2608.28986v1 Announce Type: new
Abstract: LLMs often struggle with modern Korean poetry, producing outputs that resemble "line-broken prose." We address two coupled tasks: detecting whether a K...
By Keunhyeung Park, Seunguk Yu, YoungBin Kim
arXiv:2608. 11452v1 Announce Type: cross Abstract: Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image renders the meaning of a poem.
By Haoqi Hu, Tongji Luo, Li Zhang, Boning Zhou
arXiv:2609. 28245v1 Announce Type: cross Abstract: Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored.
By AbdulRahman A. Morsy (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States), Aya Zirikly (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States, Center for Speech and Language Processing, Whiting School of Engineering, Johns Hopkins University, Baltimore MD, United States)
The paper examines whether existing automatic methods can reliably assess creativity in text produced by large language models (LLMs). By collecting human ratings on 11 creativity dimensions for both human and AI short stories, the authors compare these judgments with automated metrics and LLM-as-a-Judge evaluations. The results show a significant misalignment: automated metrics and LLM judges favor AI-generated stories and show near-zero correlation with human assessments, revealing fundamental limitations in current computational approaches to evaluating creative text.
By Alessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
arXiv:2609.23951v1 Announce Type: new
Abstract: Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior w...
By Devangi Sharma, Sophia Judicke, Glenda Tan, Conrad Schaumburg, Shinji Watanabe
However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies controlled stylistic perturbations to fixed scientific content across three settings no context, fixed-domain context, and open-domain retrieval context, covering 600 scientific ideas and 15 style variants, with 9,000 evaluation instances per setting; (ii) Second, SciStyleMetrics, a set of quantitative measures, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), to characterize how stylistic variation affects scoring stability, substance discrimination, and ranking robustness; (iii) Third, SciStyleExtractor, a plug-and-play evaluation module that separates presentation style from scientific content by predicting style type and deviation before style-conditioned evaluation, enabling us to assess whether style awareness reduces stylistic bias.
AesCanvas is a new dataset and benchmark that evaluates image aesthetic models on two fronts: CritiqueCanvas, which contains 519,136 instruction–response pairs for long‑form, multi‑dimensional critique across photography, painting, and virtual imagery, and ContextCanvas, which offers 301 expert‑reviewed use scenarios to assess contextual aesthetic suitability. The benchmark tests closed‑source, open‑weight general, and aesthetic‑specific multimodal large language models, revealing that models excel at critique generation but lag in context‑sensitive judgment. The study shows that aesthetic specialization does not reliably transfer to contextual suitability and highlights the need for culturally situated, evidence‑grounded suitability as a distinct objective for aesthetic modeling.
By Xuanwei Hu, Haoyu Dong, Kejun Wu, Tianyi Liu, Jianjun Gao
arXiv:2310.00436v2 Announce Type: replace
Abstract: Authorship identification uses patterns in writing to infer who wrote a text, but those patterns also reflect topic, genre, and register. This surv...
By Haining Wang
The paper evaluates five large language models as zero‑shot annotators of four social constructs—self‑esteem, self‑control, seeking belonging, and seeking recognition—in English song lyrics. It examines repeated‑measurement reliability, cross‑model convergence, and the transferability of consensus labels to supervised classification. Results show varying reliability across constructs, with self‑esteem being most stable and seeking recognition least stable, and indicate that consensus labels contain learnable signal for downstream tasks.
By E. Cho Smith, Samuel Ho, Dawn Laux
arXiv:2510. 20091v3 Announce Type: replace-cross Abstract: Creativity is often seen as a hallmark of human intelligence.
By Zhaoyi Joey Hou, Bowei Alvin Zhang, Yining Lu, Bhiman Kumar Baghel, Anneliese Brei, Ximing Lu, Meng Jiang, Faeze Brahman, Snigdha Chaturvedi, Haw-Shiuan Chang, Daniel Khashabi, Xiang Lorraine Li
arXiv:2608. 01666v2 Announce Type: replace-cross Abstract: However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question.
By Fengxian Ji, Yuke Li, Jingpu Yang, Juanfan Wu, Fan Zhang, Zhexuan Cui, Yu Xie, Min Peng, Qianqian Xie, Xiuying Chen, Zhuohan Xie