arXiv Computation and Language By Yun-Ang Wu, Xanh Ho, Andre Greiner-Petter, Sunisth Kumar, Tian Cheng Xia, Florian Boudin, Akiko Aizawa

How Robust Is Multimodal Claim Verification to LLM Rewriting?

Read the original on arXiv Computation and Language →

The paper investigates how stylistic changes introduced by large language models (LLMs) affect multimodal claim verification, a task that determines whether a textual claim is supported by given evidence. Two rewriting strategies are used: natural rewriting, mimicking typical academic polishing, and controlled injection, adding a single LLM-associated word. Across 11 open‑weight models (2B–38B parameters) from five VLM families, the study finds that most models remain robust to these modifications, showing no significant accuracy drop, though consistent probability shifts—especially under hedging conditions—are observed.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 2

Evaluating Style-Personalized Text Generation: Challenges and Directions

The paper "Evaluating Style-Personalized Text Generation: Challenges and Directions" examines the difficulties of assessing text that is tailored to individual users’ styles. It critiques common metrics such as BLEU, embeddings, and LLM-as-judges, and introduces a style discrimination benchmark covering domain discrimination, authorship attribution, and LLM-generated personalized versus non-personalized discrimination across eight writing tasks. The study finds that ensembles of diverse evaluation metrics outperform single-evaluator approaches and offers guidance for reliable assessment of style-personalized generation.

By Anubhav Jangra, Bahareh Sarrafzadeh, Silviu Cucerzan, Adrian de Wynter, Sujay Kumar Jauhar
arXiv AI
Aug 11

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

arXiv:2608. 08975v1 Announce Type: cross Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions.

By Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou