arXiv Machine Learning

Computational Approaches to Understanding Large Language Model Impact on Writing and Information Ecosystems

arXiv:2506. 17467v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown significant potential to change how we write, communicate, and create, leading to rapid adoption across society.

arXiv AI
Aug 28

How LLMs Distort Our Written Language

Large language models (LLMs) are widely used to assist writing, but this study shows they alter both tone and meaning of human text. A user study found that heavy LLM use increased neutral essays by nearly 70% and made writers feel less creative and less in their voice. Even when prompted to make only grammar edits, LLMs changed the semantic content of essays and produced AI-generated scientific reviews that were less focused on clarity and significance and scored higher on average.

By Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Z. Leibo, Max Kleiman-Weiner, Natasha Jaques
arXiv AI
Sep 18

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

The paper "What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks" analyzes 14,767 arXiv submissions from 2022 to 2026 that introduce or update evaluation resources for large language models. It systematically maps changes in target systems, domains, evaluation materials, conditions, and scoring mechanisms, revealing a growing emphasis on action, interaction, and professional applications. The study also notes uneven development in model participation, with LLM-based scoring increasing in both agent and non-agent groups, while model-generated materials do not show a comparable rise.

By Chao Wang (Independent Researcher)
Hugging Face Trending Papers
Jun 30

FARS: A Fully Automated Research System Deployed at Scale

Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Automated Research System), a fully automated AI-for-AI research system designed to operate across research topics at scale.

arXiv AI
Jul 1

FARS: A Fully Automated Research System Deployed at Scale

arXiv:2606. 31651v1 Announce Type: new Abstract: Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks.

By Qiong Tang, Xiangkun Hu, Xiangyang Liu, Yiran Chen, Yunfan Shao
arXiv Computation and Language
Sep 4

Beyond Accuracy: Community Perspectives on Machine Translation

The paper examines how four stakeholder groups—AI developers, professional translators, language learners, and language service providers—discuss machine translation on social media. Using a dataset of 79,286 posts from Reddit, Facebook, Bluesky, and Mastodon (2019‑2025), the authors find frequent disagreements and strong conflicts over translation quality, efficiency, and reliability. These conflicts arise because AI communities view the issues as technical, while non‑AI users prioritize quality nuances, time savings, trust, and broader social concerns.

By Yujun Wang, Ehud Reiter, Shimei Pan, Steffen Eger, Wei Zhao
arXiv AI
Aug 28

LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

The study evaluates literature reviews produced by large language models (LLMs) using short and long context windows, assessing their quality across 15 dimensions. Results show that while larger context windows allow LLMs to incorporate more information and maintain coherence, they also increase repetition, omission of key works, and a tendency toward descriptive rather than synthetic content. Human oversight remains essential for meeting academic publishing standards, and the authors suggest future work should blend human expertise with AI to mitigate these limitations.

By Muhammad Ali Chaudhry, Xinyuan Hao, Haifa Alwahaby
arXiv AI
Jun 24

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

arXiv:2603. 20450v2 Announce Type: replace-cross Abstract: A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews.

By Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar B Shah, Danish Pruthi
Hugging Face Trending Papers
Jun 8

Beyond Accuracy: Community Perspectives on Machine Translation

Despite remarkable progress in machine translation (MT), non-AI communities have raised growing concerns about MT systems, suggesting a noticeable gap between technical advancement and the needs of real-world users. For instance, while NLP researchers focus on benchmark performance, end users care about ethical concerns, trust, reliability, costs, and more.