arXiv AI

Mapping Scientific Literature with Large Language Models and Topic Modeling

arXiv:2510. 16152v2 Announce Type: replace-cross Abstract: Scientific literature is increasingly fragmented by disciplinary boundaries, specialized terminology, and potentially sparse keyword systems, making it difficult to capture the evolving structure of modern science.

arXiv AI
Aug 10

SCALE: Scientific Concept Aggregation via LLMs and Embeddings for Fine-Grained Taxonomy Extension

arXiv:2608. 07254v1 Announce Type: cross Abstract: The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of broad disciplines and research topics but often fail to capture the fine-grained conceptual structure of contemporary science.

By Daniele Raimondi, Feichi Lu, Oliver Grun, Mariia Eremina, Andrea Perlato
arXiv AI
Sep 10

Mapping the Emerging Social Science of Large Language Models

The paper maps the nascent social‑science literature on large language models (LLMs) by analysing 198 curated papers and 47,719 field‑scale papers. It identifies three main domains—LLM as Social Minds, LLM Societies, and LLM‑Human Interactions—each containing 13 subcategories such as reasoning, bias, collective intelligence, and trust. The taxonomy is validated through clustering stability, author classification agreement, and topic mapping, revealing differing prominence across conference and journal venues.

By Yi Yang, Xiao Jia, Zeyun Dong, Chenzhang Wang, Zhanzhan Zhao
arXiv AI
Jun 18

Improving Scientific Document Retrieval with Academic Concept Index

arXiv:2601. 00567v2 Announce Type: replace-cross Abstract: Adapting general-domain retrievers to scientific domains is challenging due to the scarcity of large-scale domain-specific relevance annotations and the substantial mismatch in vocabulary and information needs.

By Jeyun Lee, Junhyoung Lee, Wonbin Kweon, Bowen Jin, Yu Zhang, Susik Yoon, Dongha Lee, Hwanjo Yu, Jiawei Han, Seongku Kang
arXiv Computation and Language
Sep 23

Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining

The paper introduces WaterBERT, a domain‑adapted encoder model trained on a 2.97‑billion‑token water treatment corpus to capture domain‑specific semantics for literature mining. Fine‑tuned versions of WaterBERT outperform general‑purpose and other domain BERT models on tasks such as treatment process classification, named entity recognition, and relation extraction. The authors also demonstrate WaterBERT’s utility in large‑scale processing, generating coherent research topics, building a structured knowledge graph from 693,211 abstracts, and creating a Water Knowledge‑Enhanced Retrieval System that surpasses text‑based baselines.

By Mudi Zhai (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia), Ruihong Qiu (School of Electrical Engineering and Computer Science, The University of Queensland, Brisbane, QLD 4072, Australia), Qingyun Zeng (Microsoft Copilot Studio AI, Redmond, WA 98052, United States, Departments of Mathematics & Department of Computer and Information Science, University of Pennsylvania, Philadelphia, PA 19104, United States), T. David Waite (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia), Bing-Jie Ni (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia), Haoran Duan (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia, Department of Civil Engineering, The University of Hong Kong, Pokfulam, Hong Kong SAR, China)