The paper introduces a modular data‑science pipeline that estimates public sentiment toward individuals using fragmented, unstructured open‑source intelligence. The pipeline combines web search, text extraction, relevance filtering, tokenisation, co‑reference resolution, and sentiment analysis to produce auditable person‑level sentiment distributions. By comparing AFINN, VADER, and the domain‑specific MINOS algorithm, the authors show that MINOS best distinguishes positive, ambiguous, and negative reputational cases, and they apply the method to the UK Honours system to support transparent, reproducible, human‑in‑the‑loop sentiment assessment for high‑stakes decisions.
By Francesca von Braun-Bates, Sunreeta Sen, Indraayudh Talukdar, Anirban Lahiri
arXiv:2609.22133v1 Announce Type: new
Abstract: In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at...
By Kentaro Nakamura, Jing Ling Tan, George Yean
The paper introduces an open, modular AI framework that automatically detects and structures evidence of social tipping points in climate literature at the passage level. It integrates a DistilBERT boundary splitter, an iteratively augmented RoBERTa classifier, a Mistral 7B rewrite model, a LLaMA 3.2 3B rating model, and a Milvus vector store, all accessible via a Streamlit interface. Evaluation on a GPT‑4.1‑labelled benchmark and expert‑reviewed set shows the splitter outperforms competitors and the RoBERTa detector achieves high accuracy and agreement.
By Kavindu Perera, Mohammad Abaeiani, Ekaterina Gilman, Lauri Loven, Mourad Oussalah, Tassos Kanellos, Beatrice Gobbo, Dante Adami, Nicol\`o Ferriani, Maximiliano Romero, Pierre Rossel, Marc Bonazountas, Christina Deligianni, Nikos Xyderis, Artur Bogucki, Lampros Argyriou, Prasasthy Balasubramanian
The paper evaluates browser-based large language models (LLMs) for extracting detailed, contextualized data from scientific papers. It presents four workflows: (1) expert-curated prompts yield good extraction but struggle with nuance; (2) LLMs can generate effective prompts from simple instructions; (3) autonomous literature discovery is challenging, with missing or hallucinated references; (4) LLMs can build new datasets from guidelines that align closely with human experts, yet still need human oversight. The study outlines a practical, auditable workflow where experts set standards, models cross-check extractions, and researchers resolve disputes, enabling scalable scientific data curation.
arXiv:2604. 17289v2 Announce Type: replace Abstract: Supervised fine-tuning of large language models relies on human-annotated data, yet annotation pipelines routinely involve multiple crowdworkers of heterogeneous expertise.
By Sajjad Ghiasvand, Mark Beliaev, Mahnoosh Alizadeh, Ramtin Pedarsani
arXiv:2607. 06482v1 Announce Type: cross Abstract: Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings.
By So Hasegawa, Shailaja Keyur Sampat, Lei Liu, Wei-Peng Chen