Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Computation and Language
Sep 22

Analyzing Public Discourse on Urbanism: Topic Clustering, Sentiment Analysis and Retrieval-Augmented Generation using YouTube Comments

The paper introduces a pipeline and conversational system that processes 22,788 YouTube transcript and comment chunks from 309 North American cities to analyze public discourse on urbanism. It combines geographic entity resolution, topic modeling, sentiment analysis, and Retrieval-Augmented Generation (RAG), and reports empirical findings on model performance, such as a Twitter-tuned RoBERTa classifier outperforming VADER and dense retrieval surpassing TF‑IDF. The study also evaluates groundedness metrics, noting limitations of BERTScore and ROUGE‑1 for short user-generated text.

By Jakob Morales, Monica Hegde, Fayeq Jeelani Syed
arXiv Computation and Language
Sep 22

Custom Named Entity Recognition and Topic Classification for Global Health Publications

This thesis explores how to select and adapt NLP models for global health literature when annotated data and computational resources are scarce. It compares skip‑gram word2vec models trained on increasingly large specialized corpora with BioWordVec for semantic tag discovery, finding that larger coverage does not always yield more useful domain associations. The study also evaluates convolutional spaCy models versus a RoBERTa transformer for named entity recognition, noting a trade‑off between higher F1 scores and longer inference time, and investigates MiniLM few‑shot versus BART‑MNLI zero‑shot classification for multi‑label topic classification, highlighting practical constraints of inference cost. "whyItMatters":"The work provides empirical guidance on balancing model accuracy and resource demands for building knowledge systems in low‑resource global health settings."

By Genis Skura, Antoine Geissb\"uhler, Jean-Luc Falcone
arXiv Computation and Language
Sep 22

MechaTerp-TRACE: A Novel Approach for Component Ablation Analysis in Language Models

MechaTerp-TRACE is a new framework that systematically ablates individual components of language models to measure their causal contribution to producing a named entity. By applying TRACE to thirteen instruction‑tuned dense decoder models, the study finds that a small set of positionally fixed components consistently carry the most influence across models and prompts, while the remaining support is evenly distributed. This suggests that entity knowledge is largely embedded in generic generation machinery rather than in isolated, findable components.

By Brandon Colelough, Davis Bartels, Madeline Bittner, Dina Demner-Fushman
arXiv Computation and Language
Sep 22

Cross-Dialect NER for Bangla Regional Dialects Using Leave-One-Dialect-Out Cross-Validation and Explainable AI

The paper introduces a cross-dialect Named Entity Recognition (NER) framework for Bangla, leveraging the ANCHOLIK-NER dataset that covers five major regional dialects. Using a Leave-One-Dialect-Out Cross-Validation strategy, eight transformer-based models were evaluated, with Multilingual-E5 Large achieving the best performance (F1 up to 97.26% on Mymensingh, 82.38% on Chattogram). Local Interpretable Model-agnostic Explanations (LIME) revealed that the models rely mainly on the surface form of entity words rather than surrounding context, suggesting a direction for future improvement.

By Shamim Rahim Refat, Faika Fairuj Preotee, Shuvashis Sarker, Shifat Islam, Bidyarthi Paul, Mohammad Ashraful Hoque
arXiv Computation and Language
Sep 22

To Consolidate or not to Consolidate? Evaluating the Impact of Consolidation in Multi-Reference Training using Peer Reviews

arXiv:2609.22805v1 Announce Type: new Abstract: Natural language generation (NLG) tasks span the spectrum of conditional entropy, ranging from highly constrained machine translation to open-ended dia...

By Maitreya Prafulla Chitale, Ketaki Mangesh Shetye, Yash More, Harshit Gupta, Manav Chaudhary, Manish Shrivastava, Vasudeva Varma