arXiv:2608.22922v1 Announce Type: new
Abstract: We present HelaBERT, a family of two BERT-based masked language models pre-trained from scratch on approximately 1 billion tokens of Sinhala text sourc...
By Thisen Ekanayake, Nisansa de Silva
The article argues that Retrieval-Augmented Generation (RAG) is only one tool in NLP, and many real-world problems—such as request classification, free‑text matching, table reading, and OCR noise cleaning—are better served by simpler, cheaper techniques. It emphasizes the importance of selecting the appropriate method for each task and highlights the engineering challenge of knowing which technique to apply.
By Kezhan Shi
arXiv:2609.15369v1 Announce Type: new
Abstract: Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-l...
By Jochen Madler (Sitefire)
The article "From Words to Vectors: What Happens in Between?" explores the process of converting textual data into numerical representations, focusing on techniques such as TF-IDF and vector space models. It discusses how these representations enable text classification tasks and provides a practical overview of the underlying concepts. The piece serves as a guide for readers interested in the mechanics of text preprocessing and feature extraction for machine learning.
By Nikhil Dasari
arXiv:2607. 29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs).
By Gaetano Perrone, Simon Pietro Romano
What exactly does word2vec learn, and how? Answering this question amounts to understanding representation learning in a minimal yet interesting language modeling task.
A benchmark score means nothing without knowing what a trivial method achieves and what the best possible method could achieve. We construct both bounds for a task with a rare kind of ground truth: predicting which sentences a crowd of readers -- highlighting for their own purposes, unpaid, uninstructed, and blind to each other -- marked in 120 web documents.
arXiv:2604.08568v3 Announce Type: replace-cross
Abstract: The widespread use of LLM-based writing assistance has raised an interesting question about the homogenization of English. As LLMs tend to re...
By Nabelanita Utami, Ryohei Sasano
arXiv:2608. 07208v1 Announce Type: cross Abstract: Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities.
By Luc Hazenoot, Zhaochun Ren, Amirhossein Zohrehvand
Lit3R is a system developed by tus-nlp for the LitTraceQA shared task, which focuses on evidence-grounded question answering over scientific literature. The system combines off-the-shelf retrieval, reranking, and large language model components without task-specific training, using an iterative retrieval process that merges BM25-based sparse and dense retrieval, cross-encoder reranking, and LLM verification, along with paper-to-paper expansion. In the official test set, Lit3R achieved a 4th place ranking on the leaderboard.
By Akira Ise, Kotaro Kumagai, Yuta Yamaguchi, Hisanori Ozaki, Yukio Uematsu, Ikuya Yamada
The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.
By Valentin Romanov, Monique Bax, Steven Niederer
Enterprise Document Intelligence [Vol. 1 #2] Why the same vector search that handles synonyms and paraphrase silently fails on negation, exact identifiers, and your company’s acronyms, and what to use when it does.
By angela shi