Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv AI
Sep 17

A Lightweight CNN Integrated Compact Convolutional Transformer for Multi-Scale Feature Learning and reducing computational complexity for breast cancer mammography image detection and classification

The paper presents a lightweight CNN‑integrated Compact Convolutional Transformer (CCT) designed for multi‑scale feature learning in breast cancer mammography. With only 250,435 parameters, the model achieved 99‑100% accuracy across three datasets using 5‑fold cross‑validation, demonstrating robust generalization. Explainable AI components were added to clarify the classification process, aiming to increase clinical trust in resource‑constrained settings.

By Md Taimur Ahad (Department of Management North South University, Dhaka, Bangladesh), Ainuddin Ahmed (Department of Management North South University, Dhaka, Bangladesh)
arXiv Computation and Language
Sep 17

The Limits of BPE Tokenization in Polish: Segmentation-Flexional Forms, Grammatical Anchoring, and First-Person Stability in Inflectional Language Models

The article examines how Byte‑Pair Encoding (BPE) tokenization handles Polish, an inflectional language, and finds that BPE tends to stabilize frequent surface fragments of grammatical exponents rather than true grammatical categories. It introduces the concept of grammatical form anchoring, showing that certain Polish verb forms can signal the speaking subject without an explicit pronoun, and highlights that language models may lack a stable grammatical "I" and can shift gender or mirror user forms. The study proposes Roclawski’s segmentation‑flexional forms as a diagnostic framework and suggests that more stable Polish modeling would require sublexical stabilization, anchoring grammatical form in the inflectional system, and maintaining the grammatical "I" in dialogue.

By Elzbieta Dawidek (University of Lower Silesia DSW Ideis)
arXiv Computer Vision
Sep 17

Think Before You Move: Latent Motion Reasoning for Text-to-Motion Generation

The paper introduces Latent Motion Reasoning (LMR), a two‑stage approach that separates text‑to‑motion generation into a planning phase and an execution phase. LMR uses a Dual‑Granularity Tokenizer to create a compressed, semantically rich reasoning latent for global trajectory planning and a high‑frequency execution latent for detailed motion fidelity. Experiments on T2M‑GPT and MotionStreamer show that this architecture improves both semantic alignment and physical plausibility compared to direct translation methods.

By Yijie Qian, Juncheng Wang, Yuxiang Feng, Chao Xu, Wang Lu, Yang Liu, Baigui Sun, Yiqiang Chen, Yong Liu, Shujun Wang
arXiv AI
Sep 17

A visual large language foundational model for medical image recognition using clinician-contributed online resources

The paper introduces ThoughtMed-1M, a large-scale medical visual question answering dataset built from de‑identified images and clinician‑generated commentaries, designed to capture structured clinical reasoning and image‑text alignment. Using this dataset, the authors train FOLTMed, a foundational large language model that achieves state‑of‑the‑art performance on 42 medical VQA benchmarks, with a macro accuracy of 85.4% and improved factuality and similarity metrics over existing models.

By Lingxuan Hou, Yuhua Xie, Yue Hu, Yan Zhuang, Junqi Li, Chengzhi Xia, Binh Phu Nguyen, Abubakar Siddique, Minh Nguyen, Yao Hou, Yanju Bao, Kexin Liu, Ke Chen, Jianjun Sun, Zeqi Li, Trung Nguyen, Jiangli Lin
arXiv AI
Sep 17

One Size Does Not Fit All! Dynamic Retriever and Generator Selection for RAG

The paper introduces DRAG, a query‑adaptive framework that jointly selects retriever and generator configurations for Retrieval‑Augmented Generation (RAG) systems. Two variants are presented: DRAG_QPP, a training‑free routing method using Query Performance Prediction and perplexity signals, and DRAG_SFT, a supervised approach that fine‑tunes an LLM to predict configurations. Experiments on three LLM families and four QA benchmarks show that DRAG_QPP matches strong static baselines while cutting inference latency, and DRAG_SFT consistently outperforms both static and training‑free adaptive baselines, demonstrating a better effectiveness‑efficiency trade‑off.

By Neeraj Anand, Payel Santra, Partha Basuchowdhuri, Debasis Ganguly, Sumit Bhatia
arXiv Computation and Language
Sep 17

A Taxonomy of Programming Languages for Code Generation

The paper introduces the first reproducible taxonomy for programming languages based on resource availability, categorizing 646 languages into four tiers. It finds that a small fraction (1.9%) of high-resource languages (Tier 3) generate the majority (74.6%) of tokens in major corpora, while the majority of languages (71.7%) are scarce and contribute only 1.0% of tokens. Statistical analysis confirms the extreme and systematic imbalance across tiers.

By Nishat Raihan, Christian Newman, Marcos Zampieri
arXiv Computation and Language
Sep 17

Divide and Conquer: Mixture-of-Bottleneck Experts in Informative Ordinal Space for Video-based Multimodal Sentiment Analysis

The paper proposes a Mixture-of-Bottleneck (MoB) framework for video-based multimodal sentiment analysis that treats sentiment as an ordinal regression problem, splitting it into polarity recognition and intensity prediction. MoB assigns modality‑specific latent experts to each sub‑task, learns compact, task‑relevant representations via an information bottleneck, and fuses these experts with a multimodal bottleneck routing module and hard mining strategy. Experiments on four datasets and language models demonstrate that MoB captures fine‑grained intra‑ and inter‑modal dynamics, improving performance and enabling more trustworthy localization of nuanced sentiment signals.

By Ronghao Lin, Qiaolin He, Zefeng Lu, Yichu Liu, Li Huang, Sijie Mai, Haifeng Hu, Yap-peng Tan
arXiv AI
Sep 17

Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering

The paper compares Knowledge-Graph Based Augmentation (Graph-RAG) with Retrieval-Augmented Generation (RAG) for answering culturally specific questions. Using the LatamQA dataset, Graph-RAG, built automatically from Wikipedia via KGGen, matches RAG performance and reduces the base LLM’s error by 72% with a standard KG and 78% with a benchmark-aware variant. The approach also transfers zero‑shot to Portuguese, showing multilingual applicability.

By Pablo Poulenard, Yannis Karmim, Valentin Barri\`ere
arXiv Computation and Language
Sep 17

DANTINOX: A Unified Framework for Multi-Paradigm Language Modeling

The paper introduces DantinoX, an open‑source JAX/Flax library that unifies autoregressive decoding, discrete masked diffusion, and continuous flow‑matching language modeling under a single modular Transformer backbone. By keeping the backbone, tokenizer, initialization, and training infrastructure consistent, users can switch between generation paradigms, attention mechanisms, or hardware topologies with only a configuration change. This design enables controlled cross‑paradigm comparisons within one API for training, streaming inference, and benchmarking.

By Marco Simoni, Aleksandar Fontana, Giulio Rossolini, Andrea Saracino
arXiv AI
Sep 17

Rethinking Domain Specialization for Open-Ended Scientific Reasoning in Astronomy Language Models

The paper investigates whether domain-specific fine‑tuning benefits open‑ended scientific reasoning in astronomy. Using a curated 300‑question QA benchmark from 2017–2026 Olympiad‑style materials, the authors compare open‑weight, API‑served general‑purpose, multimodal, and astronomy‑specialized language models. Results show that strong general‑purpose models set the highest correctness baseline, but variations in metric agreement, judge sensitivity, benchmark composition, and modality suggest that domain specialization is task‑ and deployment‑dependent and that domain‑specific evaluation is crucial for scientific workflows.

By Vanessa Lama, Sanjay Das, Emily Herron, Yuan-Sen Ting, Tijmen de Haan, Junqi Yin, Tirthankar Ghosal, Feiyi Wang
arXiv Machine Learning
Sep 17

Instrument Classification of Solo Sheet Music Images

This paper investigates instrument classification using solo sheet music images rather than audio. It converts images into sequences of musical words via bootleg score representation and treats the task as text classification, training AWD‑LSTM, GPT‑2, and RoBERTa models on IMSLP data for eight instruments. Pretraining on unlabeled data and fine‑tuning improves RoBERTa’s accuracy from 34.5% to 42.9%, and two proposed data‑augmentation methods raise accuracy by an additional 15%.

By Kevin Ji, Daniel Yang, TJ Tsai
arXiv Computer Vision
Sep 17

Sim-to-Real Traffic Scene Understanding by Decoupling Semantics from Caption Generation with V-JEPA

The paper presents a decoupled framework for sim-to-real traffic scene understanding, separating semantic fact extraction from caption generation. It uses a frozen V-JEPA encoder for predictive scene representations and a lightweight Llama-based predictor for VQA, followed by a training-free structured refinement that leverages statistical priors, inter-question relationships, and temporal consistency. The refined facts are then fed to Qwen3-VL-8B to produce pedestrian and vehicle descriptions, achieving top performance on the 2026 AI City Challenge Track 2 benchmark with 87.09% VQA accuracy and an overall S2 score of 60.0853.

By Nguyen Hoai Thuong Bui, Thanh Nguyen Vo, Trinh Tra Giang Nguyen, Ha Duc Bui
arXiv AI
Sep 17

Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment

The paper reports the first systematic audit of open‑weight large language models (LLMs) in hiring contexts, examining how job‑posting language influences recruiter and job‑seeker simulations across six models. It finds that agentic language lowers recruiter scores for female candidates while communal language mitigates this effect, and that coded‑exclusion language sharply reduces recruiter scores for non‑White candidates and discourages non‑White personas from applying. The study also identifies the explicit demographic label as the main causal factor and proposes a concrete pre‑deployment audit protocol aligned with EU and U.S. regulatory requirements.

By Kosuke Kitahara, Nobuhiro Yamaguchi