Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Computation and Language
Sep 25

Parts-of-Speech as Emergent Categories in SAE Latent Space

The study investigates how part‑of‑speech (PoS) categories are represented in the latent space of Sparse AutoEncoders (SAEs) applied to language models. Results show that PoS distinctions can be reliably recovered from SAE activations, but the mapping is not one‑to‑one; instead, PoS categories are supported by compact, distributed groups of sparse latents that vary across tags and remain stable on held‑out data. The findings suggest that SAEs encode morpho‑syntactic information in a distributed, category‑dependent manner rather than through isolated grammatical features.

By Alessandro Bondielli, Lucia Passaro, Serena Auriemma, Alessandro Lenci
arXiv Computation and Language
Sep 25

Benchmarking Arabic--Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap

The study evaluates Arabic–Russian machine translation by comparing seven fine‑tuned neural machine translation (NMT) models with four few‑shot large language models (LLMs) on a new 15.47 million‑pair corpus split into 20k/5k/5k. Fine‑tuned NLLB‑1.3B achieves the best performance (BLEU 16.3, COMET 0.738), while the best few‑shot LLM, Aya‑Expanse 8B, scores only BLEU 1.7 on 500 sentences. Error analysis shows that low lexical overlap between Arabic and Russian is the main source of failures, and statistical tests confirm significant performance gaps between most models.

By Mullosharaf K. Arabov
arXiv Computation and Language
Sep 25

Evaluation of OpenAI o1: Opportunities and Challenges of AGI

The study evaluates OpenAI’s o1‑preview large language model on a wide range of complex reasoning tasks across domains such as computer science, mathematics, natural sciences, medicine, linguistics, and social sciences. It reports high success rates, including 83.3% on competitive programming, 100% on high‑school math reasoning, and superior performance in radiology reporting, chip design, anthropology, geology, quantitative investing, and social media analysis. While the model excels at intricate reasoning and knowledge integration, it still shows occasional errors on simpler problems and struggles with some highly specialized concepts.

By Tianyang Zhong, Zhengliang Liu, Yi Pan, Yutong Zhang, Zeyu Zhang, Yifan Zhou, Shizhe Liang, Zihao Wu, Yanjun Lyu, Peng Shu, Xiaowei Yu, Chao Cao, Hanqi Jiang, Hanxu Chen, Yiwei Li, Junhao Chen, Huawen Hu, Yiheng Liu, Huaqin Zhao, Shaochen Xu, Haixing Dai, Lin Zhao, Ruidong Zhang, Wei Zhao, Zhenyuan Yang, Jingyuan Chen, Peilong Wang, Wei Ruan, Hui Wang, Huan Zhao, Jing Zhang, Yiming Ren, Shihuan Qin, Tong Chen, Jiaxi Li, Arif Hassan Zidan, Afrar Jahin, Minheng Chen, Sichen Xia, Jason Holmes, Yan Zhuang, Jiaqi Wang, Bochen Xu, Weiran Xia, Jichao Yu, Kaibo Tang, Yaxuan Yang, Bolun Sun, Lifeng Chen, Tao Yang, Guoyu Lu, Xianqiao Wang, Lilong Chai, He Li, Jin Lu, Xin Zhang, Bao Ge, Xintao Hu, Lian Zhang, Hua Zhou, Lu Zhang, Shu Zhang, Zhen Xiang, Yudan Ren, Jun Liu, Xi Jiang, Yu Bao, Wei Zhang, Xiang Li, Gang Li, Wei Liu, Dinggang Shen, Andrea Sikora, Xiaoming Zhai, Dajiang Zhu, Tuo Zhang, Tianming Liu
arXiv AI
Sep 25

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition

Cross-Modal Emotion Understanding: A Transformer-GAT Approach for Dialogue Emotion Recognition proposes a hybrid framework that combines a Transformer and a Graph Attention Network to capture both global semantic information and fine-grained relationships between modalities. The model is evaluated on the IEMOCAP and MELD datasets, achieving weighted F1 scores of 72.45% and 77.37%, respectively, and surpasses state‑of‑the‑art methods. These results suggest that integrating multimodal features with balanced global and local context modeling can provide deeper emotional insights for dialogue emotion recognition.

By Jiaqi Qiao, Yifan Lyu, Xiujuan Xu
arXiv AI
Sep 25

Can Classical Semantic-Extractive Summarization Be Evaluated in Hindi? A Replication Study

The study replicates a distributional‑semantics extractive summarisation method for Hindi, adapting all language‑specific components to Devanagari. Evaluated on the Hindi portions of XL‑Sum and FIRE ILSUM 2.0 with a Devanagari‑aware ROUGE scorer, the replicated system performs significantly worse than a simple three‑sentence lead baseline. Feature ablation shows that sentence position alone reproduces the lead baseline, while other features only steer extraction toward long, entity‑dense body sentences, and TextRank performs identically. "whyItMatters":"The results indicate that current Hindi summarisation benchmarks cannot reward non‑lead content selection, highlighting the need for purpose‑built evaluation resources."

By Showket Ahmad Khan, Mudasir Mohd, Nasrullah Sheikh, Mohsin Altaf Wani, Abid Hussain Wani, Hilal Ahmad Khanday, Niyaz Ahmad Wani
Hugging Face Trending Papers
Sep 24

From Text Decisions to Pixels: An Study of Jev-Style Visual Choice Model

The paper introduces PixelJev, a native-image decision interface that combines an image, a task instruction, and a runtime candidate set to produce a structured choice and candidate-conditioned probabilities using small open multimodal models. It unifies recognition and multiple-choice visual question answering via a language-model readout, offering options for frozen inference, language-side adaptation, and held-out calibration. Across seven benchmarks, 64-shot source adaptation significantly boosts Pets accuracy from 60.13% to 92.40%, and the model supports VQA tasks without target fitting, though calibration and cross-family transfer remain challenges.

arXiv Machine Learning
Sep 24

Less Language, More Latents: Annotation-Efficient VLAs for Driving

The paper introduces Latent Action Driving Annotations (LADA), a three‑stage pipeline that converts large amounts of unlabelled observation‑trajectory data into a language‑conditioned driving model. First, a latent action model with a vector‑quantised bottleneck learns a compact codebook of vehicle intents. Then, a small set of language‑annotated examples trains a vision‑language translator to map observations and instructions into this codebook, and finally a VLA is trained on observation‑latent‑action pairs across the full corpus. Using less than 5% of language annotations, LADA attains a Driving Score of 87.98 and a Success Rate of 70.46% on Bench2Drive, matching or surpassing fully supervised baselines.

By Alexey Zakharov, Kemal Oksuz, Puneet K. Dokania
arXiv Machine Learning
Sep 24

From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015-2026

This scoping review examines 421 studies (2015‑2026) on natural language processing applied to student evaluation of teaching comments. It maps the technical evolution from lexicons and classifiers to transformers and large language models, and evaluates four value dimensions. The review identifies a significant gap between actionable outputs (61.3%) and intended‑user evaluation (11.6%), highlighting limited progress in educational value and robustness.

By Jeff Eicher, Rafael da Silva
arXiv Machine Learning
Sep 24

Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following

Fine‑tuning large language models on parallel data can improve translation quality but also causes catastrophic forgetting of general capabilities. The study evaluates several forgetting‑mitigation methods—anchored to auxiliary data, model outputs, and base model parameters—using Llama 3.2 1B Instruct and Llama 3.1 8B Instruct on Arabic‑English and Spanish‑English translation tasks. Elastic Weight Consolidation best preserves general benchmark performance, yet only data mixing with control‑task examples maintains instruction‑following abilities such as formality and grammatical gender control, though these gains do not generalize to unseen prompts.

By Niklas Scholz, David Thulke, Abdallah Nasir, Will Allred, Evgeny Matusov, Hermann Ney
arXiv Machine Learning
Sep 24

Learning to Approximate Uniform Facility Location via Graph Neural Networks

The paper introduces a fully differentiable message‑passing neural network (MPNN) designed to approximate the Uniform Facility Location (UniFL) problem. Unlike many learning‑based approaches that require supervision or reinforcement learning, this model incorporates principles from classical approximation algorithms, providing provable approximation guarantees. Empirical results show that it outperforms standard approximation algorithms and reduces the performance gap to integer linear programming solutions.

By Chendi Qian, Christopher Morris, Stefanie Jegelka, Christian Sohler
arXiv Computation and Language
Sep 24

Can One Adapted Model Do It All? Fine-Tuning Strategy Selection for Customer Support LLMs

The study evaluates whether a single large language model (LLM) can handle multiple customer‑support tasks or if separate specialist models are preferable. Using 13 models from five families and 200+ checkpoints across eight datasets, the authors find that multi‑task full fine‑tuning consistently outperforms other strategies. They also show that sequential LoRA and model merging can preserve earlier skills and improve off‑task robustness, offering practical guidelines for real‑world deployment.

By Md Tahmid Rahman Laskar, Xue-Yong Fu, Shashi Bhushan TN
arXiv Computation and Language
Sep 24

Cross-Lingual Legal QA for Vietnamese Labour Law: Retrieval, Translation, and Verifier-Guided Correction

The paper presents a cross‑lingual legal QA system for Vietnamese labour law, introducing a bilingual evaluation suite of 231 Vietnamese–English question–answer pairs, 75 of which are annotated for five complex legal reasoning phenomena. It evaluates a verifier‑guided pipeline that decomposes answers into claims, checks citation reachability and entailment, and corrects citation failures and contradictions, and introduces six automatic diagnostics for faithfulness to retrieved evidence. Experiments show that dense retrieval outperforms sparse and hybrid retrieval, translation placement has no significant effect on diagnostics, and verifier‑guided correction modestly improves citation preservation but not other dimensions, with human evaluation indicating a gap between automatic diagnostics and human judgments.

By Nguyen Minh Chi, Mo El-Haj, Nguyen Ha Thanh, Dawn Knight, Paul Rayson
arXiv Computation and Language
Sep 24

MetaHOPE: A Metaphor-Oriented Evaluation Framework for Analysing MT and LLM Translation Errors

MetaHOPE is an error‑severity‑aware annotation framework designed to evaluate how well machine translation (MT) and large language models (LLMs) translate metaphors. The authors applied MetaHOPE to three state‑of‑the‑art systems—GoogleMT, GPT5.4, and Hunyuan‑7b—using two human‑annotated metaphor corpora (VUAMC and PSUCMC) for English‑to‑Chinese and Chinese‑to‑English translation. They also produced a bilingual post‑edited gold reference, creating a new resource for metaphor translation research.

By Jiahui Liang, Lifeng Han
arXiv Computer Vision
Sep 24

VIVAS: Vitalizing Visual Perception in VLM Pre-training via Vision-language Unified Autoregressive Supervision

VIVAS is a new Vision‑Language Model pre‑training framework that addresses the lack of fine‑grained visual perception in existing VLMs. It introduces a unified token space and a dense‑structural‑semantic vision tokenizer that expands the textual vocabulary with visual tokens, enabling vision‑language unified autoregressive supervision over both visual details and linguistic content. Trained on 12.4 T tokens, VIVAS achieves state‑of‑the‑art results on 7 tasks and 39 multimodal benchmarks.

By Zhehan Kan, Yubo Zhu, Xinghua Jiang, Zhixiang Wei, Shifeng Liu, Wei Tong, Sheng Zhong, Qingmin Liao, Wenming Yang, Xin Li, Yinsong Liu, Deqiang Jiang, Xing Sun
arXiv Computer Vision
Sep 24

ASAP: Visual Analytics for Identifying and Analyzing Image Patterns in AI-generated Images

ASAP is an interactive visualization system that helps users identify and analyze deceptive patterns in AI‑generated images. It uses a CLIP‑adapted image encoder to produce interpretable representations and generates masks that highlight influential pixel regions, enabling influence measurement of key deceptive features. The system integrates these techniques into a dashboard for quantifying authenticity‑indicative patterns across collections of authentic and AI‑generated images, supporting comparative analysis of different generative models such as GANs and diffusion models, and its effectiveness is demonstrated through a user study and benchmark applications.

By Jinbin Huang, Yuki Ueno, Chen Chen, Aditi Mishra, Bum Chul Kwon, Zhicheng Liu, Chris Bryan
arXiv Computer Vision
Sep 24

Copy-Move Forgery Detection and Question Answering for Remote Sensing Image

The paper introduces the Remote Sensing Copy-Move Question Answering (RSCMQA) task, aimed at interpreting complex tampering scenarios in land resource monitoring and national defense. It presents five datasets—RS-CMQA, RS-CMQA-B, Real-RSCM, RS-TQA, and RS-TQA-B—spanning 29 regions in 14 countries, and a region-discrimination-guided multimodal copy-move forgery perception framework (CMFPF) that improves question answering accuracy on tampered images. Experiments show CMFPF outperforms general VQA and RSVQA models, establishing a new benchmark for RSCMQA.

By Ze Zhang, Enyuan Zhao, Di Niu, Jie Nie, Xinyue Liang, Lei Huang
arXiv AI
Sep 24

Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions

The study evaluates how well pre‑trained models can classify the Bloom level of AI‑generated educational questions, a task that is crucial for ensuring pedagogical quality. Traditional machine‑learning models perform poorly on out‑of‑distribution data, whereas transformer and large‑language models achieve higher accuracy, especially after feature‑engineering techniques such as text splicing and appending learning objectives. Retraining the models yields the most significant performance gains across all datasets.

By Michael Lawrence Castanares, Princess Ventures, Allan Tan
arXiv AI
Sep 24

Hard Negatives Reveal What Easy Negatives Hide: Cross-Lingual Harmfulness Representations Degrade with Resource Tier Under Hard Negatives

The study investigates how safety alignment in large language models, trained mainly in English, transfers to other languages. While models show near-perfect harmfulness detection (AUROC > 0.98) using unrelated harmless prompts (easy negatives), performance drops sharply in low‑resource languages when using surface‑similar benign prompts (hard negatives). This degradation persists across multiple languages and models, indicating that easy‑negative evaluation alone cannot confirm cross‑lingual harmfulness representation quality.

By Paras Balani, Subhrakanta Panda
arXiv Machine Learning
Sep 24

InsurTech innovation using natural language processing

The paper "InsurTech innovation using natural language processing" outlines how insurance firms are adopting NLP to convert unstructured text into structured data for actuarial analysis. It presents case studies that use alternative data from an InsurTech partner to demonstrate feature de‑biasing, compression, and new industry classification methods in commercial insurance. These techniques enhance traditional rating factors and provide fresh risk assessment perspectives, positioning NLP as a core component of contemporary insurance analytics.

By Panyi Dong, Zhiyu Quan
arXiv Computer Vision
Sep 24

Gender Bias in Vision-Language In-Context Learning

The paper investigates how in‑context learning (ICL) in large vision‑language models (LVLMs) can amplify gender bias. Using the VL‑BICLE framework, the authors show that gendered ICL demonstrations shift model bias toward the demonstrated gender, especially in tasks involving gendered language such as image captioning and pronoun prediction. They find that similarity‑based retrieval does not mitigate this bias and that replacing real images with synthetic ones from stable diffusion reduces bias without hurting caption quality.

By Tong Xiang, Noa Garcia, Yuta Nakashima