Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Machine Learning
6d ago

GyroNovo: Error-Guided Fragment Imputation with Mass-Aware Attention for \textit{De Novo} Peptide Sequencing

GyroNovo is a new framework for de novo peptide sequencing that improves fragment imputation by guiding the process with decoder errors observed during training. It introduces mass-aware attention using rotary embeddings to encode pairwise mass differences between spectral peaks, and creates easy and hard augmented views of spectra to train the decoder under varying corruption levels. Experiments on NovoBench demonstrate significant gains, with about 9 percentage points higher peptide-level precision and 7 percentage points higher amino-acid-level precision compared to the state-of-the-art baseline.

By Abdellah El Mekki, Laks V. S. Lakshmanan, Muhammad Abdul-Mageed
arXiv Machine Learning
6d ago

Reliability-aware Cross-sample Enhancement for Robust Multimodal Sentiment Analysis

Reliability-aware Cross-sample Enhancement (RCE) is a framework for multimodal sentiment analysis that tackles noise and missing modalities by first applying an adaptive variational information bottleneck to compress unreliable modality information. It then retrieves high‑confidence, semantically consistent neighbors from a large candidate pool to enrich current representations, and finally fuses cross‑modal interactions through a multilevel reliability‑aware mechanism. Experiments show RCE consistently outperforms state‑of‑the‑art methods in full, noisy, and missing‑modality scenarios.

By Menghua Jiang, Haokai Gao, Xiangui Kang, Haifeng Hu, Sijie Mai
arXiv Machine Learning
6d ago

NAC: Neural Action Codec for Vision-Language-Action Models

The paper introduces the Neural Action Codec (NAC), a convolutional encoder‑decoder architecture that treats short robot action trajectories as multi‑channel 1D signals and compresses them using a multi‑scale residual vector quantization (RVQGAN) model. NAC replaces traditional discrete action tokenizers with a compact, ordered token space via offset codebooks, allowing standard autoregressive policies to operate over short, structured sequences while a Vocos‑style decoder reconstructs the actions. Experiments on LIBERO‑10, RoboMimic, and real‑world manipulation tasks show that NAC achieves higher reconstruction fidelity and better average success rates than existing binning, FAST, and VQ‑based tokenizers at comparable or improved compression rates.

By Ahad Jawaid, Yu Xiang
arXiv Computer Vision
Sep 28

Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding

The paper introduces SGMA, a structure‑guided masked autoencoding framework designed for ultra‑high‑resolution scientific images. SGMA combines a content‑adaptive quadtree tokenizer that reduces gigapixel images to a fixed‑length sequence with a structure‑conditioned masking process that focuses reconstruction on spatially informative regions. The method, enhanced by Damped Accumulation to stabilize multi‑scale signals, achieves superior performance over standard MAE baselines on electron microscopy, whole‑slide optical microscopy, and X‑ray CT datasets, delivering significant accuracy gains and up to a 24.8× inference speedup.

By Enzhi Zhang, Du Wu, Rui Zhong, Cong Ma, Isaac Lyngaas, Amir Koushyar Ziabari, Xiao Wang, Peng Chen, Tao Luo, Toshio Endo, Fumiyoshi Shoji, Kento Sato, Kentaro Uesugi, Takayuki Nonoyama, Ryuji Kiyama, Masahiro Yoshida, Masaru Tezuka, Tetsuya Ishikawa, Satoshi Matsuoka, Masaharu Munetomo, Mohamed Wahib
arXiv AI
Sep 28

Improving Visual Sensitivity of LLMs on Multimodal Machine Translation with Metric-based Loss Weighting

The paper proposes Metric-based Loss Weighting to enhance visual grounding in multimodal machine translation. By increasing loss for tokens that benefit from image context—identified via the Point-wise Cross-mutual Information (PCXMI) metric and its Congruency-based variant—the method improves translation accuracy on the CoMMuTE dataset by over 7 percentage points. Experiments fine-tune three pretrained multimodal LLMs across three language directions, showing superior performance compared to standard fine-tuning while preserving overall translation quality.

By Pawe{\l} M\k{a}ka, Piotr Andruszkiewicz, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis
arXiv AI
Sep 28

ViSTA: A Simple Bridge Extends Visual Alignment to Clinical Time-Series Understanding in Multimodal LLMs

ViSTA is a lightweight adapter that adds irregular numerical measurements to a pretrained vision‑language model’s chart representations, enabling accurate clinical time‑series prediction without altering the base model’s parameters. On the MIMIC‑IV dataset, ViSTA outperforms other adaptations across four metrics for acute kidney injury and mortality prediction, achieving an AUC of 0.7376 with only 0.516 million trainable parameters. It also delivers strong temporal question‑answering performance, reaching 69.27% accuracy with significantly fewer trainable parameters than low‑rank adaptation methods.

By Junyi Gao, Yu Shi, Pingzhao Hu, Ewen M Harrison
arXiv AI
Sep 28

Achieving Tokenizer Flexibility in Language Models through Heuristic Adaptation and Supertoken Learning

The paper introduces TokenAdapt, a model‑agnostic tokenizer transplantation method that uses a hybrid heuristic to initialize new token embeddings, and a novel pre‑tokenization learning approach for multi‑word Supertokens to improve compression. TokenAdapt combines local subword decomposition and global semantic similarity to preserve semantics while reducing retraining needs. Empirical results show that TokenAdapt outperforms existing baselines such as Transtokenizer and ReTok, achieving lower perplexity ratios and significant compression gains.

By Shaurya Sharthak, Vinayak Pahalwan, Adithya Kamath, Adarsh Shirawalmath
arXiv Computation and Language
Sep 28

ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

ToolSearcher is a reinforcement learning framework designed to improve large‑scale tool selection for large language models. It introduces category‑constrained discrimination, event‑level search modeling, and trajectory‑aligned credit allocation to better distinguish similar tools, optimize multi‑turn search, and provide fine‑grained rewards. Experiments on large‑scale benchmarks show that ToolSearcher outperforms strong baselines in iterative search and complex tool composition scenarios.

By Zhenlong Dai, Xujie Song, Zitong Wang, Tong Niu, Jian liu, Weiqiang Wang, Xiu Tang, Sai Wu, Chang Yao, Jingyuan Chen
arXiv Computer Vision
Sep 28

STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence

STORM-Bench is a new benchmark for online video question answering that evaluates models’ ability to track state transitions and selectively abstain when visual evidence is insufficient. It contains 5,736 questions across 630 short, change‑dense episodes in five egocentric domains and two simulation subsets, with questions stratified by change intensity and answerability. The benchmark introduces STORM‑BR, a harmonic metric that reveals abstention failures and overconfidence on uncertain queries, showing that traditional accuracy masks gaps in epistemic reliability and state tracking.

By Siru Zhong, Shenghan Tan, Rihong Yan, Xiaohui Lv, Yuzheng Zhuang, Shuai Tao, Wulong Liu, Haohuan Fu, Yuxuan Liang
arXiv Computation and Language
Sep 28

Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents

The paper introduces HiCoMER, a framework that manages hierarchical collaborative memory and performs validity-aware retrieval for large language model agents. HiCoMER distinguishes between team and individual memories, updates their validity, and retrieves only those that remain valid, rather than treating all memories as a flat pool. Experiments on two new collaborative QA datasets show that HiCoMER reduces outdated retrieval, preserves current team consensus, and improves downstream question‑answering quality.

By Yufei Shi, Rujing Yao, Ang Li, Yang Wu, Zhuoren Jiang, Xiaozhong Liu
arXiv Computation and Language
Sep 28

Effects of Transcript Compression on LLM-based Medical Misinformation Detection in Japanese YouTube Videos

The study investigates how different transcript compression methods affect large language model (LLM) detection of medical misinformation in Japanese YouTube videos. Four input designs were compared: full transcripts, LLM-generated summaries, RAPTOR-based retrieval‑augmented generation (RAG), and a Screening approach that extracts candidate medical sentences. Results show that full transcripts yield the best classification accuracy, while all compressed inputs increase false negatives, with summaries causing the largest performance drop and Screening performing best among compressed methods yet still missing many relevant sentences. Linguistic analysis indicates that compression reduces affective, social, temporal, cognitive, and conversational cues, and increases the prominence of institutional and technical terms, thereby making fake videos appear more coherent and authoritative.

By Yuya Wake, Sho Tsugawa, Toshiyuki Amagasa
arXiv Computation and Language
Sep 28

Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries

The study investigates sycophancy in Chinese large language models (LLMs) by analyzing 364,941 responses from DeepSeek, Qwen, and Doubao to 12,165 yes/no factual questions derived from real-world search queries. It examines how user beliefs, reasoning, and anti-sycophancy prompts affect the distribution of correct, incorrect, and uncertain answers, finding that anti-sycophancy instructions can reduce belief-aligned errors but often increase uncertainty. The results show that preventing agreement with false beliefs does not necessarily preserve factual accuracy, underscoring the need for transition-level evaluation in Chinese-language factual QA.

By Geng Liu, Feng Li, Mengxiao Zhu, Francesco Pierri
arXiv Computation and Language
Sep 28

Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models

The paper introduces a symbiotic architecture that equips large language models with audio‑understanding abilities without fine‑tuning their weights. It uses an injector module to write audio‑conditioned vectors into the LLM’s key‑value cache, allowing the model to act as an audio language model while keeping the backbone unchanged. The approach improves scalability—since injection cost depends on the injector width—and preserves the LLM’s original text performance, outperforming conventional frozen‑LLM methods and approaching fine‑tuned ALM results on audio tasks.

By Yotaro Kubo, Qi Sun, Yujin Tang
arXiv Computation and Language
Sep 28

Why Better Cross-Lingual Alignment Fails for Better Cross-Lingual Transfer: Case of Encoders

The paper challenges the assumption that improving cross‑lingual alignment automatically enhances cross‑lingual transfer. Using XLM‑R models aligned on token, sentence, and masked‑language‑modeling objectives across four language pairs, the authors evaluate zero‑shot transfer on part‑of‑speech tagging and sentence classification. They find that embedding‑based alignment metrics poorly predict downstream performance and that alignment and task gradients are often nearly orthogonal, especially when operating at different representational levels.

By Yana Veitsman, Yihong Liu, Hinrich Sch\"utze
arXiv AI
Sep 28

Do LLMs Understand Context? A Knowledge Graph-Based Evaluation Framework

The paper introduces a knowledge‑graph‑based evaluation framework, S3KG, to assess whether large language models truly understand context in question answering tasks. S3KG combines structural and semantic signals into a single similarity score and is paired with a diagnostic analysis that pinpoints reasoning errors at the triplet level. Across nine benchmarks, the method outperforms existing baselines, achieving up to +7.6 F1 points and an AUROC of 0.973.

By Subavarshana Arumugam, Mamta Nallaretnam, Kithuni Wickramasinghe, Chamath Gunapala, Pragatheeswaran Vipulanandan, Kamal Premaratne, Uthayasanker Thayasivam
arXiv AI
Sep 28

Atelier: Learning Local Self-Supervised Features for CryoEM Volumes via Hypernetworks

Atelier is a self‑supervised framework that uses a transformer‑based hypernetwork to generate implicit neural representations (INRs) for cryo‑EM maps, enabling efficient, scale‑agnostic, coordinate‑conditioned feature extraction. Trained on 5,439 maps from the Electron Microscopy Data Bank, the pretrained INR provides continuous local feature fields that can be used as auxiliary channels for a 3D nested U‑Net, improving voxel‑level property prediction across eight tasks compared to a volume‑only baseline. The approach demonstrates that amortized INRs can serve as a geometry‑aware primitive for large‑scale cryo‑EM analysis.

By Phillip Lo, Sudarshan Babu, Dari Kimanius, Aly A. Khan
arXiv AI
Sep 28

Learning What to Skip: Counterfactual Credit Assignment for Efficient Multi-Agent LLM Workflows

The paper introduces Learning What to Skip (LW2S), a method that learns when to omit components in multi‑agent LLM workflows by treating omission as counterfactual credit assignment. LW2S builds action‑specific safety models from controlled skip interventions and uses calibration plus domain‑native guards to decide which steps to skip. Experiments on mathematical reasoning, multiple‑choice QA, and code generation show that LW2S cuts token cost while maintaining or improving overall accuracy, and further studies reveal component redundancy and limitations of agreement‑based skip selection.

By Jinfeng Xu, Zheyu Chen, Ziyue Peng, Zheng Lin, Shuo Yang, Jinze Li, Zheng Xing, Mengran Li, Victor C. M. Leung
arXiv AI
Sep 28

Toward AI-Augmented Cooperative Engineering Workflows: Requirements and Architecture the European Rover Challenge

The paper explores how Artificial Intelligence can enhance cooperative engineering workflows, focusing on the European Rover Challenge where student teams design complex rover systems under tight deadlines. A 40‑question survey of 14 teams revealed common bottlenecks such as poor documentation, unclear requirements, fragmented communication, informal task monitoring, and significant integration rework. Based on these findings, the authors propose requirements for AI‑augmented workflows and outline an assistant system architecture that integrates user interfaces, credential management, service selection, specialized AI services, and external engineering tools to support task clarification, requirement compliance, communication summarization, integration risk detection, and continuous knowledge capture.

By Ahmed R. Sadik, Frank Joublin, Mariusz Bujny, Antonello Ceravola, Joan Smith