Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv Computation and Language
Sep 22

SG-FSM: A Self-Guiding Zero-Shot Prompting Paradigm for Multi-Hop Question Answering Based on Finite State Machine

arXiv:2410.17021v2 Announce Type: replace Abstract: Large Language Models with chain-of-thought prompting, such as OpenAI-o1, have shown impressive capabilities in natural language inference tasks. H...

By Xiaochen Wang, Liang Chen, Reza Haf Zhe Yang, Yiru Wang, Xiangdi Meng, Kunhao Pan, Zhifang Sui, Junqing He
arXiv Computation and Language
Sep 22

Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

arXiv:2607.01733v2 Announce Type: replace Abstract: Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech reco...

By Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren, Keqi Deng, Xiaoyang Chen, Ali Zare, Bo Ren, Yuxuan Hu, Junkun Chen, Yan Huang, Yelong Shen, Jinyu Li
arXiv Machine Learning
Sep 22

A discrete generative model of neuronal spiking activity on microelectrode arrays

The paper presents a discrete generative model for neuronal spiking activity recorded on microelectrode arrays. It uses a shared vocabulary of spatiotemporal motifs learned by a residual vector‑quantized autoencoder and predicts motif occurrence with a factorized masked transformer. Evaluated on 31 assays from human brain organoids and ex vivo hippocampal tissue, the model achieves superior reconstruction and generation performance compared to baselines and shows that motifs are largely reused across assays.

By Md Sayed Tanveer, Mohammed A. Mostajo-Radji, Ge Wang
arXiv Computation and Language
Sep 22

Representation-guided in-context learning for medical image interpretation with multimodal large language models

arXiv:2609.24057v1 Announce Type: cross Abstract: Medical image interpretation is central to diagnosis and care, yet adapting general-purpose multimodal large language models (MLLMs) often requires r...

By Minda Zhao, Fangyu Hu, Yan Luo, Yutong Yang, Jiahui Cai, Kaichen Zhou, Manling Li, Paul Liang, Yilun Du, Lucy Q. Shen, Mengyu Wang
arXiv Machine Learning
Sep 22

MolSC: Leveraging Substituent Contributions to Enhance Fine-grained Molecular Understanding in LLMs

MolSC is a new dataset of 181,000 substituent-level examples that captures how attaching specific substituents to molecular scaffolds changes properties such as bioactivity and physicochemical descriptors. The authors also provide MolSC-Bench, a held‑out benchmark of 1,541 examples that are disjoint from MolSC at scaffold, substituent, and molecule levels. Experiments show that training molecular large language models on MolSC markedly improves their ability to predict substituent contributions, outperforming existing models on a range of downstream chemistry tasks.

By Hyuntae Park, Sooyeon Kim, Jiwon Park, SangKeun Lee
arXiv Computer Vision
Sep 22

What Makes a Good Medical Image Tokenizer? Rethinking Reconstruction and Generation in Medical Image Tokenization

arXiv:2609.24691v1 Announce Type: new Abstract: Latent diffusion models now dominate medical image generation, and every such pipeline rests on a \emph{tokenizer} that compresses images into the late...

By Niklas Bubeck, Yundi Zhang, Vasiliki Sideri-Lampretsa, Julian McGinnis, Jiancheng Yang, Daniel Rueckert, Jiazhen Pan