Natural language processing

Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.

2,564 stories · RSS feed

arXiv AI
3d ago

SkillFM: Generating Skills for LLM Agents via Latent Flow Matching

SkillFM is a generative framework that creates task‑conditioned textual skills for large language model agents without relying on manual skill banks or reinforcement learning. It encodes skills into a continuous latent space using a codec and trains a conditional flow model with improved MeanFlow, allowing single‑step latent sampling at inference. The sampled latent is decoded by an LLM into textual guidance, and the method outperforms other vector‑based skill approaches on ALFWorld, Search‑QA, and other tasks.

By Zuming Zhang, Jie He, Yizhe Zhang, Jeff Z. Pan
arXiv AI
3d ago

Personalized State-Transition-Aware Memory for Clinical Agents

The paper introduces STAM, a state‑transition‑aware memory framework for large language model agents that process clinical records. STAM records changes in a patient’s state as new entries arrive, using semantic retrieval and typed clinical relations to separate current information (Active) from superseded or resolved information (History). During retrieval, a query‑dependent gate selects the appropriate historical memory, enabling accurate question answering and state‑maintenance diagnostics across four longitudinal clinical benchmarks.

By Maryam Haghifam, Zahra Rajabi, Yizhou Sun, Carlos Morato
arXiv AI
3d ago

Prompt2Skill: Unsupervised Skill Optimization From Natural Language Instructions

Prompt2Skill is an unsupervised framework that constructs skills for Large Language Models directly from natural‑language task descriptions. It automatically derives task specifications, discovers or synthesizes datasets, and refines the skill through a reflective editing loop. In experiments across question answering, reading comprehension, spreadsheet manipulation, and mathematical reasoning, Prompt2Skill outperforms direct prompting, improving performance by an average of 10.8 points on both open‑source and frontier models.

By Bo Ni, Li Li, Ryan A. Rossi, Franck Dernoncourt, Tyler Derr
arXiv AI
3d ago

HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning

HiWE is a hierarchical world knowledge model that enables zero‑shot 3D path planning by linking visual grounding with language‑based planning through a point‑based interface. It uses PointVLM to map task‑relevant objects to image coordinates, lifts these predictions into a semantic 3D representation with depth data, and then a language planner (3DLLM) generates end‑effector waypoints and gripper commands. The system is evaluated on 14 simulated manipulation tasks and four physical‑robot tasks, with ablations on visual training data, spatial inputs, and grasp selection.

By Guoqing Ma, Mingqi Yuan, Chen Gao, Jiayu Chen, Shan Yu
arXiv AI
3d ago

TTLab at Daleel 2026: STAR-Ar, Sequence Tagging for Argument Recognition in Arabic

The paper introduces STAR‑Ar, a BERT‑BiLSTM‑CRF model designed for the Daleel 2026 Arabic argument mining shared task. It treats argument discourse unit detection and classification as a token‑level sequence labeling problem, achieving an F1‑score of 72.69 on validation and 73.7 on test data. Analysis shows that models trained only on editorial texts perform worse than those trained on debates, mainly due to the smaller editorial dataset.

By Bhuvanesh Verma, Ali Abusaleh, Alexander Mehler
arXiv AI
3d ago

LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

LEAP is a framework for long audio‑video question answering that avoids encoding entire recordings by dividing them into fixed‑duration blocks. It performs a lightweight localization pass on each block to score short candidate windows, then pools the highest‑ranked windows for a single bounded answer pass, keeping the answer input and peak context independent of recording length. The method trains both a localization LoRA and an answer LoRA, supports causal streaming queries, and achieves significant performance gains over baseline models on multiple AVQA benchmarks.

By Juyi Lin, Zhiqiang Lao, Jiali Cui, Lin Zhao, Pu Zhao, Dichang Zhang, Arman Akbari, Yu Qi, Xinru Jiang, Yanzhi Wang, Heather Yu, Liang Peng
arXiv Computation and Language
3d ago

Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

arXiv:2410.02343v2 Announce Type: replace Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer int...

By Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov
arXiv Computer Vision
3d ago

Vision-Language-Action Autonomous Driving Agent with Language-based Memory

arXiv:2609.38641v1 Announce Type: new Abstract: Vision-Language-Action (VLA) foundation models have recently emerged as one of the prevailing solutions for autonomous driving, as they can utilize kno...

By Kai Yan, Xiangyu Chen, Yulong Cao, Alex Naumann, Peter Karkus, Yan Wang, Jef Packer, Alex Schwing, Yuxiong Wang, Boris Ivanovic, Wenjie Luo, Marco Pavone
arXiv AI
3d ago

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

arXiv:2609.40219v1 Announce Type: cross Abstract: World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experienc...

By Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan
arXiv Computer Vision
3d ago

Learning Where to Look: Anatomical Grounding and Guided Attention for Cardiac MRI Vision-Language Models

arXiv:2609.39899v1 Announce Type: new Abstract: Cardiac magnetic resonance imaging (CMR) enables assessment of cardiac anatomy, ventricular function, and myocardial tissue characteristics. Clinicians...

By Bangwei Guo, Xiao Chen, Boris Mailhe, Jia Yao, Yiqing Wang, Ankush Mukherjee, Yikang Liu, Zheyuan Zhang, Hang Yu, Terrence Chen, Shanhui Sun
arXiv Computer Vision
4d ago

InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning

InsightMap is a framework that uses top‑down maps as explicit spatial memory and action‑conditioned prediction targets for language‑guided navigation. It links historical views to labeled map locations and employs a shared multimodal backbone to jointly learn navigation action prediction and post‑action map generation, providing auxiliary training supervision. The approach supports a unified RGB‑D pipeline for navigation, visual question answering, situated reasoning, and 3D grounding, achieving state‑of‑the‑art results on R2R‑CE, RxR‑CE, ScanQA, SQA3D, ScanRefer, and outperforming baselines on the Unitree Go2 platform.

By Hongpei Zheng, Hujun Yin