arXiv AI By Tianyang Zhou, Wenbo Chen, Pierre Jinghong Liang, Leman Akoglu

Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text

Read the original on arXiv AI →

arXiv:2605. 29076v2 Announce Type: replace-cross Abstract: LLMs have advanced text classification, yet existing paradigms face a trade-off: supervised (label only) fine-tuning is scalable but offers limited reasoning on complex text and lacks broader model transparency, while discrete prompt optimization offers human-readable instructions but struggles with performance and scalability.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
4d ago

Layer-Informed Fine-Tuning via Three-Stage Functional Segmentation of LLMs

The paper proposes Layer-Informed Fine-Tuning (LIFT), a method that identifies and updates only the most functionally critical layers of large language models (LLMs) using a bottleneck identification mechanism based on sensitivity analysis. By focusing on layers that handle conceptualization, reasoning, and textualization, LIFT aims to accelerate training and enhance performance on reasoning tasks. Experiments demonstrate that this selective fine-tuning approach both speeds up the training process and yields significant performance gains.

By Junning Shao, Siwei Wang, Zhixuan Fang
arXiv AI
3d ago

StateTree: Enhancing Long-Term Dialogue Reasoning via Reinforcement Learning

StateTree is a reinforcement learning approach that improves long‑term dialogue reasoning by building a tree‑structured auxiliary task from limited dialogue data. The method embeds key‑value records across multiple sessions into a binary tree, requiring the model to traverse from root to leaf, retrieve records, compare timestamps, and identify a target question among distractors. Curriculum RL training increases tree depth, and a compositional variant trains the model to combine partial reasoning fragments, enabling cross‑session retrieval, temporal reasoning, knowledge updates, and multi‑hop reasoning while generalizing from 10K‑token to 128K‑token contexts.

By Naen Xu, Wanqing Cui, Yibo Hu, Shixin Hong, Hengyu An, Meiguang Jin, Junfeng Ma, Tianyu Du