Hugging Face Trending Papers

Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning

Read the original on Hugging Face Trending Papers →

Multi-turn tool-using agents must coordinate long-horizon tool sequences while tracking dialogue state and policy constraints. Existing approaches often separate inference-time orchestration from parameter-level learning, leaving tool selection weakly structured and preference updates vulnerable to train--deployment prompt mismatch.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.