arXiv AI By Darsh Kachroo, Arjun Prasaath Anbazhagan, Adriana Caraeni, Brennan Lagasse, Kevin Zhu

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs

Read the original on arXiv AI →

arXiv:2604. 20140v2 Announce Type: replace Abstract: Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles with complex reasoning tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
3d ago

Offline Guidance, Online Reasoning: Reusing LLM Feedback for Small Language Models

arXiv:2609.39346v1 Announce Type: new Abstract: Large language models (LLMs) offer strong reasoning capabilities but are often costly to access through commercial APIs, while small language models (S...

By Bohan Zhang (Southeast University), Linan Yue (Southeast University), Weibo Gao (Hong Kong Polytechnic University), Pengyu Chen (Southeast University), Hong Guo (Southeast University), Yanqi Hao (ZTE Corporation)