← Back to all news
Hugging Face Blog January 18, 2024

Preference Tuning LLMs with Direct Preference Optimization Methods

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • llms

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 19

Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings

arXiv:2606. 19744v1 Announce Type: cross Abstract: Aligning language models with human preferences often requires optimising multiple behavioural objectives.

By Pranav Bhandari, Nicolas Fay, Amitava Datta, Usman Naseem, Mehwish Nasim
llmsreinforcement-learningfine-tuningsafety
More like this →
arXiv AI
Jun 12

Boosting Direct Preference Optimization with Penalization

arXiv:2606. 12505v1 Announce Type: cross Abstract: Offline preference optimization has become a practical substitute for reinforcement learning from human feedback, but pairwise objectives such as Direct Preference Optimization (DPO) and its variants use only the chosen and rejected responses stored in a static dataset.

By Pengwei Sun
llmsreinforcement-learning
More like this →
Hugging Face Blog
Sep 15, 2023

Optimizing your LLM in production

llms
More like this →
Hugging Face Blog
Jun 3

Direct Preference Optimization Beyond Chatbots

More like this →
Hugging Face Blog
Jan 10, 2024

Make LLM Fine-tuning 2x faster with Unsloth and 🤗 TRL

llmsfine-tuning
More like this →
arXiv AI
Jun 11

Autoregressive Direct Preference Optimization

arXiv:2602. 09533v2 Announce Type: replace Abstract: Direct preference optimization (DPO) has emerged as a promising approach for aligning large language models (LLMs) with human preferences.

By Masanari Oi, Mahiro Ukai, Masahiro Kaneko, Naoaki Okazaki, Nakamasa Inoue
llmsreinforcement-learning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e