Aligning LLMs with Direct Preference Optimization
Thu, Feb 8 · Oasis
Sign in to save, follow, or mark going.
About
We will discuss a powerful alignment technique called Direct Preference Optimisation (DPO) which was used to train Zephyr.
Tickets: Free