DPO
by Stanford · free · Last verified 2026-04-24
Direct Preference Optimization aligning LLMs from preferences without a reward model.
https://arxiv.org/abs/2305.18290 ↗D
D—Poor
Adoption: C+Quality: B+Freshness: ACitations: FEngagement: F
Specifications
- License
- Proprietary
- Pricing
- free
- Capabilities
- Integrations
- Use Cases
- API Available
- No
- Tags
- alignment, fine-tuning, stanford, rlhf-free
- Added
- 2026-04-24
- Completeness
- 60%
Index Score
34Adoption
50
Quality
70
Freshness
80
Citations
0
Engagement
0