Skip to main content

Paperacademic-papersv1.0

DPO

by Stanford · free · Last verified 2026-04-24

Direct Preference Optimization aligning LLMs from preferences without a reward model.

https://arxiv.org/abs/2305.18290 ↗

D

D—Poor

Adoption: C+Quality: B+Freshness: ACitations: FEngagement: F

Specifications

License: Proprietary
Pricing: free
Capabilities
Integrations
Use Cases
API Available: No
Tags: alignment, fine-tuning, stanford, rlhf-free
Added: 2026-04-24
Completeness: 60%

Index Score

34

Adoption

50

Quality

70

Freshness

80

Citations

0

Engagement

0

Need this tool deployed for your team?

Get a Custom Setup

Explore the full AI ecosystem on Agents as a Service