DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization logo

DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

Free

Reliable LLM self-verification through dual preference optimization

FreeFree tier
Type
Open Source

About DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

DuPO is a method for enabling reliable self-verification in large language models through dual preference optimization, as described in the research paper available on arXiv.