PreprintarXiv.org2024
Mitigating LLM Jailbreaks with Few Examples
Alwin Peng, Julian Michael, Henry Sleight, et al.
Proposes rapid response defenses that block entire classes of LLM jailbreaks after observing only a few examples, achieving a 240x reduction in attack success rate.
15Nov 12, 2024Large Language ModelsFine Tuning
arXiv