PreprintarXiv.org2025
Inoculation Prompting (IP)
Nevan Wichers, Aram Ebtekar, Ariana Azarbal, et al.
Inoculation Prompting (IP) prevents learning of undesired behaviors in LLMs by modifying training prompts to explicitly request the undesired behavior, reducing reward hacking and sycophancy without harming desired capabilities.
29Oct 6, 2025Large Language ModelsFine Tuning
arXiv