ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2024
Year
… study if model editing hurts the general abilities of LLMs. This work studies model editing in the … 2013) are employed to understand the impact of model editing on the general abilities of …
Model editing has emerged as a cost-effective alternative to full fine-tuning for updating LLMs with new knowledge or correcting errors. However, the side effects of editing on the model's general abilities have been largely underexplored. This paper addresses this gap by systematically investigating whether editing harms the model's performance on unrelated tasks. The findings are crucial for practitioners who rely on editing for rapid updates, as they reveal a potential trade-off between targeted edits and overall model quality.
The paper's significance is amplified by the growing deployment of LLMs in production environments where continuous updates are necessary. If editing degrades general abilities, it could lead to unpredictable behavior in real-world applications. The proposed regularization approach offers a simple yet effective remedy, making this work directly applicable to AI engineers and researchers seeking safe editing practices.
The paper reports that without regularization, model editing leads to noticeable degradation in general abilities, with performance drops varying across tasks and editing methods. For instance, certain editing techniques cause more harm than others. When regularization is applied, the degradation is significantly reduced, often bringing performance close to the pre-editing level, while maintaining high edit success rates. The results suggest that regularization acts as a protective mechanism, ensuring that edits are more localized and less disruptive to the model's overall knowledge.
This work has broad implications for the AI community. It underscores the importance of evaluating side effects when modifying LLMs, not just the immediate edit success. The regularization approach is model-agnostic and can be integrated into existing editing pipelines, making it a practical tool for developers. Moreover, it opens up new research directions, such as developing editing methods that inherently preserve general abilities or adaptive regularization strategies. As LLMs become more ubiquitous, ensuring their reliability post-edit is paramount, and this paper provides a foundational step toward that goal.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba