ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
0
Citations
0
Influential Citations
—
Venue
2025
Year
… examples, discuss existing approaches to hallucination detection and mitigation with a focus on … in this section require access to internal weights of the model for hallucination detection. …
Hallucinations in foundation models pose a critical challenge when these models are used for decision-making, where incorrect outputs can have serious consequences. While much research has focused on hallucination in natural language generation, this paper specifically addresses the decision-making context, where the stakes are higher and the definition of hallucination may differ. By proposing a flexible definition, the authors aim to unify disparate notions of hallucination and provide a common ground for future research.
The paper also emphasizes the role of internal model access in detection methods. Many state-of-the-art techniques rely on probing internal representations or gradients, which are not available in black-box settings. This review highlights the trade-offs between white-box and black-box approaches, offering guidance for practitioners who may not have full access to model internals.
As a review paper, it does not present new experimental metrics. Instead, it synthesizes findings from existing studies, noting that white-box methods generally achieve higher detection accuracy but require model transparency. The paper does not provide specific numerical comparisons, but it qualitatively assesses the strengths and weaknesses of different approaches, such as the reliability of internal uncertainty estimates versus external consistency checks.
This paper fills a gap in the literature by focusing on hallucination detection specifically for decision-making systems, which is increasingly important as foundation models are integrated into high-stakes applications like healthcare, finance, and autonomous systems. The proposed flexible definition and taxonomy will help standardize evaluation and encourage the development of more robust detection methods. By highlighting the role of internal model access, it also sparks discussion about the trade-offs between model transparency and performance, potentially influencing future model design and regulation.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba