A Statistical Approach to LLM Evaluation
E. Miller
Applies statistical experiment design and analysis from other sciences to LLM evaluations, providing formulas and recommendations to reduce noise and improve informativeness.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
E. Miller
Applies statistical experiment design and analysis from other sciences to LLM evaluations, providing formulas and recommendations to reduce noise and improve informativeness.
Fabrizio Sebastiani
This survey reviews machine learning approaches to automated text categorization, covering document representation, classifier construction, and evaluation.
Jonathon Luiten, Aljos̆a Os̆ep, Patrick Dendorfer, et al.
Introduces HOTA, a unified metric for multi-object tracking that balances detection, association, and localization evaluation.
Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, et al.
A comprehensive survey of deep learning-based image captioning, covering visual encoding, text generation, training strategies, datasets, and evaluation metrics.
Ajay N. Jain, Anthony Nicholls
This paper proposes standards for evaluating computational methods in drug design, including statistical reporting, data sharing, and benchmark best practices.
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, et al.
A comprehensive survey of bias evaluation and mitigation techniques for LLMs, proposing taxonomies for metrics, datasets, and mitigation methods.
Trevor E. Carlson, Wim Heirman, Stijn Eyerman, et al.
This paper evaluates high-level mechanistic core models, introducing the IW-centric model that balances simulation speed and accuracy for many-core processors.
Md Ekrim Hossin, Sulaiman M.N
This paper reviews evaluation metrics for classification, highlighting accuracy's weaknesses and proposing five aspects for constructing better discriminator metrics.
Simon Baker, Daniel Scharstein, John Lewis, et al.
Proposes a new benchmark and evaluation methodology for optical flow algorithms, including diverse datasets and improved error metrics, to address challenges in complex natural scenes.
Jakub Swacha, Michał Gracel
A survey of 47 papers on RAG chatbots in education, analyzing their character, target support, knowledge scope, LLM, and evaluation.
Ajay Bandi, Bhavani Kongari, Roshini Naguru, et al.
A comprehensive review of agentic AI systems covering definitions, frameworks, architectures, evaluation metrics, and challenges, based on 143 primary studies.
Nourhan Ibrahim, Samar AboulEla, Ahmed Ibrahim, et al.
This survey classifies LLM-KG integration into three paradigms—KG-augmented LLMs, LLM-augmented KGs, and synergized frameworks—and evaluates their methodologies, metrics, benchmarks, and challenges.