All Documents
3,528 documents available
RAG Evaluation Frameworks: Comprehensive Analysis for VERA
Surveys RAG evaluation frameworks and metrics, focusing on grounding verification for a verification engine called VERA.
Air-Gapped RAG: Grounding, Citations, and Evaluation
Defines a three-layer quality loop for air-gapped RAG: grounding prompts with citations, post-generation faithfulness checks, and offline evaluation using local LLM judges.
Epic 3: Working Memory, Evaluation & Production Readiness
Defines 12 stories to bring a cognitive memory MCP system into production with monitoring, resilience, cost control, and stability validation.
Knowledge MCP Query Reference for Evaluation Timing
Teaches you to query a Knowledge MCP server for evaluation placement, methods, anti-patterns, and decision support, with query templates and result interpretation.
Evaluation Methodology — Practitioner Reference
Distills evaluation methods for open-ended AI systems into actionable guidance with code examples and decision trees.
LLM Evaluation — Deep Dive
Covers why LLM evaluation is hard, benchmark taxonomy, LLM-as-judge methodology, contamination, robustness, and designing an eval suite for a product.
LLM Evaluation & Benchmarking
Covers why and how to evaluate generative LLMs, from offline metrics to production A/B testing.
Lesson 01: Evaluation Frameworks Overview
Teaches a taxonomy of LLM evaluation approaches and compares five frameworks for offline and online testing.
🎯 Phase 4 特別計畫:次世代學術評估引擎實作計畫書 (v1.0)
Defines a multi-dimensional 1-10 scoring system with weighted metrics for automated academic answer evaluation, including retry logic and integration points.
PRD-038: Evaluation, Safety, And Rollout
Defines evaluation metrics, gate criteria, and rollout phases for a face-recognition matcher upgrade, with safety invariants and required tests.
5_Evaluation
Explains why evaluation matters for search and RAG systems, covers metrics like nDCG and Recall@K, and describes golden test sets, demos, and user feedback.
LLM Evaluation — Interview Grill
Presents 115 active-recall questions across 14 sections to test and reinforce knowledge of LLM evaluation concepts.
Evalyn Roadmap
Tracks planned and completed features for the Evalyn observability and evaluation framework, organized by category.
The Evaluation & Optimization Loop
Describes a closed-loop system that generates synthetic data, applies human review, optimizes routing with DSPy, evaluates results, and annotates telemetry for continuous improvement.
Evaluating Retrieval Augmented Generation - a framework for assessment
Proposes a four-level framework for evaluating RAG systems: model, data ingestion, semantic retrieval, and end-to-end.
Comprehensive Evaluation Plan - MVP
Defines 8 core metrics, a 50-case golden test suite, user testing protocols, regression categories, and a 12-week measurement timeline for an AI video censoring MVP.
LLM as a Judge
Shows how to configure Promptfoo for LLM-as-a-judge evaluations using rubric prompts, model-graded scoring, multi-judge voting, and injection-safe judge prompts.
🤖 機械学習・データ可視化プレビューガイド
Documents a preview system for TensorFlow MNIST training, matplotlib visualizations, and pandas analysis inside a Docker container.
Evaluating the RAG answer quality
Walks through deploying an evaluation model, generating ground truth, and running bulk evaluations on RAG answer quality.
ProtoExtract — Evaluation Approach Using OmniDocBench Methodology
Defines a multi-level evaluation framework for clinical protocol table extraction, adapting OmniDocBench metrics and adding domain-specific accuracy measures.
Agent and LLM Evaluation Practices
Organises LLM and agent evaluation into six dimensions, three metric families, five techniques, and AWS-native tooling.
DR-ICU App Review Guide
Guides testers through reviewing a medical reference app for emergency and ICU professionals, covering navigation, features, and device compatibility.
Evaluation Plan v2 — Cognify
Defines metrics, methods, golden set, user testing protocol, and timeline for evaluating a course-material-to-summary/quiz/glossary tool.
Agent Quality & Evaluation
Defines a quality evaluation framework for AI agents covering failure modes, four pillars, evaluation hierarchy, observability, and a continuous improvement flywheel.