All Documents

3,528 documents available

EVALS.md

RAG Evaluation Frameworks: Comprehensive Analysis for VERA

Surveys RAG evaluation frameworks and metrics, focusing on grounding verification for a verification engine called VERA.

aillmrag
0
3
manutej
EVALS.md

Air-Gapped RAG: Grounding, Citations, and Evaluation

Defines a three-layer quality loop for air-gapped RAG: grounding prompts with citations, post-generation faithfulness checks, and offline evaluation using local LLM judges.

aillmrag
0
3
agentpatterns-ai
EVALS.md

Epic 3: Working Memory, Evaluation & Production Readiness

Defines 12 stories to bring a cognitive memory MCP system into production with monitoring, resilience, cost control, and stability validation.

aievalmcp
0
1
ethrdev
EVALS.md

Knowledge MCP Query Reference for Evaluation Timing

Teaches you to query a Knowledge MCP server for evaluation placement, methods, anti-patterns, and decision support, with query templates and result interpretation.

aillmeval
0
0
philbeliveau
EVALS.md

Evaluation Methodology — Practitioner Reference

Distills evaluation methods for open-ended AI systems into actionable guidance with code examples and decision trees.

aillmrag
0
1
odewahn
EVALS.md

LLM Evaluation — Deep Dive

Covers why LLM evaluation is hard, benchmark taxonomy, LLM-as-judge methodology, contamination, robustness, and designing an eval suite for a product.

aiagentllm
0
14
ffaisal93
EVALS.md

LLM Evaluation & Benchmarking

Covers why and how to evaluate generative LLMs, from offline metrics to production A/B testing.

aillmrag
0
6
spawn08
EVALS.md

Lesson 01: Evaluation Frameworks Overview

Teaches a taxonomy of LLM evaluation approaches and compares five frameworks for offline and online testing.

aillmrag
0
16
ribatshepo
PROMPTS.md

🎯 Phase 4 特別計畫:次世代學術評估引擎實作計畫書 (v1.0)

Defines a multi-dimensional 1-10 scoring system with weighted metrics for automated academic answer evaluation, including retry logic and integration points.

aiagentllm
0
1
DeadMark70
EVALS.md

PRD-038: Evaluation, Safety, And Rollout

Defines evaluation metrics, gate criteria, and rollout phases for a face-recognition matcher upgrade, with safety invariants and required tests.

airageval
0
0
NolanFox
EVALS.md

5_Evaluation

Explains why evaluation matters for search and RAG systems, covers metrics like nDCG and Recall@K, and describes golden test sets, demos, and user feedback.

airageval
0
0
navneetkrc
EVALS.md

LLM Evaluation — Interview Grill

Presents 115 active-recall questions across 14 sections to test and reinforce knowledge of LLM evaluation concepts.

aiagentllm
0
4
ffaisal93
EVALS.md

Evalyn Roadmap

Tracks planned and completed features for the Evalyn observability and evaluation framework, organized by category.

aiagentllm
0
1
shihongDev
EVALS.md

The Evaluation & Optimization Loop

Describes a closed-loop system that generates synthetic data, applies human review, optimizes routing with DSPy, evaluates results, and annotates telemetry for continuous improvement.

aiprompteval
0
2
amit-jain
RAG.md

Evaluating Retrieval Augmented Generation - a framework for assessment

Proposes a four-level framework for evaluating RAG systems: model, data ingestion, semantic retrieval, and end-to-end.

aillmrag
0
0
superlinked
EVALS.md

Comprehensive Evaluation Plan - MVP

Defines 8 core metrics, a 50-case golden test suite, user testing protocols, regression categories, and a 12-week measurement timeline for an AI video censoring MVP.

aievalsafety
0
5
daatoo
EVALS.md

LLM as a Judge

Shows how to configure Promptfoo for LLM-as-a-judge evaluations using rubric prompts, model-graded scoring, multi-judge voting, and injection-safe judge prompts.

aillmrag
0
11
promptfoo
PROMPTS.md

🤖 機械学習・データ可視化プレビューガイド

Documents a preview system for TensorFlow MNIST training, matplotlib visualizations, and pandas analysis inside a Docker container.

ai
0
0
suetaketakaya
EVALS.md

Evaluating the RAG answer quality

Walks through deploying an evaluation model, generating ground truth, and running bulk evaluations on RAG answer quality.

airageval
0
1
bhavesh-chainani
EVALS.md

ProtoExtract — Evaluation Approach Using OmniDocBench Methodology

Defines a multi-level evaluation framework for clinical protocol table extraction, adapting OmniDocBench metrics and adding domain-specific accuracy measures.

aieval
0
2
cryogenic22
EVALS.md

Agent and LLM Evaluation Practices

Organises LLM and agent evaluation into six dimensions, three metric families, five techniques, and AWS-native tooling.

aiagentllm
0
10
luisalbertogh
SPEC.md

DR-ICU App Review Guide

Guides testers through reviewing a medical reference app for emergency and ICU professionals, covering navigation, features, and device compatibility.

ai
0
0
Anhthuhai
EVALS.md

Evaluation Plan v2 — Cognify

Defines metrics, methods, golden set, user testing protocol, and timeline for evaluating a course-material-to-summary/quiz/glossary tool.

ragprompteval
0
0
BEKATX
EVALS.md

Agent Quality & Evaluation

Defines a quality evaluation framework for AI agents covering failure modes, four pillars, evaluation hierarchy, observability, and a continuous improvement flywheel.

aiagentllm
0
6
mguinada
Page 51 of 147