Data & Analytics
EVALS.md Β· 41 documents
Day 20: Evaluation & Benchmarks π
root((Day 20: Evaluation & Benchmarks π))
Using Performance Metrics to Evaluate RAG Systems
title: "Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI"
Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI
url: "https://qdrant.tech/blog/qdrant-relari/"
Instructions for Claude Code: n8n Meal Feedback LLM Evaluation Workflow
Create a plan to build an n8n workflow that evaluates multiple LLM prompts for generating meal feedback using a **thinking model to generate ground truth** for comparison.
Using Performance Metrics to Evaluate RAG Systems
title: "Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI"
Client Data Mapping to Existing System
This document maps the client's provided data structure to our existing Airtable tables and identifies new tables that need to be created.
LLM Benchmark Report
This document presents the benchmark and evaluation of the LLM agent
apt_juror_5
STATUS: Non-authoritative research notes. Superseded where conflicts with architecture_source_of_truth.md and architecture_decisions_and_naming.md.
EVALS.md β LLM & RAG Evaluation Playbook
title: EVALS.md β LLM & RAG Evaluation Playbook
LLM Evaluation & Metrics β Complete Guide
> This is one of the top 5 topics tested in LLM/AI engineer interviews in 2026. Every production LLM system needs evaluation β and most candidates only know RAGAS. This guide covers the full spectrum.
LLM-as-Judge Reliability Patterns
Status: Knowledge reference
13-02-PLAN
phase: 13-retrieval-evaluation-framework
Repository Intelligence: Building the Next Generation of Agent Evaluation Data
Source: https://potpie.ai/blog/the-agent-evaluation-gap
Exact match
title: "LLM Evaluation Cheat Sheet"
Knowledge MCP Query Reference for Evaluation Timing
This document provides a practical reference for using the Knowledge MCP to research evaluation placement, methods, and anti-patterns. It shows which queries to run and what to expect from each.
5_Evaluation
1. [Importance and Challenges of Evaluation](#why-is-evaluation-so-critical-when-developing-search-and-rag-systems-with-embeddings-and-rerankers-and-what-are-the-main-challenges-involved)
Golden Set β Medical-Retrieval Probes
This document describes the construction protocol for the **medical-retrieval
Input Data Format Reference
This document describes the data formats used by the Growth Agents system for tracking experiments, hypotheses, and creative variants.
Marketing Analytics Dashboard - User Guide
- **Shows:** Total marketing spend ($36.00M)
ANALYTICS CENTER
**ΠΠΎΠ΄ΡΠ»Ρ:** 05-ANALYTICS
After-Action Report: Sports Analytics Framework
**Branch:** `claude/analyze-sports-stats-isrOh`
erikbern.ann-benchmarks
Benchmarking nearest neighbors
unsigned_map_benchmarks
Benchmark using linear keys, 0 to 2500000, no duplicates
soumith.convnet-benchmarks
Easy benchmarking of all public open-source implementations of convnets.