Back to EVALS.md

Data & Analytics

EVALS.md Β· 41 documents

EVALS.md

Day 20: Evaluation & Benchmarks πŸ“

root((Day 20: Evaluation & Benchmarks πŸ“))

aillmrag
0
3
Ravikiran-Bhonagiri
EVALS.md

Using Performance Metrics to Evaluate RAG Systems

title: "Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI"

aillmrag
0
0
AlexisBalayre
EVALS.md

Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI

url: "https://qdrant.tech/blog/qdrant-relari/"

aillmrag
0
0
Kohnnn
EVALS.md

Instructions for Claude Code: n8n Meal Feedback LLM Evaluation Workflow

Create a plan to build an n8n workflow that evaluates multiple LLM prompts for generating meal feedback using a **thinking model to generate ground truth** for comparison.

aillmrag
0
1
B-vR
EVALS.md

Using Performance Metrics to Evaluate RAG Systems

title: "Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI"

aillmrag
0
0
qdrant
EVALS.md

Client Data Mapping to Existing System

This document maps the client's provided data structure to our existing Airtable tables and identifies new tables that need to be created.

ai
0
0
jstnrme77
EVALS.md

LLM Benchmark Report

This document presents the benchmark and evaluation of the LLM agent

aiagentllm
0
2
LucasTechAI
EVALS.md

apt_juror_5

STATUS: Non-authoritative research notes. Superseded where conflicts with architecture_source_of_truth.md and architecture_decisions_and_naming.md.

airageval
0
1
bcdannyboy
EVALS.md

EVALS.md β€” LLM & RAG Evaluation Playbook

title: EVALS.md β€” LLM & RAG Evaluation Playbook

aillmrag
0
0
framersai
EVALS.md

LLM Evaluation & Metrics β€” Complete Guide

> This is one of the top 5 topics tested in LLM/AI engineer interviews in 2026. Every production LLM system needs evaluation β€” and most candidates only know RAGAS. This guide covers the full spectrum.

aillmrag
0
7
mdrijwan123
EVALS.md

LLM-as-Judge Reliability Patterns

Status: Knowledge reference

aillmprompt
0
0
epappas
EVALS.md

13-02-PLAN

phase: 13-retrieval-evaluation-framework

aievalclaude
0
0
sebc-dev
EVALS.md

Repository Intelligence: Building the Next Generation of Agent Evaluation Data

Source: https://potpie.ai/blog/the-agent-evaluation-gap

aiagenteval
0
1
kriegcloud
EVALS.md

Exact match

title: "LLM Evaluation Cheat Sheet"

aillmprompt
0
2
tslateman
EVALS.md

Knowledge MCP Query Reference for Evaluation Timing

This document provides a practical reference for using the Knowledge MCP to research evaluation placement, methods, and anti-patterns. It shows which queries to run and what to expect from each.

aillmeval
0
0
philbeliveau
EVALS.md

5_Evaluation

1. [Importance and Challenges of Evaluation](#why-is-evaluation-so-critical-when-developing-search-and-rag-systems-with-embeddings-and-rerankers-and-what-are-the-main-challenges-involved)

airageval
0
0
navneetkrc
EVALS.md

Golden Set β€” Medical-Retrieval Probes

This document describes the construction protocol for the **medical-retrieval

airageval
0
0
DEUS-AI
EVALS.md

Input Data Format Reference

This document describes the data formats used by the Growth Agents system for tracking experiments, hypotheses, and creative variants.

aiagent
0
0
jcolano
EVALS.md

Marketing Analytics Dashboard - User Guide

- **Shows:** Total marketing spend ($36.00M)

airag
0
0
VivianAr2409
EVALS.md

ANALYTICS CENTER

**ΠœΠΎΠ΄ΡƒΠ»ΡŒ:** 05-ANALYTICS

ai
0
0
MakRusSakh
EVALS.md

After-Action Report: Sports Analytics Framework

**Branch:** `claude/analyze-sports-stats-isrOh`

aillmclaude
0
0
quarterback
EVALS.md

erikbern.ann-benchmarks

Benchmarking nearest neighbors

aieval
0
2
DLR-SC
EVALS.md

unsigned_map_benchmarks

Benchmark using linear keys, 0 to 2500000, no duplicates

0
0
p-groarke
EVALS.md

soumith.convnet-benchmarks

Easy benchmarking of all public open-source implementations of convnets.

airag
0
0
DLR-SC
Page 1 of 2