All Documents

3,528 documents available

RAG.md

RAG Deep Dive Part 7: Evaluation and Debugging RAG Systems

Teaches you to measure and debug a RAG pipeline with retrieval metrics, generation metrics, and production monitoring, all implemented from scratch in Python.

aillmrag
0
2
Sachinchaurasiya360
SKILL.md

LLM Evaluation

Guides building an LLM eval pipeline with datasets, automated metrics, LLM-as-judge, A/B testing, and CI regression tests.

aillmrag
0
1
projectious-work
GOLDEN_SET.md

AI Tester Interview Preparation Guide

Prepares candidates for an AI Tester interview focused on LLM, API, and automation testing in the pharmaceutical industry.

aillmprompt
0
3
k21academyuk
EVALS.md

Data-Driven RAG Evaluation: Testing Qdrant Apps with Relari AI

Explains how to evaluate RAG applications using Qdrant and Relari, covering Top-K parameter tuning and Auto Prompt Optimization with code examples.

aillmrag
0
0
Kohnnn
EVALS.md

GenAI Benchmarks & Evaluation — Product-Based Companies

Understanding how to **benchmark, evaluate, and compare LLMs** is essential for roles at Google, OpenAI, Anthropic, Cohere, and AI research teams. This file covers the most important benchmarks, evaluation methodologies, and how to build custom evaluation harnesses.

aiagentllm
0
2
CodeWithDhruvX
FINE_TUNING.md

Agents and LLMs

Serves as a personal reference notebook covering LLM families, fine-tuning, RAG, agents, evaluation, and applied AI topics.

aiagentllm
0
1
btcoal
EVALS.md

Topic: Evaluation & Benchmarking

Surveys evaluation and benchmarking for LLMs, covering metrics, methods, pipelines, and trade-offs across offline and online settings.

aiagentllm
0
8
linhvuquach
EVALS.md

Day 20: Evaluation & Benchmarks 📏

Teaches evaluation of generative AI outputs using n-gram metrics, LLM-as-judge, and RAGAS, with runnable code examples for each technique.

aillmrag
0
3
Ravikiran-Bhonagiri
SKILL.md

LLM Evaluation

Guides building an LLM evaluation pipeline with automated metrics, LLM-as-judge, A/B testing, and CI regression tests.

aillmrag
0
4
projectious-work
MONITORING.md

🧠 Big Picture

Explains why API Gateway non-proxy integration with mapping templates is the correct answer for a GenAI model routing exam question, contrasting it with three incorrect options.

aillmrag
0
0
emilyg888
EVALS.md

Research Report: Using LLMs as Oracle for Entity Matching Ground Truth

Summarises research on using LLMs (DeepSeek, GPT-4, Claude) for entity matching ground truth, covering accuracy, prompt engineering, ensembles, cost analysis, validation, and regression test conversion.

aillmrag
0
2
ClaudioLutz
EVALS.md

Evaluation Framework

Defines a three-layer evaluation system for an AI investment assistant: guardrails, online RAGAS scoring, and offline golden dataset regression tests.

aiagentllm
0
2
yussaaa
EVALS.md

After LangGraph node execution, convert messages

Explains how to integrate RAGAS evaluation metrics into a LangGraph RAG pipeline, covering offline and online evaluation strategies.

aiagentllm
0
2
lowkaihon
EVALS.md

RAG System Testing Methodologies: A Comprehensive Guide

Surveys RAG evaluation frameworks, metrics, and testing methodologies with implementation guidance and tool comparisons.

aillmrag
0
5
destefani
GOLDEN_SET.md

Understanding the Sources of Uncertainty - and Why Our Evals are Biased

Catalogues nine sources of error in AI evaluations, bias and noise, with concrete examples and statistical intuition for each.

aiagentrag
0
1
reliableai
RAG.md

Project Memory

Documents a production-ready RAG system for Paris cultural events with hybrid search, LLM generation, and Streamlit UI.

aillmrag
0
0
shah-data-scientist
RUBRIC.md

Reddit Virality Grading Rubric

Defines a weighted scoring system for predicting Reddit post virality, used by an LLM to grade rumours before simulation.

aillmprompt
0
0
Riden28
DEPLOYMENT.md

Project 2: Full-Stack Application

Defines a graded full-stack project assignment requiring AI modalities, authentication, CI/CD, and Agile documentation.

airageval
0
0
JasonIngersoll9000
EVALS.md

Using Performance Metrics to Evaluate RAG Systems

Walks through evaluating RAG systems with Qdrant and Relari, covering Top-K parameter tuning and auto prompt optimization.

aillmrag
0
0
qdrant
GOLDEN_SET.md

Thesis Falsifier

Defines a RAG-based web app that generates 19-point falsification assessments of research papers from uploaded PDFs.

aillmrag
0
0
fobert789
RUBRIC.md

QualRubric

Defines a qualitative rubric for PhD qualifying exam presentations, focusing on context, organization, and knowledge demonstration rather than numerical scoring.

aieval
0
3
adsarwate
RUBRIC.md

Judging Rubric

Defines a 100-point scoring rubric for an AI-for-social-good hackathon, covering five criteria with detailed score bands and tie-breaking rules.

aievalworkflow
0
2
rudra496
RUBRIC.md

Criteria 1: Quality of Exploratory Data Analysis (20%) [20]

Defines a 3-criteria marking rubric for a logistic regression assignment, each with four grade bands and specific descriptors.

ai
0
0
iftikharafridi
GOLDEN_SET.md

Module 6: Synthesis

Guides you through synthesizing a RAG evaluation project into portfolio artifacts, stakeholder presentations, and technical handoffs.

airageval
0
0
natnew
Page 13 of 147