All Documents

3,528 documents available

EVALS.md

Analysis of Paper 50

Structured analysis of a research paper's use of BigCloneBench, answering 12 specific questions about dataset usage and validity.

aieval
0
0
jkrinke
EVALS.md

Day 4: Agent Quality - Observability & Evaluation

Teaches observability and evaluation for AI agents, covering logs, traces, metrics, plugins, and automated testing with the ADK framework.

aiagenteval
0
0
Sid10501
EVALS.md

evaluating-llm

Introduces evaluation concepts for LLMs and agents, then provides a TypeScript framework with scorers, experiment tracking, and JSON persistence.

aiagentllm
0
0
gmotyl
AGENTS.md

Day 4 - Podcast Transcript

Summarises a white paper on evaluating AI agent quality, covering pillars, observability, and a continuous improvement flywheel.

aiagenteval
0
0
donbr
CLAUDE.md

Code indexing for AI agents: summarization strategies and evaluation systems

Synthesises 2024-2025 research on code indexing for AI agents, covering summarisation strategies, hybrid retrieval architectures, and evaluation benchmarks.

aiagentllm
0
10
MadAppGang
EVALS.md

Training Readiness Checklist

Lists five gates to verify before launching a training or tuning run, with checkboxes for each item.

airageval
0
0
m-cahill
ARCHITECTURE.md

Page Architecture: Claru

Defines a seven-section landing page layout with animated hero, comparison grid, testimonials, and a terminal-style progressive disclosure form.

aieval
0
2
claruai
EVALS.md

From Heuristics to Hybrid: A Methodology for Building a Testable, Sequential Log Anomaly Detection Engine

Describes a sequential rule-then-ML log anomaly detector and a manual golden set evaluation to avoid circular validation.

aieval
0
2
Shreyansh1812
EVALS.md

Quick start: Evaluation sets

Walks through creating an evaluation set of questions and optional ground-truth answers to measure RAG application quality before stakeholder review.

aillmrag
0
1
epec254
FINE_TUNING.md

-PLUIE: Personalisable metric with Llm Used for Improved Evaluation

Lists 16 arXiv papers on LLM evaluation, fine-tuning, and interpretability, each with title, authors, link, and abstract.

aillmrag
0
0
CSQianDong
EVALS.md

apt_juror_5

Defines metrics, ground truth, audit, and regression plans for evaluating an apartment listing crawler and ranking system.

airageval
0
1
bcdannyboy
RAG.md

lib-ai-app-community-rag

Collects community insights, leaderboard links, and production war stories for building RAG systems, especially with multi-modal and large-scale documents.

aiagentllm
0
3
uptonking
EVALS.md

Evaluation and Benchmarking

Explains why LLM evaluation is hard, defines 10+ standard benchmarks and metrics, and provides code for pass@k, BERTScore, and Elo simulation.

aillmrag
0
2
spawn08
EVALS.md

Evaluating AI Agent Systems: Metrics, Benchmarks, and Quality Assurance (2024-2026)

Surveys 2024-2026 metrics, benchmarks, and monitoring tools for evaluating AI agent systems, with recommendations for a self-improving coding agent.

aiagentllm
0
18
oddurs
EVALS.md

Claim Extraction Evaluation Matrix

Defines a matrix evaluation framework for comparing YouTube claim extraction quality across multiple LLM models and video types.

aillmrag
0
0
GitCmurf
EVALS.md

── 数据结构定义 ──────────────────────────────────────────────

Explains why agent evaluation is harder than model evaluation and provides a five-dimension framework plus a runnable Python evaluation class.

aiagentllm
0
1
xuqi2024
EVALS.md

Chapter 12 — Verification: evaluation inside the loop

Explains why self-verification fails and defines a six-layer verification stack with binary verifiers and score-based evaluators for agent loops.

aiagentprompt
0
2
dustinober1-archive
AGENTS.md

Agent Eval - Complete UI Flow Visualization

Maps every screen of an agent evaluation application to its URL, UI layout, and the evaluation features it implements.

aiagentrag
0
4
Poornima-Bhupanagouda
EVALS.md

AI System Evaluation & Testing / Đánh Giá và Kiểm Thử Hệ Thống AI

Teaches a four-level evaluation framework, three LLM testing patterns, regression CI pipeline, and production monitoring for AI features.

aillmrag
0
5
Nhi4912
EVALS.md

Configuration Reference

Guides an LLM Evaluator agent through designing evaluation frameworks, test plans, and quality gate decisions for AI systems.

aiagentllm
0
3
philbeliveau
EVALS.md

LLM Evaluation Overview

Surveys evaluation methods for LLMs: automated metrics, human eval, AI-as-judge, and task-specific checks, with guidance on dataset design and cadence.

aillmrag
0
1
armoutihansen
EVALS.md

Traditional function - easy to test

Introduces AI model evaluation concepts, metrics, and code examples for measuring accuracy, reliability, robustness, efficiency, and safety.

aillmeval
0
1
josephstreeter
EVALS.md

Superpipe Studio

Introduces a free observability and experimentation app for Superpipe pipelines, with logging, dataset management, and experiment tracking.

eval
0
1
villagecomputing
EVALS.md

Evaluation Fundamentals

Explains why LLM evaluation is hard and how to build a golden dataset, an eval pipeline, and offline plus online monitoring.

aiagentllm
0
4
dipakkr
Page 49 of 147