All Documents
3,528 documents available
Comprehensive Evaluation Plan (Version 2)
Defines a multi-week evaluation plan for an AI receipt-scanning app, covering accuracy, latency, user testing, safety audits, and cost tracking.
Context-Aware RAG Agent: Flow & Evaluation Plan
Documents the execution flow, offline evaluation pipeline, and resource constraints for a context-aware RAG system using RAGAs metrics.
Evaluation Plan v2 — Cognify
Defines metrics, methods, golden set, user testing protocol, and timeline for evaluating a course-material-to-summary/quiz/glossary tool.
UX Review Guide for SaaS
Provides a structured UX audit checklist for SaaS apps, based on Nielsen heuristics and modern design patterns.
Evaluation and Testing LLM Systems
Lays out a four-layer evaluation stack for LLM systems, from automated assertions through human review, with concrete testing practices.
Example: LLM-as-judge for answer quality
Teaches evaluation strategies for AI agents, covering offline/online evals, RAGAS metrics, LLM-as-judge, and a complete feedback loop with tracing.
LLM Evaluation
Defines automated metrics, human evaluation, LLM-as-judge patterns, A/B testing, and regression detection for LLM applications.
Evaluating Retrieval Augmented Generation - a framework for assessment
Proposes a four-level framework for evaluating RAG systems: model, data ingestion, semantic retrieval, and end-to-end.
Instagram App Review Submission Guide
Walks through the full Instagram app review submission process with checklists, permission justifications, screencast scripts, and test scenarios.
ProtoExtract — Evaluation Approach Using OmniDocBench Methodology
Defines a multi-level evaluation framework for clinical protocol table extraction, adapting OmniDocBench metrics and adding domain-specific accuracy measures.
Evaluating the RAG answer quality
Walks through deploying an evaluation model, generating ground truth, and running bulk evaluations on RAG answer quality.
Agent and LLM Evaluation Practices
Organises LLM and agent evaluation into six dimensions, three metric families, five techniques, and AWS-native tooling.
UI Preview Guide
Lists six ways to preview UI components locally, including a dedicated mock-data preview page and optional Storybook setup.
prompt-eval-designer
Guides an LLM through designing a complete evaluation framework: criteria, rubrics, test cases, judge prompt, and decision thresholds for any application.
EVAL-001: Evaluation Contract — Flash Evaluations (FEATURE-053)
Locks scoring definitions, ground truth strategy, thresholds, and failure taxonomies for Flash Evaluations before any code is written.
Manual Review Guide for OCR Outputs
Guides manual correction of OCR errors in CSV files from scanned PDFs, with a step-by-step workflow and common error patterns.
Hybrid Preview Implementation - Complete! ✅
Documents a hybrid preview system that auto-compiles React Native mobile apps to web for instant browser previews via WebContainer.
🤖 AI-Powered Permit Review - Complete Guide
Implements an AI-powered permit review system using Anthropic Claude to analyze documents, score compliance, and suggest improvements before submission.
Document Preview & Download Feature - Complete Guide
Adds document preview and download endpoints that retrieve files from MinIO and serve them through the application with caching and security.
PR #37 Review Guide
Documents a pull request that adds graceful degradation, fixes a sandbox argument mutation bug, and upgrades FastMCP to 3.2.4.
🔍 Pull Request Review Guide
Defines a 7-step framework for reviewing Android/Kotlin pull requests, covering correctness, completeness, compatibility, consistency, clarity, and edge cases.
🚀 Guia Completo para Revisão da Meta - Garantindo Aprovação
Guides you through preparing a Facebook app review for WhatsApp Business API approval, with checklists, templates, and deployment steps.
Docker Setup and Website Preview Guide
Explains how to run a Jekyll-based research lab website locally via Docker and preview it on localhost:8080.
Tutorial Review Guide
Lists 12 tutorials across three phases with run commands, key files, and review checkpoints for a genomic visualization platform.