All Documents
3,528 documents available
proj1rubric
Scores a software project against a 90-item rubric covering documentation, testing, community, and licensing practices.
Model Evaluation using Gemini (GPT-Eval)
Evaluates a locally trained Gemma model's story generation using Gemini API to score grammar, creativity, consistency, plot, and instruction adherence.
Sokuji Evaluation Framework
Defines a structured evaluation framework for AI translation quality using LLM-as-Judge scoring and instruction override A/B testing.
Pull Request Review Rubric
Defines a multi-tier rubric for reviewing pull requests against AWS Well-Architected standards across seven architectural dimensions.
QuantumPDF Chat App - System Architecture
Maps the full request flow for a PDF chat app: ingestion, caching, guardrails, vector search, model gateway, and monitoring.
JP Tasks
Runs Japanese language model evaluations across 12 tasks, covering QA, NLI, classification, and summarization benchmarks.
File Processing API Specification
Defines a Go API for secure file processing with validation, error handling, and optional performance optimizations.
Code Challenge 4 Sanitized Rubric
Provides a rubric for grading a React code challenge, with scoring criteria for props, state, code structure, and rendering.
BIOF-309, Spring 2020 Rubric
Defines a grading scheme for a Python course with homework, a mini-project, and a final project, plus a detailed rubric for the final project.
Hybrid RAG Chatbot Tasks
Tracks a RAG chatbot's development backlog with 30+ completed and 20+ planned tasks across monitoring, retrieval, and UX.
API Reference Documentation
Documents the functions, parameters, and usage of a Qdrant-based RAG pipeline with CLI and utility modules.
CDR Evaluation
> How we measure output quality, and what the numbers mean.
Comprehensive Agentic AI Learning Plan for ML Engineers
Outlines an 8-week curriculum for ML engineers to build agentic AI systems, from basic agents to production deployment.
POS Tagging with LLMs: The Hard Parts
Guides students through building and comparing classical and LLM-based POS taggers, including error analysis and segmentation challenges.
Product Requirements Document: Unleash MCP Server
Defines the product vision, goals, and phased delivery plan for an MCP server that creates and wraps code changes with Unleash feature flags.
Fission — NEAR Track
Explains how Fission's AI agent meets NEAR track requirements for autonomous governance and data analysis on the NEAR AI Agent Hub.
QuantTradeAI - LLM Agent Guide
Defines architecture, development workflow, and coding standards for a quantitative trading ML framework with ensemble models and backtesting.
llm.md — LLM-Assisted Semantic Repair (DP3)
Specifies an optional LLM-assisted repair step that converts DP2 FAIL outcomes into validated PATCH or NOOP proposals while preserving determinism and safety.
Presentation Evaluator — Prototype Plan
Defines an 8-module pipeline that evaluates academic presentations from recorded video, extracting speech, prosody, body language, and language metrics against TED talk benchmarks.
DFAH: Determinism-Faithfulness Assurance Harness
A harness for measuring whether LLM agents produce consistent, auditable behavior when given the same input multiple times.
RAG-Lighter Wiki
Documents a modular Python RAG framework with support for multiple LLMs, vector stores, embeddings, and evaluation.
TODO: Update Tag Filtering Logic in Featured Agents
Outlines steps to deduplicate displayed filter tags in a React component by mapping aliases to a primary tag.
Responsible AI (RAI)
Documents where AI is used, model choices, guardrails, cost controls, and risk mitigations for an LLM-based system.
Logging & Evals
Documents a trigger-based logging system, cost tracking, token analytics, and a separate evals dashboard for reviewing LLM interactions.