All Documents

3,528 documents available

EVALS.md

proj1rubric

Scores a software project against a 90-item rubric covering documentation, testing, community, and licensing practices.

airag
0
0
anshulp2912
EVALS.md

Model Evaluation using Gemini (GPT-Eval)

Evaluates a locally trained Gemma model's story generation using Gemini API to score grammar, creativity, consistency, plot, and instruction adherence.

aillmrag
0
0
lmassaron
EVALS.md

Sokuji Evaluation Framework

Defines a structured evaluation framework for AI translation quality using LLM-as-Judge scoring and instruction override A/B testing.

aillmeval
0
0
kizuna-ai-lab
ARCHITECTURE.md

Pull Request Review Rubric

Defines a multi-tier rubric for reviewing pull requests against AWS Well-Architected standards across seven architectural dimensions.

ai
0
2
aws-samples
ARCHITECTURE.md

QuantumPDF Chat App - System Architecture

Maps the full request flow for a PDF chat app: ingestion, caching, guardrails, vector search, model gateway, and monitoring.

airagguardrails
0
1
Kedhareswer
EVALS.md

JP Tasks

Runs Japanese language model evaluations across 12 tasks, covering QA, NLI, classification, and summarization benchmarks.

aieval
0
0
arcee-ai
SPEC.md

File Processing API Specification

Defines a Go API for secure file processing with validation, error handling, and optional performance optimizations.

ai
0
1
rejot-dev
EVALS.md

Code Challenge 4 Sanitized Rubric

Provides a rubric for grading a React code challenge, with scoring criteria for props, state, code structure, and rendering.

ai
0
1
learn-co-students
ARCHITECTURE.md

BIOF-309, Spring 2020 Rubric

Defines a grading scheme for a Python course with homework, a mini-project, and a final project, plus a detailed rubric for the final project.

ai
0
2
biof309
TASKS.md

Hybrid RAG Chatbot Tasks

Tracks a RAG chatbot's development backlog with 30+ completed and 20+ planned tasks across monitoring, retrieval, and UX.

rageval
0
0
faiqhilman13
RAG.md

API Reference Documentation

Documents the functions, parameters, and usage of a Qdrant-based RAG pipeline with CLI and utility modules.

airageval
0
1
karinehei
EVALS.md

CDR Evaluation

> How we measure output quality, and what the numbers mean.

airageval
0
2
DeepRatAI
AGENTS.md

Comprehensive Agentic AI Learning Plan for ML Engineers

Outlines an 8-week curriculum for ML engineers to build agentic AI systems, from basic agents to production deployment.

aiagentllm
0
3
girishg-dh
EVALS.md

POS Tagging with LLMs: The Hard Parts

Guides students through building and comparing classical and LLM-based POS taggers, including error analysis and segmentation challenges.

aillmprompt
0
1
ItayAbuhazera
PRD.md

Product Requirements Document: Unleash MCP Server

Defines the product vision, goals, and phased delivery plan for an MCP server that creates and wraps code changes with Unleash feature flags.

aillmeval
0
0
Unleash
EVALS.md

Fission — NEAR Track

Explains how Fission's AI agent meets NEAR track requirements for autonomous governance and data analysis on the NEAR AI Agent Hub.

aiagentllm
0
0
arkenstone-lab
ARCHITECTURE.md

QuantTradeAI - LLM Agent Guide

Defines architecture, development workflow, and coding standards for a quantitative trading ML framework with ensemble models and backtesting.

aiagentllm
0
3
AKKI0511
EVALS.md

llm.md — LLM-Assisted Semantic Repair (DP3)

Specifies an optional LLM-assisted repair step that converts DP2 FAIL outcomes into validated PATCH or NOOP proposals while preserving determinism and safety.

aillmprompt
0
0
franchyi
EVALS.md

Presentation Evaluator — Prototype Plan

Defines an 8-module pipeline that evaluates academic presentations from recorded video, extracting speech, prosody, body language, and language metrics against TED talk benchmarks.

aillmeval
0
1
nejohnson2
EVALS.md

DFAH: Determinism-Faithfulness Assurance Harness

A harness for measuring whether LLM agents produce consistent, auditable behavior when given the same input multiple times.

aiagentllm
0
4
ibm-client-engineering
RAG.md

RAG-Lighter Wiki

Documents a modular Python RAG framework with support for multiple LLMs, vector stores, embeddings, and evaluation.

aiagentllm
0
5
stargazerwh
EVALS.md

TODO: Update Tag Filtering Logic in Featured Agents

Outlines steps to deduplicate displayed filter tags in a React component by mapping aliases to a primary tag.

aiagentllm
0
0
MIAJIA
GUARDRAILS.md

Responsible AI (RAI)

Documents where AI is used, model choices, guardrails, cost controls, and risk mitigations for an LLM-based system.

aillmrag
0
0
AsadMir10
SPEC.md

Logging & Evals

Documents a trigger-based logging system, cost tracking, token analytics, and a separate evals dashboard for reviewing LLM interactions.

aiagentllm
0
2
bradwmorris
Page 120 of 147