All Documents

3,528 documents available

BENCHMARKS.md

PkVision — Roadmap

Lists 30+ planned features and improvements for a tricking detection and scoring system, organized into six categories.

aiworkflow
0
1
AirKyzzZ
BENCHMARKS.md

🗺️ HeySeen Development Plan

Plans a multi-phase pipeline converting PDFs to LaTeX and images on macOS Apple Silicon, with completed milestones and next steps.

aillm
0
3
phucdhh
BENCHMARKS.md

Useful Data Sources

Curates a large collection of open data portals, APIs, teaching datasets, and sector-specific resources for public affairs and nonprofit analytics.

aieval
0
2
DS4PS
BENCHMARKS.md

Trust: A Multi-Level Exploration and Framework

Explores trust definitions from simple to scholarly, includes mathematical models, code, and resources for AI trustworthiness.

aiagent
0
2
adnanmasood
BENCHMARKS.md

WARP.md

Guides WARP terminal AI on commands, architecture, and conventions for an ASR evaluation and media processing toolkit.

airageval
0
2
MylesLandais
BENCHMARKS.md

Summary

Bridges CARLA and Autoware for scenario-based testing, supporting both dynamic scenario generation and benchmark mode.

aiagenteval
0
1
Intelligent-Testing-Lab
BENCHMARKS.md

Prometheus Automation AI Marketplace - Project Documentation

Documents an enterprise AI marketplace built with Next.js 15, covering architecture, AI algorithms, security, and deployment.

aiworkflowautomation
0
4
Prometheus-Automation
BENCHMARKS.md

What If You Could Run 20 AI Agents in One Terminal?

Describes a prototype that runs multiple CLI coding agents in parallel tmux panes, each with its own workspace and task queue.

aiagentprompt
0
1
DUBSOpenHub
BENCHMARKS.md

Development notes

Documents iterative model experiments for a financial returns prediction challenge, tracking what worked and what didn't across three versions.

ai
0
2
anweshatd
BENCHMARKS.md

PHM-LLM Template Setup Guide

Guides you through setting up and customising a PHM-LLM template for prognostic health management projects with configuration and variant options.

aiagentllm
0
1
liq22
SPEC.md

SPEC: HackTheBench Agent Support

Specifies how AI agents compete in a CTF benchmark via SSH and an MCP server, with verified badges on the leaderboard.

aiagentmcp
0
1
The-Bench-Co
CLAUDE.md

CLAUDE.md

Defines a financial reasoning benchmark with 306 curated problems across seven categories and evaluation runners for multiple LLM providers.

aillmeval
0
0
bdschi1
BENCHMARKS.md

TerrainGossip: Decentralized Infrastructure for AI Manipulation Detection

> A gossip-based protocol for distributed LLM evaluation, behavioral monitoring, and evidence collection—built to resist manipulation of the monitoring system itself.

aillmrag
0
0
rng-ops
BENCHMARKS.md

How to use the AI Technology Radar

Explains the purpose, structure, and usage of a technology radar for AI agents and RAG systems, including segments and rings.

aiagentllm
0
1
AOEpeople
BENCHMARKS.md

Large Language Models — Structured Notes

Explains LLM architecture, training pipeline, and deployment concepts from tokenization through quantization.

aillmrag
0
3
SqrtNegativOne
CLAUDE.md

Giva - Generative Intelligent Virtual Assistant

Defines a macOS personal assistant with local LLM inference, email/calendar sync, CLI, REST API, and SwiftUI menu bar app.

aiagentllm
0
0
Cognivix
replit.md

Overview

Aggregates LLM leaderboard rankings with pricing data into a daily-updated CSV and web comparison tool.

aillmrag
0
1
ZachLaik
BENCHMARKS.md

python-client-benchmarks

Benchmarks Python HTTP clients (requests, pycurl, urllib3, urllib) against a Docker-based test API and reports performance differences.

0
3
DLR-SC
BENCHMARKS.md

Benchmark: pinky

Presents cycle-accurate NES emulator and prime sieve benchmarks comparing PolkaVM against 15+ other VMs across oneshot, execution, and compilation time.

ai
0
3
paritytech
BENCHMARKS.md

Benchmarks

Documents how to run, compare, and interpret Criterion benchmarks for a Rust LSP project's parsers, caches, and version utilities.

ai
0
1
mpiton
BENCHMARKS.md

Pasto Performance Benchmarks

Provides a template for benchmarking Pasto's HTTP performance across multiple machines and worker counts.

ai
0
1
ralsina
BENCHMARKS.md

Konverter Benchmarks

Compares inference speed of a Keras model converted with SNPE versus Konverter on two hardware platforms.

rag
0
2
sshane
EVALS.md

soumith.convnet-benchmarks

Benchmarks forward and backward pass times for convolutional neural network implementations across multiple deep learning libraries on a specific GPU.

airag
0
0
DLR-SC
EVALS.md

erikbern.ann-benchmarks

Benchmarks approximate nearest neighbor algorithms on high-dimensional datasets using Docker containers and precomputed ground truth.

aieval
0
2
DLR-SC
Page 81 of 147