All Documents
3,528 documents available
PkVision — Roadmap
Lists 30+ planned features and improvements for a tricking detection and scoring system, organized into six categories.
🗺️ HeySeen Development Plan
Plans a multi-phase pipeline converting PDFs to LaTeX and images on macOS Apple Silicon, with completed milestones and next steps.
Useful Data Sources
Curates a large collection of open data portals, APIs, teaching datasets, and sector-specific resources for public affairs and nonprofit analytics.
Trust: A Multi-Level Exploration and Framework
Explores trust definitions from simple to scholarly, includes mathematical models, code, and resources for AI trustworthiness.
WARP.md
Guides WARP terminal AI on commands, architecture, and conventions for an ASR evaluation and media processing toolkit.
Summary
Bridges CARLA and Autoware for scenario-based testing, supporting both dynamic scenario generation and benchmark mode.
Prometheus Automation AI Marketplace - Project Documentation
Documents an enterprise AI marketplace built with Next.js 15, covering architecture, AI algorithms, security, and deployment.
What If You Could Run 20 AI Agents in One Terminal?
Describes a prototype that runs multiple CLI coding agents in parallel tmux panes, each with its own workspace and task queue.
Development notes
Documents iterative model experiments for a financial returns prediction challenge, tracking what worked and what didn't across three versions.
PHM-LLM Template Setup Guide
Guides you through setting up and customising a PHM-LLM template for prognostic health management projects with configuration and variant options.
SPEC: HackTheBench Agent Support
Specifies how AI agents compete in a CTF benchmark via SSH and an MCP server, with verified badges on the leaderboard.
CLAUDE.md
Defines a financial reasoning benchmark with 306 curated problems across seven categories and evaluation runners for multiple LLM providers.
TerrainGossip: Decentralized Infrastructure for AI Manipulation Detection
> A gossip-based protocol for distributed LLM evaluation, behavioral monitoring, and evidence collection—built to resist manipulation of the monitoring system itself.
How to use the AI Technology Radar
Explains the purpose, structure, and usage of a technology radar for AI agents and RAG systems, including segments and rings.
Large Language Models — Structured Notes
Explains LLM architecture, training pipeline, and deployment concepts from tokenization through quantization.
Giva - Generative Intelligent Virtual Assistant
Defines a macOS personal assistant with local LLM inference, email/calendar sync, CLI, REST API, and SwiftUI menu bar app.
Overview
Aggregates LLM leaderboard rankings with pricing data into a daily-updated CSV and web comparison tool.
python-client-benchmarks
Benchmarks Python HTTP clients (requests, pycurl, urllib3, urllib) against a Docker-based test API and reports performance differences.
Benchmark: pinky
Presents cycle-accurate NES emulator and prime sieve benchmarks comparing PolkaVM against 15+ other VMs across oneshot, execution, and compilation time.
Benchmarks
Documents how to run, compare, and interpret Criterion benchmarks for a Rust LSP project's parsers, caches, and version utilities.
Pasto Performance Benchmarks
Provides a template for benchmarking Pasto's HTTP performance across multiple machines and worker counts.
Konverter Benchmarks
Compares inference speed of a Keras model converted with SNPE versus Konverter on two hardware platforms.
soumith.convnet-benchmarks
Benchmarks forward and backward pass times for convolutional neural network implementations across multiple deep learning libraries on a specific GPU.
erikbern.ann-benchmarks
Benchmarks approximate nearest neighbor algorithms on high-dimensional datasets using Docker containers and precomputed ground truth.