All Documents
41 documents available
Testing
Documents a 3,292-test suite for a markdown vault tool, covering performance, concurrency, fuzzing, security, and retrieval benchmarks against HotpotQA and LoCoMo.
index
Presents a SQL case study analyzing Northwind Traders sales data with six business questions and their query solutions.
🎯 AGENTE LP CONVERTER - Landing Pages de Alta Conversão
Defines a complete workflow and design system for building high-conversion landing pages for Brazilian infoproducts, including benchmark analysis, copywriting formulas, and visual style guides.
PkVision — Roadmap
Lists 30+ planned features and improvements for a tricking detection and scoring system, organized into six categories.
🗺️ HeySeen Development Plan
Plans a multi-phase pipeline converting PDFs to LaTeX and images on macOS Apple Silicon, with completed milestones and next steps.
📔 AI Assistant Diary
Presents fictional diary entries from an AI coding assistant, reflecting on collaboration, debugging, and teaching moments.
The Low Hanging Fruit of AI Self Improvement
Presents a framework for identifying AI self-improvement opportunities by mapping tasks onto overlapping S-curves, with concrete categories and capability estimates.
agentmark — Benchmark AI Coding Agents on Your Codebase
Defines an open-source Python CLI that benchmarks AI coding agents on a user's own codebase and tasks, producing a terminal comparison report of pass/fail, time, cost, tokens, and LLM calls.
OABench: Benchmarking Large Language Models on the Brazilian Bar Examination
Evaluates 11 LLMs on the Brazilian Bar Exam's first phase, reporting accuracy, cost, and latency across three exam editions.
PHM-LLM Template Setup Guide
Guides you through setting up and customising a PHM-LLM template for prognostic health management projects with configuration and variant options.
Trust: A Multi-Level Exploration and Framework
Explores trust definitions from simple to scholarly, includes mathematical models, code, and resources for AI trustworthiness.
WARP.md
Guides WARP terminal AI on commands, architecture, and conventions for an ASR evaluation and media processing toolkit.
What If You Could Run 20 AI Agents in One Terminal?
Describes a prototype that runs multiple CLI coding agents in parallel tmux panes, each with its own workspace and task queue.
TerrainGossip: Decentralized Infrastructure for AI Manipulation Detection
> A gossip-based protocol for distributed LLM evaluation, behavioral monitoring, and evidence collection—built to resist manipulation of the monitoring system itself.
Useful Data Sources
Curates a large collection of open data portals, APIs, teaching datasets, and sector-specific resources for public affairs and nonprofit analytics.
Development notes
Documents iterative model experiments for a financial returns prediction challenge, tracking what worked and what didn't across three versions.
Prometheus Automation AI Marketplace - Project Documentation
Documents an enterprise AI marketplace built with Next.js 15, covering architecture, AI algorithms, security, and deployment.
Summary
Bridges CARLA and Autoware for scenario-based testing, supporting both dynamic scenario generation and benchmark mode.
How to use the AI Technology Radar
Explains the purpose, structure, and usage of a technology radar for AI agents and RAG systems, including segments and rings.
Large Language Models — Structured Notes
Explains LLM architecture, training pipeline, and deployment concepts from tokenization through quantization.
Benchmarks
Compares.NET serializer performance using an echo benchmark with a custom MessageEnvelope type, showing throughput and size trade-offs.
Speed Evaluation
Presents cycle counts, memory usage, and code size for Kyber and NewHope KEM variants on an ARM Cortex-M4 target.
Konverter Benchmarks
Compares inference speed of a Keras model converted with SNPE versus Konverter on two hardware platforms.
Benchmarks
Presents benchmark results comparing CacheManager handle performance (Dictionary, Runtime, MsMemory, Redis) and serializer throughput using BenchmarkDotNet.