All Documents

41 documents available

BENCHMARKS.md

Testing

Documents a 3,292-test suite for a markdown vault tool, covering performance, concurrency, fuzzing, security, and retrieval benchmarks against HotpotQA and LoCoMo.

airageval
0
0
velvetmonkey
BENCHMARKS.md

index

Presents a SQL case study analyzing Northwind Traders sales data with six business questions and their query solutions.

ai
0
0
absubuh
BENCHMARKS.md

🎯 AGENTE LP CONVERTER - Landing Pages de Alta Conversão

Defines a complete workflow and design system for building high-conversion landing pages for Brazilian infoproducts, including benchmark analysis, copywriting formulas, and visual style guides.

aiagent
0
1
comeca-ai
BENCHMARKS.md

PkVision — Roadmap

Lists 30+ planned features and improvements for a tricking detection and scoring system, organized into six categories.

aiworkflow
0
1
AirKyzzZ
BENCHMARKS.md

🗺️ HeySeen Development Plan

Plans a multi-phase pipeline converting PDFs to LaTeX and images on macOS Apple Silicon, with completed milestones and next steps.

aillm
0
3
phucdhh
BENCHMARKS.md

📔 AI Assistant Diary

Presents fictional diary entries from an AI coding assistant, reflecting on collaboration, debugging, and teaching moments.

ai
0
1
ewdlop
BENCHMARKS.md

The Low Hanging Fruit of AI Self Improvement

Presents a framework for identifying AI self-improvement opportunities by mapping tasks onto overlapping S-curves, with concrete categories and capability estimates.

ai
0
1
HunterJayPerson
BENCHMARKS.md

agentmark — Benchmark AI Coding Agents on Your Codebase

Defines an open-source Python CLI that benchmarks AI coding agents on a user's own codebase and tasks, producing a terminal comparison report of pass/fail, time, cost, tokens, and LLM calls.

aiagentllm
0
4
manishbabel
BENCHMARKS.md

OABench: Benchmarking Large Language Models on the Brazilian Bar Examination

Evaluates 11 LLMs on the Brazilian Bar Exam's first phase, reporting accuracy, cost, and latency across three exam editions.

aillmeval
0
7
robertotcestari
BENCHMARKS.md

PHM-LLM Template Setup Guide

Guides you through setting up and customising a PHM-LLM template for prognostic health management projects with configuration and variant options.

aiagentllm
0
1
liq22
BENCHMARKS.md

Trust: A Multi-Level Exploration and Framework

Explores trust definitions from simple to scholarly, includes mathematical models, code, and resources for AI trustworthiness.

aiagent
0
1
adnanmasood
BENCHMARKS.md

WARP.md

Guides WARP terminal AI on commands, architecture, and conventions for an ASR evaluation and media processing toolkit.

airageval
0
2
MylesLandais
BENCHMARKS.md

What If You Could Run 20 AI Agents in One Terminal?

Describes a prototype that runs multiple CLI coding agents in parallel tmux panes, each with its own workspace and task queue.

aiagentprompt
0
1
DUBSOpenHub
BENCHMARKS.md

TerrainGossip: Decentralized Infrastructure for AI Manipulation Detection

> A gossip-based protocol for distributed LLM evaluation, behavioral monitoring, and evidence collection—built to resist manipulation of the monitoring system itself.

aillmrag
0
0
rng-ops
BENCHMARKS.md

Useful Data Sources

Curates a large collection of open data portals, APIs, teaching datasets, and sector-specific resources for public affairs and nonprofit analytics.

aieval
0
2
DS4PS
BENCHMARKS.md

Development notes

Documents iterative model experiments for a financial returns prediction challenge, tracking what worked and what didn't across three versions.

ai
0
2
anweshatd
BENCHMARKS.md

Prometheus Automation AI Marketplace - Project Documentation

Documents an enterprise AI marketplace built with Next.js 15, covering architecture, AI algorithms, security, and deployment.

aiworkflowautomation
0
4
Prometheus-Automation
BENCHMARKS.md

Summary

Bridges CARLA and Autoware for scenario-based testing, supporting both dynamic scenario generation and benchmark mode.

aiagenteval
0
1
Intelligent-Testing-Lab
BENCHMARKS.md

How to use the AI Technology Radar

Explains the purpose, structure, and usage of a technology radar for AI agents and RAG systems, including segments and rings.

aiagentllm
0
1
AOEpeople
BENCHMARKS.md

Large Language Models — Structured Notes

Explains LLM architecture, training pipeline, and deployment concepts from tokenization through quantization.

aillmrag
0
3
SqrtNegativOne
BENCHMARKS.md

Benchmarks

Compares.NET serializer performance using an echo benchmark with a custom MessageEnvelope type, showing throughput and size trade-offs.

ai
0
1
ReferenceType
BENCHMARKS.md

Speed Evaluation

Presents cycle counts, memory usage, and code size for Kyber and NewHope KEM variants on an ARM Cortex-M4 target.

eval
0
3
erdemalkim
BENCHMARKS.md

Konverter Benchmarks

Compares inference speed of a Keras model converted with SNPE versus Konverter on two hardware platforms.

rag
0
2
sshane
BENCHMARKS.md

Benchmarks

Presents benchmark results comparing CacheManager handle performance (Dictionary, Runtime, MsMemory, Redis) and serializer throughput using BenchmarkDotNet.

ai
0
2
MichaCo
Page 1 of 2