Back to .md Directory

BENCHMARKS.md

Performance benchmarks and comparison metrics

41 documents available

Categories

Recent BENCHMARKS.md Documents

View all
BENCHMARKS.md

Testing

Documents a 3,292-test suite for a markdown vault tool, covering performance, concurrency, fuzzing, security, and retrieval benchmarks against HotpotQA and LoCoMo.

airageval
0
0
velvetmonkey
BENCHMARKS.md

index

Presents a SQL case study analyzing Northwind Traders sales data with six business questions and their query solutions.

ai
0
0
absubuh
BENCHMARKS.md

🎯 AGENTE LP CONVERTER - Landing Pages de Alta Conversão

Defines a complete workflow and design system for building high-conversion landing pages for Brazilian infoproducts, including benchmark analysis, copywriting formulas, and visual style guides.

aiagent
0
1
comeca-ai
BENCHMARKS.md

PkVision β€” Roadmap

Lists 30+ planned features and improvements for a tricking detection and scoring system, organized into six categories.

aiworkflow
0
1
AirKyzzZ
BENCHMARKS.md

πŸ“” AI Assistant Diary

Presents fictional diary entries from an AI coding assistant, reflecting on collaboration, debugging, and teaching moments.

ai
0
1
ewdlop
BENCHMARKS.md

The Low Hanging Fruit of AI Self Improvement

Presents a framework for identifying AI self-improvement opportunities by mapping tasks onto overlapping S-curves, with concrete categories and capability estimates.

ai
0
1
HunterJayPerson
BENCHMARKS.md

agentmark β€” Benchmark AI Coding Agents on Your Codebase

Defines an open-source Python CLI that benchmarks AI coding agents on a user's own codebase and tasks, producing a terminal comparison report of pass/fail, time, cost, tokens, and LLM calls.

aiagentllm
0
4
manishbabel
BENCHMARKS.md

OABench: Benchmarking Large Language Models on the Brazilian Bar Examination

Evaluates 11 LLMs on the Brazilian Bar Exam's first phase, reporting accuracy, cost, and latency across three exam editions.

aillmeval
0
7
robertotcestari