BENCHMARKS.md
Performance benchmarks and comparison metrics
41 documents availableCategories
Recent BENCHMARKS.md Documents
View allTesting
Documents a 3,292-test suite for a markdown vault tool, covering performance, concurrency, fuzzing, security, and retrieval benchmarks against HotpotQA and LoCoMo.
index
Presents a SQL case study analyzing Northwind Traders sales data with six business questions and their query solutions.
π― AGENTE LP CONVERTER - Landing Pages de Alta ConversΓ£o
Defines a complete workflow and design system for building high-conversion landing pages for Brazilian infoproducts, including benchmark analysis, copywriting formulas, and visual style guides.
PkVision β Roadmap
Lists 30+ planned features and improvements for a tricking detection and scoring system, organized into six categories.
π AI Assistant Diary
Presents fictional diary entries from an AI coding assistant, reflecting on collaboration, debugging, and teaching moments.
The Low Hanging Fruit of AI Self Improvement
Presents a framework for identifying AI self-improvement opportunities by mapping tasks onto overlapping S-curves, with concrete categories and capability estimates.
agentmark β Benchmark AI Coding Agents on Your Codebase
Defines an open-source Python CLI that benchmarks AI coding agents on a user's own codebase and tasks, producing a terminal comparison report of pass/fail, time, cost, tokens, and LLM calls.
OABench: Benchmarking Large Language Models on the Brazilian Bar Examination
Evaluates 11 LLMs on the Brazilian Bar Exam's first phase, reporting accuracy, cost, and latency across three exam editions.