Back to BENCHMARKS.md
Research & Science
BENCHMARKS.md · 4 documents
BENCHMARKS.md
The Low Hanging Fruit of AI Self Improvement
Presents a framework for identifying AI self-improvement opportunities by mapping tasks onto overlapping S-curves, with concrete categories and capability estimates.
ai
0
1
HunterJayPersonBENCHMARKS.md
OABench: Benchmarking Large Language Models on the Brazilian Bar Examination
Evaluates 11 LLMs on the Brazilian Bar Exam's first phase, reporting accuracy, cost, and latency across three exam editions.
aillmeval
0
7
robertotcestariBENCHMARKS.md
Large Language Models — Structured Notes
Explains LLM architecture, training pipeline, and deployment concepts from tokenization through quantization.
aillmrag
0
3
SqrtNegativOneBENCHMARKS.md
Multi-class: exactly one of the sentiment labels applies
Describes LabelFusion, a Python package that learns to combine a transformer classifier with one or more LLMs for multi-class and multi-label text classification.
aillmprompt
0
2
DataandAIReseach