Major LLMs + Data Availability
FreeA reference sheet for LLM data availability and openness
About Major LLMs + Data Availability
A comprehensive Google Sheet cataloging major large language models (LLMs) and their data availability. Includes details on model size, whether the model is open-source, playground access, inference API availability, pretraining corpus, corpus public availability, fine-tuning status and corpora, documentation completeness, and links to inference and training resources. Covers models from OpenAI (davinci, text-davinci series, code-davinci), Google (LaMDA, T5, UL2, PaLM, Flan), BigScience (BLOOM, BLOOMZ, mT0), Meta (OPT, OPT-IML), EleutherAI (GPT-J, GPT-NeoX), and Galactica. Serves as a reference for researchers and practitioners to assess LLM openness and data transparency. This spreadsheet is part of the Awesome-LLM curated list (Hannibal046/Awesome-LLM).
Key Features
Pros & Cons
- Comprehensive single-page view of many important LLMs
- Clearly indicates which models are open-source (OSS)
- Documents fine-tuning datasets and their public availability
- Provides direct links to playgrounds and inference APIs
- Useful for researchers evaluating model provenance
- Static spreadsheet may not reflect the latest model releases or updates
- Relies on manual maintenance; some models may be missing
- Not interactive – no built-in filtering or search
- Inference API reliability noted as 'sort of (very slow/unreliable)' for some models