Major LLMs + Data Availability logo

Major LLMs + Data Availability

Free

A reference sheet for LLM data availability and openness

FreeFree tier
Type
Open Source

About Major LLMs + Data Availability

A comprehensive Google Sheet cataloging major large language models (LLMs) and their data availability. Includes details on model size, whether the model is open-source, playground access, inference API availability, pretraining corpus, corpus public availability, fine-tuning status and corpora, documentation completeness, and links to inference and training resources. Covers models from OpenAI (davinci, text-davinci series, code-davinci), Google (LaMDA, T5, UL2, PaLM, Flan), BigScience (BLOOM, BLOOMZ, mT0), Meta (OPT, OPT-IML), EleutherAI (GPT-J, GPT-NeoX), and Galactica. Serves as a reference for researchers and practitioners to assess LLM openness and data transparency. This spreadsheet is part of the Awesome-LLM curated list (Hannibal046/Awesome-LLM).

Key Features

Lists 23+ major LLMs with creator, model name, and size
Columns for open-source status, playground link, and inference API
Documents pretraining corpus and whether it is publicly available
Shows fine-tuning status and sources of fine-tuning data
Indicates if the model is fully documented with links to papers
Includes direct links to inference APIs and training resources
Covers both proprietary (OpenAI, Google) and open models (BigScience, Meta, EleutherAI)

Pros & Cons

Pros
  • Comprehensive single-page view of many important LLMs
  • Clearly indicates which models are open-source (OSS)
  • Documents fine-tuning datasets and their public availability
  • Provides direct links to playgrounds and inference APIs
  • Useful for researchers evaluating model provenance
Cons
  • Static spreadsheet may not reflect the latest model releases or updates
  • Relies on manual maintenance; some models may be missing
  • Not interactive – no built-in filtering or search
  • Inference API reliability noted as 'sort of (very slow/unreliable)' for some models

Best For

Comparing LLM data transparency and openness across vendorsSelecting base models for fine-tuning based on available dataResearch on the reproducibility and documentation of major language modelsQuick reference for model sizes and access methods

FAQ

What models are covered in this spreadsheet?
The sheet includes major LLMs from OpenAI (davinci, text-davinci series, code-davinci), Google (LaMDA, T5, UL2, PaLM, Flan variants), BigScience (BLOOM, BLOOMZ, mT0), Meta (OPT, OPT-IML), EleutherAI (GPT-J, GPT-NeoX), and Galactica.
Does the spreadsheet include proprietary models?
Yes, it covers proprietary models from OpenAI and Google alongside open-source ones, with notes on their data availability and documentation.
How can I access the models listed?
For each model, the sheet provides links to playgrounds, inference APIs, and training resources where available. Some models also have Hugging Face integration.
Is the data in this sheet regularly updated?
The sheet appears to be a static snapshot (part of the Awesome-LLM repository) and may not be continuously updated with new model releases.