Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity logo

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Free

Scaling transformers to trillions of parameters with sparse mixture-of-experts

FreeFree tier
Type
Open Source

About Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Switch Transformers is a research paper that introduces a sparse mixture-of-experts (MoE) architecture for scaling transformer models to trillions of parameters with simple and efficient sparsity. The paper is available as an open-access preprint on arXiv.