AI Models

Petals Lets Users Run Large AI Models at Home Like BitTorrent

Petals is a decentralized platform that allows users to run large language models such as Llama 3.1, Mixtral, Falcon, and BLOOM on consumer-grade hardware by sharing computational resources in a peer-to-peer network. Users load only a portion of a model and join a network of others serving the remaining parts, enabling inference speeds of up to 6 tokens per second for Llama 2 70B and 4 tokens per second for Falcon 180B. The platform supports fine-tuning, custom sampling methods, and access to hidden states, combining the convenience of an API with the flexibility of PyTorch and Hugging Face Transformers.

Neura News

Neura News

Neura Market Editorial

July 23, 20262 min read
Petals Lets Users Run Large AI Models at Home Like BitTorrent

A new platform called Petals is offering a novel way to run large language models at home using a peer-to-peer approach similar to BitTorrent. The system allows users to generate text with models such as Llama 3.1 (up to 405 billion parameters), Mixtral (8x22B), Falcon (40B+), and BLOOM (176B) using only a consumer-grade GPU or Google Colab.

How Petals Works

Instead of requiring users to load an entire large language model onto their own hardware, Petals distributes the workload across a network of participants. Each user loads only a portion of the model and then joins a network of other people serving the remaining parts. This collaborative approach makes it possible to run models that would otherwise be too large for individual consumer hardware.

Performance and Capabilities

According to the project's website, single-batch inference runs at up to 6 tokens per second for Llama 2 (70B) and up to 4 tokens per second for Falcon (180B). These speeds are sufficient for chatbots and interactive applications, the developers say.

Beyond Standard APIs

Petals offers more than what typical large language model APIs provide. Users can employ any fine-tuning and sampling methods they choose, execute custom paths through the model, or inspect its hidden states. The platform provides the comforts of a standard API while maintaining the flexibility of PyTorch and Hugging Face Transformers.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Getting Started

The project is available to try now in Google Colab, with documentation hosted on GitHub. Users can also contribute their own GPU to the network to help serve model parts for others. The development team encourages following progress through Discord or email updates, which they say will be sent only once every few months and contain no spam.

Background and Recognition

Petals is part of the BigScience research workshop, a collaborative effort focused on large language model research. The project has been featured on several notable platforms, though the specific outlets are not listed on the main page.

The network status display on the project's site currently shows a loading state with an error message indicating it cannot load the network status. The page also lists top contributors, though that section is also in a loading state.

Related on Neura Market

More from Neura News

AI Models

42 Mathematicians Urge Royal Society to Warn Government and Media About AI Existential Risk

Forty-two mathematical fellows, including Fields Medal winners Martin Hairer, Peter Scholze, and Wendelin Werner, have signed an open letter urging the Royal Society to warn the UK government and media about existential risks from advanced AI. The letter follows recent breakthroughs in which leading models solved open research problems, including a Millennium Problem. None of the signatories are affiliated with AI companies. The group warns that AI labs' estimates of existential risk above ten percent must not be dismissed as hype, and that by the time the situation becomes obvious to the public, it may be too late to act.

Sep 18·2 min read
Developer

Steve Yegge Shuts Down Gas Town After Failing to Build Anything Else With It

Steve Yegge shut down Gas Town, his ultra-vibed coding agent orchestrator, after admitting he never built anything else with it despite heavy subscription spend. Databricks reported a 60% coding spend increase after rolling out GPT-6 Astra to 3,500 engineers, OpenAI published a misalignment disclosure framework with six case reports, and Xiaomi ran MiMo-V2.6 RL training in public with live telemetry.

Sep 18·21 min read