High-Speed AI Chat with GPT-OSS-120B via Cerebras Inference

This n8n workflow enables ultra-fast AI chat using OpenAI's GPT-OSS-120B model on Cerebras Inference, delivering thousands of tokens/sec and <0.5s latency for responsive apps.

This workflow provides seamless integration with Cerebras' high-performance inference platform, leveraging OpenAI's open-source GPT-OSS-120B model. It achieves industry-leading speeds of thousands of tokens per second and ultra-low latency under 0.5 seconds, allowing developers and businesses to build responsive AI applications without managing complex infrastructure or enduring slow response times common in traditional setups. The workflow consists of four streamlined nodes: a trigger for inco
Platform
n8n
Category
Development & IT
Price
$14.99
Creator
Nadia Sokolov

How to import this workflow into n8n

  1. 1Purchase or download the workflow to get the n8n workflow JSON file.
  2. 2In your n8n instance, open Workflows and choose "Import from File" (or paste the JSON with Ctrl+V on the canvas).
  3. 3Open each node marked with a credential warning and connect your own accounts and API keys.
  4. 4Run the workflow once manually to verify the data flow, then toggle it to Active.

Related Development & IT workflows

More from Nadia Sokolov

Need this deployed? We'll set it up for you.

Our automation experts deploy this workflow in your stack, connect your accounts, and verify it works — or build a custom solution from scratch.