Together AI logo

Together AI

Free

Train, fine-tune-and run inference on AI models blazing fast, at low cost, and at production scale.

FreeFree tier
Inputs: text, image, audio, videoOutputs: text, image, audio, video
Type
Open Source
Company
Together AI

About Together AI

Together AI is a full-stack AI platform designed for building, training, and deploying AI models at scale, powered by cutting-edge systems research. It offers accelerated inference (2x faster), lower cost (60% savings), and faster pre-training (90% faster with Together Kernel Collection). The platform includes serverless inference, batch inference, provisioned throughput, dedicated model inference, and accelerated compute with GPU clusters. It also provides managed storage with zero egress fees, secure code sandboxes for AI agents, and fine-tuning capabilities to improve model accuracy and reduce hallucinations. Trusted by leading companies, Together AI supports a wide range of open-source models for chat, vision, image, audio, video, transcription, embeddings, and moderation tasks.

Key Features

Serverless Inference for on-demand open-source model deployment without infrastructure management
Batch Inference to cost-effectively process massive asynchronous workloads up to 30 billion tokens per model
Provisioned Throughput with token-based pricing, reserved capacity, and 99% uptime SLA
Dedicated Model Inference on dedicated infrastructure for speed and control
Accelerated Compute with Together Kernel Collection for 90% faster pre-training
Secure code sandboxes at scale for AI apps and agents
Managed Storage with object storage and parallel filesystems optimized for AI, zero egress fees
Fine-tuning open-source models using latest research techniques to reduce hallucinations and control behavior
Supports multiple modalities: chat, vision, image, audio, video, transcription, embeddings, rerank, and moderation
Cutting-edge research integration from Together Research for production AI

Pros & Cons

Pros
  • 2x faster inference and 90% faster pre-training through research-optimized kernels
  • Up to 60% lower cost through workload-specific optimization and batch pricing
  • Full-stack platform covering inference, compute, fine-tuning, storage, and sandboxes
  • Supports a wide range of open-source models and modalities (text, image, audio, video)
  • 99% uptime SLA for provisioned throughput with token-based pricing
  • Zero egress fees on managed storage for AI-native workloads

Best For

Deploying open-source models for production inference at scaleFine-tuning models to improve accuracy and reduce hallucination for specific workloadsPre-training large language models on large datasets with optimized kernel collectionBuilding AI agents with secure code sandboxes for development and executionGenerating images, videos, and audio with state-of-the-art modelsCost-effectively processing massive inference workloads asynchronously via batch APIRunning serverless inference without infrastructure management for rapid prototyping