Modular

Modular

AI engineering teams and companies that need high-performance, portable inference for production workloads, including AI coding tool providers, speech synthesis companies, and enterprises deploying custom models.

Sunnyvale, California, USA
modular.com
Founded
2022
Headquarters
Sunnyvale, California, USA
Team size
51-200
Type
Private company
Pricing
Usage-based
Visit Website

Overview

Modular is an AI infrastructure company that builds a unified inference platform for deploying and running generative AI models across heterogeneous hardware, from GPUs to CPUs. The company was founded in 2022 and is headquartered in Sunnyvale, California. It has raised $380M in funding from investors including General Catalyst, Google Ventures, and Greylock, with a last known valuation of $1.6B. In a recent development, Qualcomm completed its acquisition of Modular, making it a Qualcomm company.

What it does

Modular provides a unified AI inference stack that spans from GPU kernels to API endpoints. Its MAX framework is a high-performance, hardware-agnostic serving framework that automatically optimizes kernels and request execution across accelerators. The platform supports running 1000+ models, including DeepSeek and Kimi, and offers deployment options in Modular's managed cloud, in a customer's VPC, or fully self-hosted. Modular also develops the Mojo programming language, a systems language designed for AI, which enables writing custom GPU kernels for maximum performance.

Key features

  • Unified inference stack from kernel to cloud
  • Hardware portability across NVIDIA, AMD, Intel, ARM, and Apple Silicon
  • Compiler-native speculative decoding for code generation
  • Custom GPU kernels written in Mojo
  • Shared and dedicated endpoints with per-token or per-minute pricing
  • Deployment options: Modular Cloud, Your Cloud (VPC), and Self-Hosted
  • Support for 1000+ open models out of the box
  • Streaming-aware scheduler for low time-to-first-token (TTFT)

Use cases

  • Serving AI coding assistants with low-latency code completion
  • Deploying text-to-speech models with reduced latency and cost
  • Running custom or fine-tuned models on optimized infrastructure
  • Agentic AI workflows requiring high-throughput batch inference
  • On-device inference for air-gapped or offline environments

Pricing

Usage-basedFree tier
PlanPrice
Free Forever Self HostedFree
Our Cloud
Your Cloud

Pricing for cloud endpoints is per token (shared) or per minute (dedicated). Specific model prices listed on the pricing page. Enterprise contracts available.

View current pricing

Pricing is gathered from the company's public pages and may change. Check the vendor's site before buying.

Pros and cons

Strengths

  • 2x performance improvement over vLLM on diverse hardware
  • Up to 70% cost savings reported by customers
  • True hardware portability with a single codebase
  • Small runtime footprint (under 700MB) enabling faster rollouts
  • Flexible deployment options including self-hosted and VPC
  • SOC 2 Type 2 certified

Limitations

  • Pricing is not fully public; requires contacting sales for enterprise plans
  • No free tier for cloud endpoints; only self-hosted is free
  • Limited to NVIDIA and AMD GPUs in Modular's cloud (more hardware coming soon)
  • Custom kernels require expertise in Mojo

What sets it apart

  • Only platform offering a unified stack from kernel to cloud
  • Compiler-native speculative decoding across all GPU targets
  • Cloud-to-device portability for coding models
  • Mojo language enables full-stack programmability

Ecosystem

Notable customers

InworldFireworksSourcegraph

Frequently asked questions

What is Modular's pricing model?

Modular offers a free self-hosted tier, and cloud endpoints are priced per token (shared) or per minute (dedicated). Specific model prices are listed on the pricing page, and enterprise contracts are available.

Can I run models on my own hardware?

Yes, Modular offers a self-hosted option where you deploy MAX and Mojo in a container on your own infrastructure, supporting NVIDIA, AMD, Intel, ARM, and Apple Silicon.

What hardware does Modular support?

Modular supports NVIDIA, AMD, Intel, ARM, and Apple Silicon for self-hosted deployments. In Modular's cloud, NVIDIA and AMD GPUs are available, with more hardware types coming soon.

Does Modular support custom models?

Yes, you can bring your own custom or fine-tuned models and deploy them on Modular's optimized infrastructure with per-minute pricing.

Is Modular SOC 2 certified?

Yes, Modular is SOC 2 Type 2 certified across all deployment options.

Sources

This profile was compiled from the company's own pages and public web research.

Last researched August 11, 2026.

Listed on:claude

Manage this listing

Own this company? Update your listing information.

Update Listing