Modular
AI engineering teams and companies that need high-performance, portable inference for production workloads, including AI coding tool providers, speech synthesis companies, and enterprises deploying custom models.
Overview
Modular is an AI infrastructure company that builds a unified inference platform for deploying and running generative AI models across heterogeneous hardware, from GPUs to CPUs. The company was founded in 2022 and is headquartered in Sunnyvale, California. It has raised $380M in funding from investors including General Catalyst, Google Ventures, and Greylock, with a last known valuation of $1.6B. In a recent development, Qualcomm completed its acquisition of Modular, making it a Qualcomm company.
What it does
Modular provides a unified AI inference stack that spans from GPU kernels to API endpoints. Its MAX framework is a high-performance, hardware-agnostic serving framework that automatically optimizes kernels and request execution across accelerators. The platform supports running 1000+ models, including DeepSeek and Kimi, and offers deployment options in Modular's managed cloud, in a customer's VPC, or fully self-hosted. Modular also develops the Mojo programming language, a systems language designed for AI, which enables writing custom GPU kernels for maximum performance.
Key features
- Unified inference stack from kernel to cloud
- Hardware portability across NVIDIA, AMD, Intel, ARM, and Apple Silicon
- Compiler-native speculative decoding for code generation
- Custom GPU kernels written in Mojo
- Shared and dedicated endpoints with per-token or per-minute pricing
- Deployment options: Modular Cloud, Your Cloud (VPC), and Self-Hosted
- Support for 1000+ open models out of the box
- Streaming-aware scheduler for low time-to-first-token (TTFT)
Use cases
- Serving AI coding assistants with low-latency code completion
- Deploying text-to-speech models with reduced latency and cost
- Running custom or fine-tuned models on optimized infrastructure
- Agentic AI workflows requiring high-throughput batch inference
- On-device inference for air-gapped or offline environments
Pricing
| Plan | Price |
|---|---|
| Free Forever Self Hosted | Free |
| Our Cloud | — |
| Your Cloud | — |
Pricing for cloud endpoints is per token (shared) or per minute (dedicated). Specific model prices listed on the pricing page. Enterprise contracts available.
View current pricingPricing is gathered from the company's public pages and may change. Check the vendor's site before buying.
Pros and cons
Strengths
- 2x performance improvement over vLLM on diverse hardware
- Up to 70% cost savings reported by customers
- True hardware portability with a single codebase
- Small runtime footprint (under 700MB) enabling faster rollouts
- Flexible deployment options including self-hosted and VPC
- SOC 2 Type 2 certified
Limitations
- Pricing is not fully public; requires contacting sales for enterprise plans
- No free tier for cloud endpoints; only self-hosted is free
- Limited to NVIDIA and AMD GPUs in Modular's cloud (more hardware coming soon)
- Custom kernels require expertise in Mojo
What sets it apart
- Only platform offering a unified stack from kernel to cloud
- Compiler-native speculative decoding across all GPU targets
- Cloud-to-device portability for coding models
- Mojo language enables full-stack programmability
Ecosystem
Notable customers
Frequently asked questions
What is Modular's pricing model?
Modular offers a free self-hosted tier, and cloud endpoints are priced per token (shared) or per minute (dedicated). Specific model prices are listed on the pricing page, and enterprise contracts are available.
Can I run models on my own hardware?
Yes, Modular offers a self-hosted option where you deploy MAX and Mojo in a container on your own infrastructure, supporting NVIDIA, AMD, Intel, ARM, and Apple Silicon.
What hardware does Modular support?
Modular supports NVIDIA, AMD, Intel, ARM, and Apple Silicon for self-hosted deployments. In Modular's cloud, NVIDIA and AMD GPUs are available, with more hardware types coming soon.
Does Modular support custom models?
Yes, you can bring your own custom or fine-tuned models and deploy them on Modular's optimized infrastructure with per-minute pricing.
Is Modular SOC 2 certified?
Yes, Modular is SOC 2 Type 2 certified across all deployment options.
Sources
This profile was compiled from the company's own pages and public web research.
- https://www.modular.com/
- https://www.modular.com/case-studies/inworld
- https://www.modular.com/solutions/code-generation
- https://www.modular.com/pricing
- https://www.modular.com/company/report-issue
Last researched August 11, 2026.
Manage this listing
Own this company? Update your listing information.