NeMo Framework logo

NeMo Framework

Free

Generative AI framework built for researchers and PyTorch developers working on Large Language Models (LLMs), Multimodal Models (MMs), Automatic Speech Recognition (ASR), Text to Speech (TTS), and Computer Vision (CV) domains.

FreeFree tier
Inputs: text, audioOutputs: text, audio
Type
Open Source
Company
NVIDIA

About NeMo Framework

NVIDIA NeMo Speech is a scalable, open-source generative AI framework built for researchers and PyTorch developers working on speech and multimodal AI. It provides tools and pretrained models for Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech LLMs, including support for streaming inference with controllable latency. The framework has recently released models like Nemotron-3.5-ASR-Streaming (40 languages, 80ms-1s latency), Parakeet-unified-en-0.6b (offline and streaming), and Nemotron 3 VoiceChat for full-duplex conversations. It also includes MagpieTTS for 9 languages and offers early access to voice chat demos. NeMo Speech is designed to help developers create, customize, and deploy AI models efficiently using existing code and pretrained checkpoints available on HuggingFace.

Key Features

Open-source framework built for PyTorch developers
Supports Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech LLMs
Streaming inference with controllable latency (80ms-1s)
Pretrained models available on HuggingFace and NGC containers
Multimodal LLM support including voice chat and full-duplex conversations
Model releases: Nemotron, Parakeet, Canary, MagpieTTS
Language support: up to 40 languages for ASR, 9 for TTS
Early access to Nemotron 3 VoiceChat demo
Integration with NVIDIA NIM for deployment

Pros & Cons

Pros
  • Comprehensive open-source framework backed by NVIDIA
  • Extensive collection of pretrained models with state-of-the-art performance
  • Supports both offline and streaming inference
  • Active development with frequent model updates and demos
  • Strong community and integration with HuggingFace and NGC
Cons
  • Requires significant GPU resources for training and deployment
  • Documentation may be spread across multiple sources (GitHub, NGC, HuggingFace)
  • Recently split from full NeMo repository, limited to speech/audio modalities
  • Some features like VoiceChat are in early access

Best For

Real-time speech recognition and transcriptionText-to-speech synthesis for multilingual applicationsVoice chat and conversational AI with low latencyMultimodal AI research combining speech, text, and languageBuilding custom speech AI models for domain-specific tasks

FAQ

What modalities does NeMo Speech support?
NeMo Speech focuses on Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech LLMs, including multimodal LLMs that combine speech, audio, and text.
Is NeMo Speech free to use?
Yes, NeMo Speech is open-source and free to use under the provided license. Pretrained checkpoints are available on HuggingFace and NGC.
Can I use NeMo Speech for real-time streaming?
Yes, models like Nemotron-3.5-ASR-Streaming and Parakeet-unified-en-0.6b support streaming inference with configurable latency as low as 80ms.
What languages are supported?
ASR supports up to 40 languages (e.g., Nemotron-3.5-ASR-Streaming). TTS supports 9 languages including English, Spanish, German, French, Vietnamese, Italian, Chinese, Hindi, and Japanese.
Where can I find pretrained models?
Pretrained models and demos are available on NVIDIA's HuggingFace collection. Stable releases are also provided via NGC containers.