Nemotron 3 Nano Omni (free)
by NVIDIAtext+image+audio+video->text1 endpoint
Overview
NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...
Capabilities
Text generation
Image understanding
Audio processing
Video understanding
Tool use / Function calling
Extended reasoning
Long context
Top-P sampling
Deterministic seed
Modalities
Input
TextAudioImageVideo
Output
Text
Technical Specifications
Context Window
256K tokens
Max Output
65.5K tokens
Tokenizer
Other
Uptime (24h)
8913.7%
Supported Parameters
include_reasoningmax_tokensreasoningseedtemperaturetool_choicetoolstop_p