OptionalDevOpsVersion 1.0.0

Run 150+ AI Apps from the Terminal with inference.sh CLI

Run 150+ AI apps (image, video, LLM) via inference.sh CLI.

Written by Neura Market from the official Hermes Agent documentation for Inference Sh Cli. Commands, paths, and version numbers are reproduced from the source unchanged.

Read the official documentation

The inference.sh CLI puts more than 150 AI applications at your fingertips, all from a single terminal command. No GPU, no per-provider API keys, no complex setup. If you need an image, a video, a search result, or a talking avatar, you can call it directly from your shell. This guide walks through the essential commands, the workflow that keeps you out of trouble, and the pitfalls that trip up most first-time users.

What it does

The infsh CLI is a unified front end for a large catalog of hosted AI models. Instead of signing up for a dozen different services and juggling their individual SDKs, you install one tool, authenticate once, and then run any app by its ID. The catalog covers image generation (FLUX, Reve, Seedream, Grok, Gemini image), video generation (Veo, Wan, Seedance, OmniHuman), AI-powered search (Tavily, Exa), and more. You pass a JSON input with your prompt or files, and the CLI returns structured JSON containing URLs to the generated media.

This is the tool to reach for when a user asks for something visual or generative and you do not want to manage the underlying infrastructure. It is also the answer when someone asks about inference.sh or infsh directly. The CLI handles the heavy lifting, so you can focus on the task at hand.

Before you start

Two things must be true before any command works: the infsh CLI is installed, and you are authenticated. The quickest way to check is to run:

infsh me

If the command is not found, install the CLI and log in:

curl -fsSL https://cli.inference.sh | sh
infsh login

Full setup details, including API key configuration, live in references/authentication.md. The CLI runs on Linux, macOS, and Windows, so platform support is not a concern.

Workflow

1. Always Search First

Never guess app names. The catalog changes often, and app IDs are not always intuitive. Always search to find the correct app ID before running anything:

infsh app list --search flux
infsh app list --search video
infsh app list --search image

2. Run an App

Once you have the exact app ID from the search results, run the app with the --input flag. Always include --json for machine-readable output:

infsh app run <app-id> --input '{"prompt": "your prompt here"}' --json

3. Parse the Output

The JSON output contains URLs to the generated media. Present these to the user with MEDIA: for inline display.

Common Commands

Image Generation

Image generation is the most common use case. Search first, then run the specific model you need:

# Search for image apps
infsh app list --search image

# FLUX Dev with LoRA
infsh app run falai/flux-dev-lora --input '{"prompt": "sunset over mountains", "num_images": 1}' --json

# Gemini image generation
infsh app run google/gemini-2-5-flash-image --input '{"prompt": "futuristic city", "num_images": 1}' --json

# Seedream (ByteDance)
infsh app run bytedance/seedream-5-lite --input '{"prompt": "nature scene"}' --json

# Grok Imagine (xAI)
infsh app run xai/grok-imagine-image --input '{"prompt": "abstract art"}' --json

Video Generation

Video generation follows the same pattern, but expect longer runtimes. The CLI will wait for the job to finish, so warn the user that it may take a moment:

# Search for video apps
infsh app list --search video

# Veo 3.1 (Google)
infsh app run google/veo-3-1-fast --input '{"prompt": "drone shot of coastline"}' --json

# Seedance (ByteDance)
infsh app run bytedance/seedance-1-5-pro --input '{"prompt": "dancing figure", "resolution": "1080p"}' --json

# Wan 2.5
infsh app run falai/wan-2-5 --input '{"prompt": "person walking through city"}' --json

Local File Uploads

The CLI automatically uploads local files when you provide a path in the input. This is handy for upscaling, image-to-video, or avatar generation:

# Upscale a local image
infsh app run falai/topaz-image-upscaler --input '{"image": "/path/to/photo.jpg", "upscale_factor": 2}' --json

# Image-to-video from local file
infsh app run falai/wan-2-5-i2v --input '{"image": "/path/to/image.png", "prompt": "make it move"}' --json

# Avatar with audio
infsh app run bytedance/omnihuman-1-5 --input '{"audio": "/path/to/audio.mp3", "image": "/path/to/face.jpg"}' --json

Search & Research

For AI-powered search, the CLI offers Tavily and Exa. These are useful when you need up-to-date information without scraping the web yourself:

infsh app list --search search
infsh app run tavily/tavily-search --input '{"query": "latest AI news"}' --json
infsh app run exa/exa-search --input '{"query": "machine learning papers"}' --json

Other Categories

The catalog extends beyond images and video. Search for 3D generation, text-to-speech, or social media automation:

# 3D generation
infsh app list --search 3d

# Audio / TTS
infsh app list --search tts

# Twitter/X automation
infsh app list --search twitter

Pitfalls

  1. Never guess app IDs, always run infsh app list --search first. App IDs change and new apps are added frequently.
  2. Always use --json, raw output is hard to parse. The --json flag gives structured output with URLs.
  3. Check authentication, if commands fail with auth errors, run infsh login or verify INFSH_API_KEY is set.
  4. Long-running apps, video generation can take 30-120 seconds. The terminal tool timeout should be sufficient, but warn the user it may take a moment.
  5. Input format, the --input flag takes a JSON string. Make sure to properly escape quotes.

Limits and gotchas

The biggest gotcha is the app ID. The catalog is not static; IDs change and new apps appear frequently. That is why the search-first rule is non-negotiable. Another common failure is forgetting the --json flag, which makes output much harder to parse programmatically. Authentication errors are usually a sign that your session expired or INFSH_API_KEY is not set. Finally, video generation is slow, so set expectations with the user before you kick off a job.

Reference Docs

For deeper dives, the skill ships with these references:

  • references/authentication.md, Setup, login, API keys
  • references/app-discovery.md, Searching and browsing the app catalog
  • references/running-apps.md, Running apps, input formats, output handling
  • references/cli-reference.md, Complete CLI command reference

More DevOps skills