Ollama Local
Manage and use local Ollama models. Use for model management (list/pull/remove), chat/completions, embeddings, and tool-use with local LLMs. Covers OpenClaw sub-agent integration a…
Timverhoogt
@timverhoogt
What This Skill Does
Command-line interface and Python integration for managing and using local Ollama models. Supports model listing, pulling, removing, chatting, generating completions, computing embeddings, and tool/function calling with local LLMs. Also enables spawning local model sub-agents in OpenClaw for parallel or collaborative tasks.
Replaces reliance on cloud-hosted LLM APIs by providing a local, private alternative for inference, embeddings, and agentic workflows.
When to Use It
- List all locally available Ollama models and inspect their details
- Pull a new model (e.g., llama3.1:8b) for offline use
- Chat with a local model to get quick answers without internet
- Generate text completions or embeddings for local processing
- Use function-calling models to perform tool-based tasks like weather lookups
- Spawn multiple local LLM sub-agents for parallel code review or collaborative reasoning
Install
$ openclaw skills install @timverhoogt/ollama-localOllama Local
Work with local Ollama models for inference, embeddings, and tool use.
Configuration
Set your Ollama host (defaults to http://localhost:11434):
export OLLAMA_HOST="http://localhost:11434"
# Or for remote server:
export OLLAMA_HOST="http://192.168.1.100:11434"
Quick Reference
# List models
python3 scripts/ollama.py list
# Pull a model
python3 scripts/ollama.py pull llama3.1:8b
# Remove a model
python3 scripts/ollama.py rm modelname
# Show model details
python3 scripts/ollama.py show qwen3:4b
# Chat with a model
python3 scripts/ollama.py chat qwen3:4b "What is the capital of France?"
# Chat with system prompt
python3 scripts/ollama.py chat llama3.1:8b "Review this code" -s "You are a code reviewer"
# Generate completion (non-chat)
python3 scripts/ollama.py generate qwen3:4b "Once upon a time"
# Get embeddings
python3 scripts/ollama.py embed bge-m3 "Text to embed"
Model Selection
See references/models.md for full model list and selection guide.
Quick picks:
- Fast answers:
qwen3:4b - Coding:
qwen2.5-coder:7b - General:
llama3.1:8b - Reasoning:
deepseek-r1:8b
Tool Use
Some local models support function calling. Use ollama_tools.py:
# Single request with tools
python3 scripts/ollama_tools.py single qwen2.5-coder:7b "What's the weather in Amsterdam?"
# Full tool loop (model calls tools, gets results, responds)
python3 scripts/ollama_tools.py loop qwen3:4b "Search for Python tutorials and summarize"
# Show available example tools
python3 scripts/ollama_tools.py tools
Tool-capable models: qwen2.5-coder, qwen3, llama3.1, mistral
OpenClaw Sub-Agents
Spawn local model sub-agents with sessions_spawn:
# Example: spawn a coding agent
sessions_spawn(
task="Review this Python code for bugs",
model="ollama/qwen2.5-coder:7b",
label="code-review"
)
Model path format: ollama/<model-name>
Parallel Agents (Think Tank Pattern)
Spawn multiple local agents for collaborative tasks:
agents = [
{"label": "architect", "model": "ollama/gemma3:12b", "task": "Design the system architecture"},
{"label": "coder", "model": "ollama/qwen2.5-coder:7b", "task": "Implement the core logic"},
{"label": "reviewer", "model": "ollama/llama3.1:8b", "task": "Review for bugs and improvements"},
]
for a in agents:
sessions_spawn(task=a["task"], model=a["model"], label=a["label"])
Direct API
For custom integrations, use the Ollama API directly:
# Chat
curl $OLLAMA_HOST/api/chat -d '{
"model": "qwen3:4b",
"messages": [{"role": "user", "content": "Hello"}],
"stream": false
}'
# Generate
curl $OLLAMA_HOST/api/generate -d '{
"model": "qwen3:4b",
"prompt": "Why is the sky blue?",
"stream": false
}'
# List models
curl $OLLAMA_HOST/api/tags
# Pull model
curl $OLLAMA_HOST/api/pull -d '{"name": "phi3:mini"}'
Troubleshooting
Connection refused?
- Check Ollama is running:
ollama serve - Verify OLLAMA_HOST is correct
- For remote servers, ensure firewall allows port 11434
Model not loading?
- Check VRAM: larger models may need CPU offload
- Try a smaller model first
Slow responses?
- Model may be running on CPU
- Use smaller quantization (e.g.,
:7binstead of:30b)
OpenClaw sub-agent falls back to default model?
- Ensure
ollama:defaultauth profile exists in OpenClaw config - Check model path format:
ollama/modelname:tag
Top skills in this category
Humanizer
@biostartechnologyRemove signs of AI-generated writing from text. Use when editing or reviewing text to make it sound more natural and human-written. Based on Wikipedia's comprehensive "Signs of AI writing" guide. Detects and fixes patterns including: inflated symbolism, promotional language, superficial -ing analyses, vague attributions, em dash overuse, rule of three, AI vocabulary words, negative parallelisms, and excessive conjunctive phrases.
Gmail
@byungkyuGmail API integration with managed OAuth. Read, send, and manage emails, threads, labels, and drafts. Use this skill when users want to interact with Gmail. For other third party apps, use the api-gateway skill (https://clawhub.ai/byungkyu/api-gateway).
Summarize Pro
@mkpareek0315When user asks to summarize text, articles, documents, meetings, emails, YouTube transcripts, books, PDFs, reports, conversations, or any long content. Also...
Session-logs
@guogang1024Search and analyze your own session logs (older/parent conversations) using jq.
WeChat MP CN
@guohongbin-git微信公众号监控 - 文章监控、阅读量追踪、舆情分析(WeChat Official Account)