Py Gpt logo

Py Gpt

Free

Desktop AI Assistant powered by GPT-5, GPT-4, o1, o3, Gemini, Claude, Ollama, DeepSeek, Perplexity, Grok, Bielik, chat, vision, voice, RAG, image and video generation, agents, tools, MCP, plugins, speech synthesis and recognition, web search, memory, presets, assistants,and more. Linux, Windows, Mac

FreeFree tier
Inputs: text, image, audio, video, codeOutputs: text, image, video
Type
Open Source

About Py Gpt

PyGPT is an open-source desktop AI assistant for Windows, macOS, and Linux, designed to run locally like a ChatGPT client but with extensive customization and multi-model support. It integrates with numerous large language models including OpenAI GPT-5, GPT-4, o1, o3, o4, Google Gemini, Anthropic Claude, xAI Grok, DeepSeek V3/R1, Perplexity Sonar, Mistral AI, and any model accessible through Ollama and LlamaIndex. The assistant offers 12 modes of operation: Chat, Chat with Files, Realtime + Audio, Research (via Perplexity), Completion, Image and Video Generation, Vision, Assistants, Experts, Computer Use, Agents, and Autonomous Mode. It supports RAG (retrieval-augmented generation) with built-in vector databases and LlamaIndex for chatting with local files (txt, pdf, csv, html, md, docx, json, epub, xlsx, xml, webpages, and more). Additional features include voice interaction (speech synthesis and recognition via multiple providers), image and video generation (DALL-E, Imagen, Gemini, Veo 3, Sora 2), internet search (Google, Bing, DuckDuckGo), plugin system, MCP support, real-time Python code interpreter, task scheduler, custom commands, calendar, notepad, painter tool, Agents Builder, themes, and accessibility support. The full source code is available on GitHub, and it requires no prior AI knowledge to use.

Key Features

Supports multiple LLMs: GPT-5, GPT-4, o1, o3, o4, Gemini, Claude, Grok, DeepSeek, Perplexity, Mistral, and any Ollama/LlamaIndex model
12 modes of operation: Chat, Chat with Files, Realtime + Audio, Research, Completion, Image/Video generation, Vision, Assistants, Experts, Computer Use, Agents, Autonomous Mode
Built-in RAG with LlamaIndex and vector databases for chatting with local files (txt, pdf, csv, html, md, docx, json, epub, xlsx, xml, webpages, etc.)
Voice interaction: speech recognition (Whisper, Google, Microsoft) and speech synthesis (Azure, Google, Eleven Labs, OpenAI TTS)
Image and video generation via DALL-E, gpt-image, Imagen, Gemini, Nano Banana, Veo 3, Sora 2
Internet search integration with Google, Bing, DuckDuckGo
Plugin system with built-in plugins (Files I/O, Code Interpreter, Web Search, Google, Facebook, X/Twitter, Slack, Telegram, GitHub) and MCP support
Real-time Python code interpreter with syntax highlighting
Task scheduler (crontab) and custom command execution
Long-term memory with context history and revert capability

Pros & Cons

Pros
  • 100% free and open-source (MIT license)
  • Supports a wide range of leading AI models from multiple providers
  • Runs locally on desktop (Windows, macOS, Linux) for privacy
  • Comprehensive local file RAG with vector database support
  • Extensive plugin ecosystem and MCP support for integration
  • Built-in voice interaction and speech synthesis
  • Highly customizable with presets, themes, and model editor
  • Includes tools like Python code interpreter, task scheduler, and calendar
Cons
  • Requires API keys for most cloud-based models (OpenAI, Anthropic, etc.)
  • Local installation and configuration may be complex for non-technical users
  • Some advanced features (e.g., Computer Use, Autonomous Mode) require familiarity with the tool
  • Dependency on external model providers for best performance
  • Only available as a desktop application; no web or mobile version

Best For

Natural conversation with AI using multiple modelsChat with personal documents and files (PDFs, spreadsheets, code, etc.)Image and video analysis through vision models and camera captureAutonomous task execution via Agents and Computer Use modeAudio-based interactions and voice-controlled assistanceIn-depth research using Perplexity and OpenAI research modelsGenerating images and videos with DALL-E, Imagen, Sora 2, and othersReal-time code generation and execution with Python interpreter

FAQ

What models does PyGPT support?
PyGPT supports OpenAI GPT-5, GPT-4, o1, o3, o4, Sora2, Google Gemini, Anthropic Claude, xAI Grok, DeepSeek V3/R1, Perplexity Sonar, Mistral AI, and any model accessible through Ollama and LlamaIndex.
Is PyGPT free?
Yes, PyGPT is completely free and open-source. The source code is available on GitHub.
Which operating systems does PyGPT support?
PyGPT runs on Windows, macOS, and Linux.
Can I use PyGPT without an internet connection?
Some features (like chatting with local files and using local Ollama models) work offline, but many models (GPT, Claude, Gemini) require an internet connection and API keys.
How do I get API keys for the supported models?
You need to obtain API keys from each provider (e.g., OpenAI, Anthropic, Google). PyGPT provides an interface to enter and manage these keys.