Analyze Images, Videos, Docs & Audio with Gemini + Qwen Agent
Multimodal workflow analyzes uploaded images, videos, audio, and documents using Google Gemini tools and a text-only Qwen LLM agent for cost-efficient insights.
This workflow enables comprehensive multimodal file analysis through a chat interface. Users upload images, videos, audio files, or documents, which are automatically uploaded to Google Gemini to generate accessible URLs. A lightweight text-only Qwen 32B LLM agent (via Groq) then dynamically creates contextual prompts and invokes specialized Gemini tools based on file types and user queries, delivering concise, helpful responses without relying on expensive full multimodal models.
Key benefits
- Platform
- n8n
- Category
- Other
- Price
- $24.99
- Creator
- Andre Vermeulen
- AI
- Gemini
- Qwen
- Multimodal
- Image Analysis
- Video Processing
- Audio Analysis
- Document OCR
- LLM Agent
- Groq
How to import this workflow into n8n
- 1Purchase or download the workflow to get the n8n workflow JSON file.
- 2In your n8n instance, open Workflows and choose "Import from File" (or paste the JSON with Ctrl+V on the canvas).
- 3Open each node marked with a credential warning and connect your own accounts and API keys.
- 4Run the workflow once manually to verify the data flow, then toggle it to Active.
Related Other workflows
- Build your first AI agent$9.99
- Automate Daily Lead Outreach from Google Sheets via Outlook$9.99
- Automate Lead Generation by Scraping Emails from Google Maps to Google Sheets$14.99
- Real-time crypto market analysis with GPT-4o-mini and CoinMarketCap APIs via Telegram$2.99
- Automate Comprehensive Jira Project Setup with AI and Form Integration$24.99
- Book and manage appointments with Google Calendar and Gmail$14.99
More from Andre Vermeulen
Need this deployed? We'll set it up for you.
Our automation experts deploy this workflow in your stack, connect your accounts, and verify it works — or build a custom solution from scratch.