Data & Analytics Workflows
Data processing and analytics
Build Academic Citation Networks with PDF Vector API for Gephi Visualization
This workflow contains community nodes that are only compatible with the self-hosted version of n8n. ## Build Citation Networks from Research Papers Automatically build and visualize citation networks by fetching papers and their references. Discover influential works and research trends in any field. ### Workflow Features: - Start with seed papers (DOIs, PubMed IDs, etc.) - Fetch cited and citing papers recursively - Build network graph data - Export to visualization tools (Gephi, Cytoscape) - Identify key papers and research clusters ### Process Flow: 1. **Input**: Seed paper identifiers 2. **Fetch Papers**: Get paper details and references 3. **Expand Network**: Fetch cited papers (configurable depth) 4. **Build Graph**: Create nodes and edges 5. **Analyze**: Calculate metrics (centrality, clusters) 6. **Export**: Generate visualization-ready data ### Applications: - Research trend analysis - Finding seminal papers in a field - Grant proposal background research
n8n$9.99Daily Insight Email from Structured Web Data with Firecrawl
**Daily Web Scraper & AI Summary with Firecrawl + Email Automation** Need to extract and summarize web content from a site that doesn't have an API? This workflow runs daily to scrape a web page using Firecrawl, summarize the content with OpenAI, and send it directly to your email - fully automated. **Watch Full Video Step-by-step Tutorial Here:** [https://www.youtube.com/@Automatewithmarc](https://www.youtube.com/@Automatewithmarc) **How It Works** - **Daily Trigger** - Starts the workflow every 24 hours. - **Firecrawl Node** - Crawls and extracts structured data from any web page you specify. - **OpenAI Node (Optional)** - Processes and summarizes the raw content using a prompt you control. - **Gmail Node** - Sends the final summary or content snapshot to your email inbox. **Perfect For** - Business analysts tracking daily market or industry news - Researchers and founders automating competitive intelligence - Anyone who wants web data delivered without coding or scraping scripts **Setup Instructions** - **Firecrawl API Key** - Sign up and insert your key in the credentials. - **Update Target URL** - Edit the URL in the Firecrawl node to your desired site. - **Customize the Prompt** - Tailor the OpenAI prompt to extract the insights you want. - **Connect Gmail** - Add your Gmail credentials and set your recipient email. **Built With** - Firecrawl (Web scraping without code) - OpenAI (For summarizing and insight extraction) - Gmail (Automated notifications) - n8n (Workflow automation engine)
n8n$14.99Automate Business Data Enrichment for Company Domains with Perplexity AI and Google Sheets
This workflow automates the enrichment of company domain lists with detailed business information using Perplexity AI, storing results in Google Sheets for streamlined analysis.
n8n$9.99Automate LinkedIn Company Data Extraction to Airtable
Effortlessly extract and structure detailed company insights from LinkedIn using Airtop, ideal for investors, sales teams, and market researchers.
n8n$4.99Automate YouTube Video Statistics Collection to Google Sheets
Efficiently gather YouTube video statistics such as views, likes, and comments, and store them in Google Sheets for analysis and reporting.
n8n$9.99Extract, Summarize & Analyze Amazon Price Drops with Bright Data & Google Sheets
### Notice Community nodes can only be installed on self-hosted instances of n8n. ### Who this is for This n8n-powered automation uses Bright Data's MCP Client to extract real-time data from a price drop site listing Amazon products, including price changes and related product details. The extracted data is enriched with structured data transformation, content summarization, and sentiment analysis using Google Gemini LLM. The Amazon Price Drop Intelligence Engine is designed for: - **Ecommerce Analysts** who need timely updates on competitor pricing trends - **Brand Managers** seeking to understand consumer sentiment around pricing - **Data Scientists** building pricing models or enrichment pipelines - **Affiliate Marketers** looking to optimize campaigns based on dynamic pricing - **AI Developers** automating product intelligence pipelines ### What problem is this workflow solving? This workflow solves several key pain points: **Reliable Scraping**: Uses Bright Data MCP, a managed crawling platform that handles proxies, captchas, and site structure changes automatically. **Insight Generation**: Transforms unstructured HTML into structured data and then into human-readable summaries using Google Gemini LLM. **Sentiment Context**: Goes beyond raw pricing data to reveal how customers feel about the price change, helping businesses and researchers measure consumer reaction. **Automated Reporting**: Aggregates and stores data for easy access and downstream automation (e.g., dashboards, notifications, pricing models). ### What this workflow does **Scrape price drop site with Bright Data MCP** The workflow begins by scraping targeted price drop site for Amazon listings using Bright Data's Model Context Protocol (MCP). You can configure this to target: **Structured Data Extraction** Once the HTML content is retrieved, Google Gemini is employed to: - Parse and structure the product information (title, price, discount, brand, ratings) **Summarization & Sentiment Analysis** The extracted data is passed through an LLM chain to: - Generate a concise summary of the product and its recent price movement - Perform sentiment analysis on user reviews and public perception **Store the Results** - Save to disk for archiving or bulk processing - Updated in a Google Sheet, making it instantly shareable with your team or integrated into a BI dashboard ### Pre-conditions 1. Knowledge of Model Context Protocol (MCP) is highly essential. Please read this blog post - [model-context-protocol](https://www.anthropic.com/news/model-context-protocol) 2. You need to have the [Bright Data](https://brightdata.com/) account and do the necessary setup as mentioned in the **Setup** section below. 3. You need to have the Google Gemini API Key. Visit [Google AI Studio](https://aistudio.google.com/) 4. You need to install the Bright Data MCP Server [@brightdata/mcp](https://www.npmjs.com/package/@brightdata/mcp) 5. You need to install the [n8n-nodes-mcp](https://github.com/nerding-io/n8n-nodes-mcp) ### Setup 1. Please make sure to set up n8n locally with MCP Servers by navigating to [n8n-nodes-mcp](https://www.youtube.com/watch?v=NUb73ErUCsA) 2. Please make sure to install the Bright Data MCP Server [@brightdata/mcp](https://www.npmjs.com/package/@brightdata/mcp) on your local machine. 3. Sign up at [Bright Data](https://brightdata.com/). 4. Create a Web Unlocker proxy zone called mcp_unlocker on Bright Data control panel. 5. Navigate to Proxies & Scraping and create a new Web Unlocker zone by selecting Web Unlocker API under Scraping Solutions. 6. In n8n, configure the Google Gemini (PaLM) API account with the Google Gemini API key (or access through Vertex AI or proxy). 7. In n8n, configure the credentials to connect with MCP Client (SDIO) account with the Bright Data MCP Server as shown below.  Make sure to copy the Bright Data API_TOKEN within the Environments textbox above as API_TOKEN=<your-token> ### How to customize this workflow to your needs - **Target different platforms**: Switch Amazon for Walmart, eBay, or any e-commerce source using Bright Data's flexible scraping infrastructure. - **Enrich with more LLM tasks**: Add brand tone analysis, category classification, or competitive benchmarking using Gemini prompts. - **Visualize output**: Pipe the Google Sheet to Looker Studio, Tableau, or Power BI. - **Notification integrations**: Add Slack, Discord, or email notifications for price drop alerts.
n8n$14.99Hacker News Job Listing Scraper and Parser
This automated workflow scrapes and processes the monthly Who is Hiring thread from Hacker News, transforming raw job listings into structured data for analysis or integration with other systems. Perfect for job seekers, recruiters, or anyone looking to monitor tech job market trends. ## How it works - Automatically fetches the latest Who is Hiring thread from Hacker News - Extracts and cleans relevant job posting data using the HN API - Splits and processes individual job listings into structured format - Parses key information like location, role, requirements, and company details - Outputs clean, structured data ready for analysis or export ## Set up steps 1. Configure API access to [Hacker News](https://github.com/HackerNews/API) (no authentication required) 2. Follow the steps to get your cURL command from [https://hn.algolia.com/](https://hn.algolia.com/) 3. Set up desired output format (JSON structured data or custom format) 4. Optional: Configure additional parsing rules for specific job listing information 5. Optional: Set up integration with preferred storage or analysis tools The workflow transforms unstructured job listings into clean, structured data following this pattern: - Input: Raw HN thread comments - Process: Extract, clean, and parse text - Output: Structured job listing data This template saves hours of manual work collecting and organizing job listings, making it easier to track and analyze tech job opportunities from Hacker News's popular monthly hiring threads.
n8n$14.99Scrape Latest 20 TechCrunch Articles
# Retrieve 20 Latest TechCrunch Articles ## Who is this for? This workflow is designed for developers, content creators, and data analysts who need to scrape recent articles from TechCrunch. It's perfect for anyone looking to aggregate news articles or create custom feeds for analysis, reporting, or integration into other systems. ## What problem is this workflow solving? This workflow automates the process of scraping recent articles from TechCrunch. Manually collecting article data can be time-consuming and inefficient, but with this workflow, you can quickly gather up-to-date news articles with relevant metadata, saving time and effort. ## What this workflow does This workflow retrieves the latest 20 news articles from TechCrunch’s “Recent” page. It extracts the article URLs, metadata (such as titles and publication dates), and main content for each article, allowing you to access the information you need without any manual effort. ## Setup 1. Clone or download the workflow template. 2. Ensure you have a working n8n environment. 3. Configure the HTTP Request nodes with your desired parameters to connect to the TechCrunch API. 4. (Optional) Customize the workflow to target specific sections or topics of interest. 5. Run the workflow to retrieve the latest 20 articles. ## How to customize this workflow to your needs - Modify the HTTP request to pull articles from different pages or sections of TechCrunch. - Adjust the number of articles to retrieve by changing the selection criteria. - Add additional processing steps to further filter or analyze the article data. ## Workflow Steps 1. **Send an HTTP request** to the TechCrunch Recent page. 2. **Parse a posts box** that holds the list of articles. 3. **Parse all posts** to extract all articles. 4. **Split out posts** for each article. 5. **Extract the URL and metadata** from each article. 6. **Send an HTTP request** for each article using its URL. 7. **Locate and parse** the main content of each article. --- **Note:** Be sure to update the HTTP Request nodes with any necessary headers or authentication to work with TechCrunch's website.
n8n$9.99Send Google Analytics Data to AI to Analyze, Then Save Results in Baserow
## Who's this for? - If you own a website and need to analyze your Google Analytics data - If you need to create an SEO report on which pages are getting the most traffic or how your Google search terms are performing - If you want to grow your site based on suggestions from data   ## Use case Instead of hiring an SEO expert, I run this report weekly. It checks and compares the data from this week to the week before: - Views based on countries - The top performing pages - Google Search Console performance [Watch YouTube tutorial here](https://www.youtube.com/watch?v=KlWFhz9M9g) [Get my SEO A.I. agent system here](https://2828633406999.gumroad.com/l/rumjahn) ## How it works - The workflow gathers Google Analytics data for the past 7 days, then it gathers the data for the week before for comparison. - It does this 3 times to get: views per country, engagement per page, and Google Search Console results for organic search results. - The Google Analytics nodes have already chosen the correct dimensions and metrics. - At the end, it passes the data to openrouter.ai for A.I. analysis. - Finally, it saves to Baserow. ## How to use this - Input your Google Analytics credentials - Input your property ID - Input your Openrouter.ai credentials - Input your Baserow credentials - You will need to create a Baserow database with columns: Name, Country Views, Page Views, Search Report, Blog (name of your blog). Created by [Rumjahn](https://rumjahn.com/)
n8n$14.99Automate Web Data Search, Summarization, and Delivery via Webhooks
This workflow automates the process of searching web data using Perplexity, cleaning and summarizing the results with Gemini AI, and delivering structured insights to a webhook.
n8n$14.99Automate Multi-Search Engine Queries and Extract Structured Data
Simulate human-like searches across Google, Bing, and Yandex using Bright Data MCP and extract structured insights with Google Gemini.
n8n$14.99Efficient Image Retrieval with BrightData Web Unblocker Fallback
This workflow ensures reliable image retrieval from any web source by using a cost-effective strategy. It first attempts to fetch images using a standard HTTP request and switches to BrightData Web Unblocker for challenging cases.
n8n$4.99Automate Pinterest Content Scraping with AI and BrightData
This n8n workflow automates the process of scraping Pinterest content based on user-defined keywords using BrightData's API and the Claude Sonnet 4 AI model. It efficiently manages the scraping process, monitors progress, and organizes the extracted data into Google Sheets.
n8n$14.99Automate Upwork Job Scraping and Daily Email Reporting
This n8n workflow automates the process of scraping Upwork job listings using Apify, storing data in Google Sheets, and sending daily email reports. It ensures data quality through cleaning and deduplication, providing a streamlined solution for job tracking and market analysis.
n8n$14.99Collect Company Social Media Profiles with Extruct AI to Google Sheets
**Who's it for:** Sales teams, marketers, and analysts who need to quickly access all the social media and public profile links for any company. **How it works / What it does:** When you enter a company into the form, this workflow automatically searches for and collects all available links to the company's social media accounts, review sites, and public profiles from sources like Crunchbase and Zoominfo. All discovered URLs are added directly to your Google Sheet. **How to set up:** 1. Create an Extruct account at [www.extruct.ai/](https://www.extruct.ai/). 2. Open the Extruct table template, find the table ID in your browser's address bar, and copy it. 3. Make a copy of the provided Google Sheets template to your own Google Drive. 4. In n8n, paste the table ID into the variables node of your flow. 5. Set up Bearer authentication in every HTTP Request node using your Extruct API token (found on the API page in Extruct). 6. In the Google Sheets node, paste the link to your copied template and connect your Google account. 7. Run the flow once to load the fields, then map the output fields to the correct columns in your sheet. 8. Activate the flow and start adding companies via the form. **Requirements:** - Extruct account and API token - Extruct table template - Google account with Google Sheets **How to customize the workflow:** You can add your own columns to the Extruct table and your Google Sheet. Just add the new column in both places and map it in the Google Sheets node in n8n.
n8n$9.99Pull Square Sales Summary Reports for Automated Reporting and Analysis
## Programmatically Pull Square Report Data Into N8N ## What It Does This sub-workflow connects to the Square API and generates a daily sales summary report for all of your Square locations. The report matches the figures displayed in the Square Dashboard > Reports > Sales Summary. It's designed to be reused in other workflows, ideal for reporting, data storage, accounting, or automation. ## Prerequisites To use this workflow, you'll need: - Square API credentials (configured as a Header Auth credential) ## How to Set Up Square Credentials: - Go to Credentials > Create New - Choose Header Auth - Set the Name to Authorization - Set the Value to your Square Access Token (e.g., Bearer <your-api-key>) ## How It Works 1. Trigger: The workflow is triggered as a sub-workflow, requiring a report_date input. 2. Fetch Locations: An HTTP request gets all Square locations linked to your account. 3. Fetch Orders: For each location, an HTTP request pulls completed orders for the specified report_date. 4. Filter Empty Locations: Locations with no sales are ignored. 5. Aggregate Sales Data: A Code node processes the order data and produces a summary identical to Square's built-in Sales Summary report. 6. Output: A cleaned, consistent summary that can be consumed by parent workflows or other nodes. ## Example Use Cases - Automatically store daily sales data in Google Sheets, MySQL, or PostgreSQL for analysis and historical tracking - Automatically send daily email or Slack reports to managers or finance teams - Build weekly/monthly reports by looping over multiple dates - Push sales data into accounting software like QuickBooks or Xero for automated bookkeeping - Calculate commissions or rent payments based on sales volume ## How to Use - Configure both HTTP Request nodes to use your Square API credential. - If you are not in the Toronto/New York timezone, please change the start_at and end_at parameters in the second HTTP node from -05:00 to your local timezone - Use as a sub-workflow inside a main workflow. - Pass a report_date (formatted as YYYY-MM-DD) to the sub-workflow when you call it. ## Customization Options - Add pagination to handle locations with more than 1,000 orders per day. - Expand the workflow to save or send the report output via additional integrations (email, database, webhook, etc.). ## Why It's Useful This workflow saves time, reduces manual report pulling from Square, and enables smarter automation around sales data—whether for operations, finance, or performance monitoring.
n8n$9.99Automatic Scraping of Company Information Before Call
Instantly collect and organize detailed company data, allowing sales reps to focus on selling and providing every prospect with personal, in-depth attention. It helps to flag pain points and highlight relevant trends, empowering your team to deliver perfectly customized value propositions every time.
Activepieces$2.99Automate Google Reviews Analysis and Summarization with AI
This n8n workflow automates the process of analyzing Google Maps reviews for restaurants, using AI to summarize insights and identify optimization opportunities. It integrates with Google Sheets, SerpAPI, and OpenAI to streamline data collection, analysis, and reporting.
n8n$9.99Sequential Google Sheets Data Processing with Execution Control
Ensure reliable and sequential data processing from Google Sheets by preventing simultaneous workflow executions in n8n.
n8n$9.99Automate SEO Insights from Matomo Analytics with AI and Store in Baserow
This workflow automates the process of analyzing Matomo analytics data using AI and stores the insights in Baserow. It helps website owners improve visitor engagement by providing actionable SEO recommendations.
n8n$9.99Automate TrustPilot Review Analysis with Bright Data and OpenAI
Streamline the extraction, summarization, and analysis of TrustPilot reviews using Bright Data and OpenAI. This workflow automates data collection and insights generation, enhancing decision-making for product managers and marketing teams.
n8n$14.99Automate SQL Query Generation and Visualization from Natural Language
This workflow allows users to convert natural language inputs into SQL queries, execute them on a PostgreSQL database, and visualize the results using QuickChart, streamlining data analysis without manual query writing.
n8n$24.99Automate Startup Funding Data Extraction to Excel
This workflow automates the discovery and extraction of seed-funded startup data from RSS feeds, processes it with AI, and exports it to an Excel sheet on OneDrive, streamlining lead generation for sales teams.
n8n$9.99Automate Google Places Data Collection and Storage in Google Sheets
Streamline your lead generation by automatically scraping Google Places data using Dumpling AI and storing it in Google Sheets. Ideal for marketers and SEO professionals seeking efficient data collection.
n8n$4.99
Related categories
Custom AI Systems & Services
Our team of experienced AI builders will help build custom AI systems, workflows, and solutions.
Request Custom Work