PulpMiner Web Scraper - Convert Any Webpage to Realtime JSON API
Convert any webpage into structured JSON data using AI. Scrape websites, extract data into custom JSON schemas, and call saved APIs programmatically. Useful for web scraping, data …
melvin2016
@melvin2016
What This Skill Does
AI-powered web scraping tool that converts any webpage into structured JSON data. Configure a URL and optional JSON template in the dashboard, then call the saved API endpoint to get clean, structured output. Supports dynamic URLs with variables, JavaScript rendering, CSS selectors, and caching.
Replaces manual web scraping scripts and traditional scraping tools by using AI to intelligently extract structured data from any webpage without writing parsing logic.
When to Use It
- Extract product prices and details from e-commerce pages for price tracking
- Scrape search results or listings from websites that lack an official API
- Monitor competitor websites for changes in content or pricing
- Build a data pipeline by scraping multiple pages and feeding structured JSON into a database
- Generate leads by extracting contact information from business directories
- Pull real-time data from dynamic JavaScript-heavy pages that traditional scrapers can't handle
Install
$ openclaw skills install @melvin2016/webscraper-pulpminerPulpMiner — AI Web Scraping & JSON API
PulpMiner converts any webpage into structured JSON using AI. You provide a URL and optionally a JSON template, and PulpMiner scrapes the page, runs it through an LLM, and returns clean structured data.
Authentication
All API calls require the apikey header:
apikey: <PULPMINER_API_KEY>
Get your API key from https://pulpminer.com/api — click "Regenerate Key" if you don't have one.
Core Workflow
PulpMiner works in two phases:
- Create a saved API — Configure a URL, scraper, LLM, and optional JSON template via the PulpMiner dashboard at https://pulpminer.com/api
- Call the saved API — Use the external endpoint with your API key to fetch structured JSON
Calling a Saved API
Static API (fixed URL)
curl -X GET "https://api.pulpminer.com/external/<apiId>" \
-H "apikey: <PULPMINER_API_KEY>"
Returns JSON extracted from the configured webpage.
Dynamic API (URL with variables)
For APIs saved with template URLs like https://example.com/search?q={{query}}&page={{page}}:
curl -X POST "https://api.pulpminer.com/external/<apiId>" \
-H "apikey: <PULPMINER_API_KEY>" \
-H "Content-Type: application/json" \
-d '{"query": "javascript frameworks", "page": "1"}'
The {{variable}} placeholders in the saved URL get replaced with the values you provide.
Response Format
Successful responses return:
{
"data": { ... },
"errors": null
}
Error responses return:
{
"data": null,
"errors": "Error message describing what went wrong"
}
Caching
- API responses are cached for 24 hours by default
- If cache is older than 15 minutes, PulpMiner serves the cached version while refreshing in the background
- Cache can be disabled per-API in the dashboard settings
Configuration Options (Set in Dashboard)
When creating a saved API at https://pulpminer.com/api, you can configure:
| Option | Description |
|---|---|
| URL | The webpage to scrape |
| JSON Template | Optional JSON structure for the LLM to follow (e.g., {"name": "", "price": ""}) |
| Render JS | Enable for SPAs and JS-heavy pages (uses headless browser) |
| CSS Selector | Extract only a specific part of the page (e.g., .product-list, #main-content) |
| Extra Instructions | Additional guidance for the AI (e.g., "Only extract items with prices above $50") |
| Dynamic URL | Enable template variables in the URL with {{variable}} syntax |
| Cache | Toggle response caching on/off |
Integration with Zapier
For async scraping in Zapier workflows:
# Static API
curl -X POST "https://api.pulpminer.com/external/zapier/get/<apiId>" \
-H "apikey: <PULPMINER_API_KEY>" \
-d '{"callbackURL": "https://hooks.zapier.com/..."}'
# Dynamic API
curl -X POST "https://api.pulpminer.com/external/zapier/post/<apiId>" \
-H "apikey: <PULPMINER_API_KEY>" \
-d '{"callbackURL": "https://hooks.zapier.com/...", "query": "value"}'
Returns 201 immediately. Sends scraped data to the callback URL when complete.
Integration with n8n
Verify authentication:
curl -X GET "https://api.pulpminer.com/external/n8n/auth" \
-H "apikey: <PULPMINER_API_KEY>"
Then use the standard /external/<apiId> endpoints for data fetching.
Credits
- Each API call costs 0.25–0.4 credits depending on the endpoint
- JavaScript rendering adds 0.1 credits extra
- New users get 5 free credits
- Purchase more at https://pulpminer.com/credits
Tips
- Use CSS selectors to narrow down the scraped content and improve accuracy
- Provide a JSON template for consistent, predictable output structures
- Enable JS rendering only when needed — static pages scrape faster and cost fewer credits
- Use extra instructions to guide the AI (e.g., "Return dates in ISO 8601 format")
- For monitoring use cases, keep caching enabled to reduce credit usage
- Use the playground first to verify a URL is scrapable before saving an API config
- Dynamic APIs are ideal for search pages, paginated content, and parameterized URLs
Links
- Website: https://pulpminer.com
- API Dashboard: https://pulpminer.com/api
Top skills in this category
Computer Use
@ram-raghav-sFull desktop computer use for headless Linux servers. Xvfb + XFCE virtual desktop with xdotool automation. 17 actions (click, type, scroll, screenshot, drag,...
Web Content Fetcher
@mrtommywu网页内容获取工具 | 当常规爬虫被过滤时,使用替代服务获取网页内容。支持:1) r.jina.ai - 最稳定 2) markdown.new - Cloudflare 专用 3) defuddle.md - 备用方案。触发词:获取网页内容、网页转markdown、内容抓取、fetch webpage、bypas...
Local Places
@steipeteSearch for places (restaurants, cafes, etc.) via Google Places API proxy on localhost.
SEO Intelligence & Competitor Analysis Pro
@qqyulePerform deep SEO competitor analysis, including keyword research, backlink checking, and content strategy mapping. Use when the user wants to analyze a website's competitors or improve their own SEO ranking by studying the competition.
System Resource Monitor
@passersssA clean, reliable system resource monitor for CPU load, RAM, Swap, and Disk usage. Optimized for OpenClaw.