Jina Reader

Web content extraction via Jina AI Reader API. Three modes: read (URL to markdown), search (web search + full content), ground (fact-checking). Extracts clea...

Eric Santos

@ericsantos

What This Skill Does

Extracts clean web content via Jina AI Reader API in three modes: read (URL to markdown), search (web search with full content), and ground (fact-checking). Supports CSS selectors, geo-proxies, and multiple output formats.

Replaces manual web scraping and server-side IP exposure by routing requests through Jina's infrastructure for clean, JavaScript-rendered content extraction.

When to Use It

  • Convert a blog post or article URL to clean markdown for offline reading
  • Search the web for recent topics and retrieve full page content for each result
  • Fact-check a specific statement by grounding it against authoritative web sources
  • Extract a specific section of a webpage using a CSS selector
  • Remove navigation, ads, or footer elements before extracting page content
  • Fetch content from a geo-restricted site using a country-specific proxy

Install

$ openclaw skills install @ericsantos/jina-reader

Jina Reader

Extract clean web content via Jina AI — without exposing your server IP.

Read a URL

{baseDir}/scripts/reader.sh "https://example.com/article"

Search the web (top 5 results with full content)

{baseDir}/scripts/reader.sh --mode search "latest AI news 2025"

Fact-check a statement

{baseDir}/scripts/reader.sh --mode ground "OpenAI was founded in 2015"

Options

FlagDescriptionDefault
--moderead, search, groundread
--selectorCSS selector to extract specific region
--waitCSS selector to wait for before extraction
--removeCSS selectors to remove (comma-separated)
--proxyCountry code for geo-proxy (br, us, etc.)
--nocacheForce fresh content (skip cache)off
--formatmarkdown, html, text, screenshotmarkdown
--jsonRaw JSON outputoff

Examples

# Extract article content
{baseDir}/scripts/reader.sh "https://blog.example.com/post"

# Extract specific section via CSS selector
{baseDir}/scripts/reader.sh --selector "article.main" "https://example.com"

# Remove nav and ads before extraction
{baseDir}/scripts/reader.sh --remove "nav,footer,.ads" "https://example.com"

# Search with JSON output
{baseDir}/scripts/reader.sh --mode search --json "AI enterprise trends"

# Read via Brazil proxy
{baseDir}/scripts/reader.sh --proxy br "https://example.com.br"

# Fact-check a claim
{baseDir}/scripts/reader.sh --mode ground "Tesla is the most valuable car company"

API Key

Resolution order:

  1. $JINA_API_KEY env var
  2. ~/.config/jina/api_key file (canonical local path)

Either is fine. Pick one:

# Option A: env var
export JINA_API_KEY="jina_..."

# Option B: config file (preferred for local installs)
mkdir -p ~/.config/jina && echo "jina_..." > ~/.config/jina/api_key
chmod 600 ~/.config/jina/api_key

Free tier: 10M tokens (no signup needed). Get key at https://jina.ai/reader/

Pricing

  • Read: ~$0.005/page (standard) | 3x for ReaderLM-v2
  • Search: 10K tokens fixed + variable per result
  • Ground: ~300K tokens/request (~30s latency)

Why Jina Reader?

  • IP protection — requests route through Jina's infra, not your server
  • Clean markdown — readability extraction + optional ReaderLM-v2
  • Dynamic content — headless Chrome renders JavaScript
  • Structured extraction — JSON schema support for data extraction

Top skills in this category