finding-data-scientists-on-twitter
Finds data scientists and ML engineers to recruit using apidojo's Twitter scrapers on Apify. Triggers when the user asks to: find data scientists on Twitter for recruiting, discove…
API Dojo
@apidojo-io
Install
$ openclaw skills install @apidojo-io/finding-data-scientists-on-twitterFinding Data Scientists And Ml Engineers on Twitter
Discovers data scientists and ML engineers on Twitter via skill keywords, portfolio/project signals, and open-to-work indicators. Twitter surfaces professionals who actively discuss their craft — a strong passive candidate signal.
Prerequisites
APIFY_TOKENenvironment variable set- Optional: Apify MCP server installed
Inputs
| Parameter | Type | Required | Default | Notes |
|---|---|---|---|---|
startUrls | array | Optional | [] | Twitter profile or tweet URLs |
twitterHandles | array | Optional | [] | Twitter usernames (without @) |
twitterUserIds | array | Optional | [] | Twitter user IDs |
getFollowers | boolean | Optional | false | Extract follower lists |
getFollowing | boolean | Optional | false | Extract following lists |
getRetweeters | boolean | Optional | false | Extract retweeters of a tweet URL |
includeUnavailableUsers | boolean | Optional | false | Include unavailable/suspended users |
maxItems | number | Optional | Unlimited | Maximum users to return |
customMapFunction | string | Optional | — | JavaScript function to transform each output object |
Workflow
Progress:
- [ ] Step 1: Search for role-specific tweets
- [ ] Step 2: Collect unique handles
- [ ] Step 3: Enrich profiles
- [ ] Step 4: Score candidate fit
- [ ] Step 5: Deliver candidate list
Step 1: Search Queries
Recommended — run_actor.js (handles waiting, output, and file saving automatically):
# Quick answer (prints table to chat)
node scripts/run_actor.js \
--actor "apidojo~twitter-user-scraper" \
--input '{"param": "value"}'
# Save as CSV
node scripts/run_actor.js \
--actor "apidojo~twitter-user-scraper" \
--input '{"param": "value"}' \
--output YYYY-MM-DD_results.csv --format csv
# Save as JSON
node scripts/run_actor.js \
--actor "apidojo~twitter-user-scraper" \
--input '{"param": "value"}' \
--output YYYY-MM-DD_results.json --format json
APIFY_TOKENmust be set in environment or.envfile.
If Apify MCP is available:
Tool: apify:run-actor
Actor: "apidojo~tweet-scraper"
Input:
{
"searchTerms": ["data scientist", "ML engineer", "LLM engineer", "machine learning open to work"],
"maxItems": 300,
"tweetLanguage": "en"
}
REST API fallback:
curl -X POST \
"https://api.apify.com/v2/acts/apidojo~tweet-scraper/runs?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"searchTerms": ["data scientist", "ML engineer", "LLM engineer", "machine learning open to work"], "maxItems": 300}'
Collect unique author.username from results.
Step 2: Enrich Profiles
If Apify MCP is available:
Tool: apify:run-actor
Actor: "apidojo~twitter-user-scraper"
Input: {"usernames": ["[username1]", "[username2]", "..."]}
REST API fallback:
curl -X POST \
"https://api.apify.com/v2/acts/apidojo~twitter-user-scraper/runs?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"usernames": ["handle1", "handle2"]}'
Step 3: Filter and Score
Skill confirmation: bio contains keywords: "data science", "machine learning", "ML", "NLP", "LLM", "AI", "neural network", "PyTorch", "TensorFlow", "scikit-learn"
research_signal = bio contains 'PhD', 'researcher', or links to papers/Google Scholar
Candidate score:
candidate_score = (skill_confirmed ? 1 : 0) * 0.35
+ (open_to_work_signal ? 1 : 0) * 0.30
+ (followerCount in 200..20000 ? 1 : 0.6) * 0.20
+ (tweeted_in_last_30_days ? 1 : 0) * 0.15
Activity: active (< 30 days) | passive (30–90 days) | dormant (> 90 days)
Step 4: Edge Cases
- Company/brand accounts in results: Filter where
followerCount > 50KAND bio contains no personal pronouns; these are likely brand accounts - < 20 candidates found: Broaden skill term; remove location or seniority filter; try adjacent skills
- Bot detection: Flag
followerCount / followingCount < 0.05ANDtweetsCount < 20as potential bot - Location not matching: Bio location is free text — use fuzzy match; accept partial city/country names
Output Format
# Data Scientists And Ml Engineers Candidates: [ML_SPECIALTY]
Profiles found: [N] | Open-to-work: [N] | Active: [N] | Date: [DATE]
## Priority: Open-to-Work Candidates
| Name | @Handle | Specialty | Location | Followers | Last Active | Score |
|------|---------|----------|---------|-----------|------------|-------|
## Passive Candidates
| Name | @Handle | Specialty | Location | Followers | Score |
|------|---------|----------|---------|-----------|-------|
## Bio Highlights (Top 5)
1. @[handle]: "[bio excerpt]"
Troubleshooting
All results are agencies/companies not individuals: Add personal pronouns filter or search "I am a [role]", "I do [skill]".
Role too generic returns too many results: Add location OR seniority qualifier.
No open-to-work signals: Most candidates don't signal publicly — treat passive candidates as warm leads with personalized outreach referencing their recent content.
Top skills in this category
Playwright (Automation + MCP + Scraper)
@ivangdavilaAutomates, tests, and debugs browsers with Playwright: locators, auto-waiting, traces, CI runs, and MCP browser control. Use when a test is flaky, times out, or fails only in CI or headless; when a locator matches multiple elements or the wrong one (strict mode violation); when clicks need force, waits become sleeps, or networkidle never settles; for storageState and login setup, request mocking and HAR replay, uploads and downloads, iframes and shadow DOM, popups and dialogs, screenshot diffs that change per machine, trace and report artifacts, sharding a slow suite, device and permission emulation, accessibility checks, driving a real browser through Playwright MCP, extracting data from JS-rendered pages, or porting a Cypress, Puppeteer, or Selenium suite to Playwright. Not for maintaining an existing Cypress or Puppeteer suite (cypress, puppeteer) or for work a plain HTTP request answers (http).
Openclaw Command Center
@jontsaiMission control dashboard for OpenClaw - real-time session monitoring, LLM usage tracking, cost intelligence, and system vitals. View all your AI agents in o...
Computer Use
@ram-raghav-sFull desktop computer use for headless Linux servers. Xvfb + XFCE virtual desktop with xdotool automation. 17 actions (click, type, scroll, screenshot, drag,...
Interview Simulator
@wscatsSimulates mock interviews for any role and experience level with tailored technical, behavioral, and case questions plus detailed feedback and scoring.
Scrapling Official Skill
@d4vinciScrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering. Use when asked to scrape, crawl, or extract data from websites; web_fetch fails; the site has anti-bot protections; write Pyth