Openclaw Memory Toolkit
Hybrid memory pipeline for OpenClaw agents — extraction, archiving, temporal decay scoring, consolidation, and hybrid search (FTS5 + sqlite-vec + RRF). Six standalone Python script…
MisterMiJarvis
@mistermijarvis
Install
$ openclaw skills install @mistermijarvis/memory-toolkitMemory Pipeline Skill
Complete memory management pipeline for OpenClaw agents: extraction, archiving, scoring, consolidation, health monitoring, and ontology — all local-first, zero external API dependencies.
Pipeline Overview
Nightly Cron (23h)
│
├─ 1. trace_extractor.py # Extract decisions/errors/facts from sessions
├─ 2. auto_archive.py # Archive daily notes >21 days
├─ 3. scoring.py # Score all memories with temporal decay
├─ 4. consolidate_advisor.py # Suggest consolidations (agent reviews)
├─ 5. memory_health.py # Periodic health check (weekly)
└─ 6. hybrid_search.py # Hybrid search: FTS5 + sqlite-vec + RRF
All scripts are standalone and composable. Run individually or as a pipeline.
Scripts
1. trace_extractor.py — Session extraction
Extracts decisions, errors, facts, and patterns from OpenClaw session transcripts and daily notes. Updates daily notes with extracted items, appends entities to the ontology graph.
# Nightly (pattern-based, fast ~5s)
python3 scripts/trace_extractor.py --days 1
# Deep extraction (LLM-powered, ~60-180s)
python3 scripts/trace_extractor.py --days 3 --llm
# With session transcripts
python3 scripts/trace_extractor.py --days 1 --llm --sessions
# Preview only
python3 scripts/trace_extractor.py --days 1 --llm --dry-run
Categories extracted:
- 🟢 DECISIONS — new choices, config changes, migrations
- 🔴 ERRORS — bugs, failures, workarounds
- 🔵 FACTS — new versions, configs, status changes
- ⬆️ PROMOTIONS — items worth promoting to MEMORY.md
Output: Daily notes updated, ontology entities added, .trace-extracted flag.
2. auto_archive.py — Daily note archiving
Moves daily notes older than N days to memory/archive/YYYY-MM/ subdirectories.
python3 scripts/auto_archive.py # Archive notes > 21 days
python3 scripts/auto_archive.py --days 30 # Custom threshold
python3 scripts/auto_archive.py --dry-run # Preview only
python3 scripts/auto_archive.py --verbose # Show each file
Idempotent. Only moves YYYY-MM-DD*.md files. Zero dependencies.
3. scoring.py — Temporal decay scoring
Scores all memory items using exponential recency decay, category weights, frequency boost, entity boost, and completion penalty.
python3 scripts/scoring.py # Score all memories
python3 scripts/scoring.py --verbose # Show top 20
python3 scripts/scoring.py --threshold 0.3 # Filter by min score
python3 scripts/scoring.py --dry-run # Don't write output
Scoring formula:
score = weight_category × recency_decay × frequency_boost × entity_boost × completion_penalty
recency_decay = exp(-ln(2) × days_old / HALF_LIFE_DAYS)
Category weights: DECISIONS ×3, ERRORS ×2, FACTS ×1.5, PATTERNS ×1.2, TRANSIENT ×1
Output: memory/scores.json — full ranking with stats and promotion candidates.
4. consolidate_advisor.py — Consolidation suggestions
Analyzes recent daily notes + scores.json to identify clusters, promotions, stale items, and duplicates. Does NOT modify files unless explicitly asked.
python3 scripts/consolidate_advisor.py # Last 7 days
python3 scripts/consolidate_advisor.py --days 14 # Custom window
python3 scripts/consolidate_advisor.py --verbose # All suggestions
python3 scripts/consolidate_advisor.py --no-llm # Skip LLM (fallback)
python3 scripts/consolidate_advisor.py --apply-promotions # Write to MEMORY.md
Output: memory/consolidation_report.json — clusters, promotions, stale items, duplicates.
LLM optional (Ollama) for cluster summaries. Falls back to text-based with --no-llm.
5. memory_health.py — System health check
Comprehensive diagnostics: trace extraction, LoCoMo benchmark, MEMORY.md size, ontology health, daily notes hygiene, index status, drift detection.
python3 scripts/memory_health.py # Full health check
python3 scripts/memory_health.py --quick # Skip benchmark & LLM
python3 scripts/memory_health.py --benchmark # Benchmark only
python3 scripts/memory_health.py --deep # LLM + sessions + benchmark
python3 scripts/memory_health.py --fix # Fix mode (archive, clean)
Output: results/YYYY-MM-DD.json — full diagnostic report.
6. hybrid_search.py — Hybrid search (FTS5 + sqlite-vec + RRF)
Hybrid memory search combining lexical (BM25 via SQLite FTS5) and semantic (vector via sqlite-vec) retrieval using Reciprocal Rank Fusion (RRF, k=60).
# Initialize DB with schema
python3 scripts/hybrid_search.py init
# Index all memory files
python3 scripts/hybrid_search.py index
# Search
python3 scripts/hybrid_search.py query "AstroCapture"
python3 scripts/hybrid_search.py query "roadmap EIIDP" --top 10
# Lexical only (BM25)
python3 scripts/hybrid_search.py query "2026-08-17" --lexical-only
# Vector only (semantic)
python3 scripts/hybrid_search.py query "memory decay scoring" --vector-only
# JSON output for programmatic use
python3 scripts/hybrid_search.py query "leadership coaching" --json
# Stats
python3 scripts/hybrid_search.py stats
# Index a single file
python3 scripts/hybrid_search.py add path/to/file.md --category skill
How it works:
query → ┬─ vector_search (nomic-embed-text, top 20) ──┐
└─ lexical_search (FTS5/BM25, top 20) ────────┤
↓
RRF(k=60) fusion
↓
min_score filter (≥0.015)
↓
temporal boost (optional)
↓
source deduplication
↓
top K results
RRF ignores raw scores and uses only ranks: rrf(d) = Σ 1/(k + rank_m(d)).
Source deduplication groups by file, returning the best chunk per source.
Gemini vigilance #1 — min_rrf_score (noise threshold):
Chunks appearing in neither top-20 list have RRF score ~0 = pure noise.
Filtered by default at 0.015. Override with --min-score 0 to disable.
Gemini vigilance #3 — temporal_boost (decay weighting):
RRF score is multiplied by (1 + 0.1 * normalized_score) where normalized_score
comes from the score column (populated by scoring.py temporal decay).
Gives slight priority to recent facts when context conflicts.
Disable with --no-temporal-boost.
Schema: Single SQLite file with three synchronized tables:
memories— content, category, layer, source, score, timestampsmemories_fts— FTS5 virtual table (external content, auto-synced via triggers)memories_vec— vec0 virtual table (float[768], nomic-embed-text)
Layers: episodic (daily notes), semantic (long-term facts, ontology), procedural (skills, config)
Requirements: sqlite-vec (pip install in venv), Ollama with nomic-embed-text
Output: hybrid-search/agent_memory.db — SQLite DB with FTS5 + vec0 indexes.
Ontology
The ontology graph stores entities and relations as JSONL. A YAML schema defines allowed types and relations.
Entity types: Person, Organization, Project, Task, Document, Event, Skill, Device, Service, Tool, Infrastructure, Concept, Location, Pet, BugFix, SecurityEvent, Integration, Feature, Software, Configuration
Relation types: reports_to, has_owner, includes, depends_on, manages, uses, integrated_with, located_at, fixes, monitors
Files:
memory/ontology/graph.jsonl— entity and relation recordsmemory/ontology/schema.yaml— type and relation definitionsmemory/ontology/graph-index.json— search index
Configuration
Environment variables with defaults:
| Variable | Default | Description |
|---|---|---|
WORKSPACE | ~/.openclaw/workspace | OpenClaw workspace path |
OLLAMA_URL | http://localhost:11434 | Ollama API URL |
OLLAMA_MODEL | glm-5.2 | Model for LLM extraction/summaries |
Scoring constants (top of scoring.py):
| Parameter | Default | Description |
|---|---|---|
HALF_LIFE_DAYS | 14 | Recency decay half-life |
MAX_SCORE | 5.0 | Score cap |
PROMOTE_THRESHOLD | 2.0 | Min score for promotion |
ARCHIVE_THRESHOLD | 0.05 | Score below = archive candidate |
Memory health thresholds (top of memory_health.py):
| Parameter | Default | Description |
|---|---|---|
MEMORY_MAX_SIZE | 5000 | MEMORY.md max size in bytes |
DAILY_NOTES_MAX_AGE | 14 | Days before archiving |
Requirements
- Python 3.10+
- Ollama (optional — LLM extraction and cluster summaries)
- No pip packages required — pure stdlib
Nightly Cron Integration
Recommended nightly pipeline (after trace extraction):
# In nightly cron (23h):
python3 scripts/trace_extractor.py --days 1
python3 scripts/auto_archive.py
python3 scripts/scoring.py
python3 scripts/consolidate_advisor.py --no-llm # quiet mode
Weekly health check (Monday, separate cron):
python3 scripts/memory_health.py --quick
Monthly deep check (manual):
python3 scripts/memory_health.py --deep
Design Principles
- Local-first — no external API, no paid dependencies
- Composable — each script is standalone, can run independently
- Safe by default — dry-run is default, writes require explicit flags
- Human-in-the-loop — consolidation suggestions, not auto-merge
- Pipeline-friendly — scripts chain naturally, outputs feed inputs
Design Principles
- Local-first — no external API, no paid dependencies
- Composable — each script is standalone, can run independently
- Safe by default — dry-run is default, writes require explicit flags
- Human-in-the-loop — consolidation suggestions, not auto-merge
- Pipeline-friendly — scripts chain naturally, outputs feed inputs
- Zero dependencies — pure Python stdlib (except sqlite-vec for hybrid search)
- Strata-aware — episodic, semantic, and procedural memory are separated
License
MIT — free to use, modify, and share.
Top skills in this category
Humanizer
@biostartechnologyRemove signs of AI-generated writing from text. Use when editing or reviewing text to make it sound more natural and human-written. Based on Wikipedia's comprehensive "Signs of AI writing" guide. Detects and fixes patterns including: inflated symbolism, promotional language, superficial -ing analyses, vague attributions, em dash overuse, rule of three, AI vocabulary words, negative parallelisms, and excessive conjunctive phrases.
Elite Longterm Memory
@nextfrontierbuildsUltimate AI agent memory system for Cursor, Claude, ChatGPT & Copilot. WAL protocol + vector search + git-notes + cloud backup. Never lose context again. Vibe-coding ready.
Model Usage
@steipeteUse CodexBar CLI local cost usage to summarize per-model usage for Codex or Claude, including the current (most recent) model or a full model breakdown. Trigger when asked for model-level usage/cost data from codexbar, or when you need a scriptable per-model summary from codexbar cost JSON.
Screenshot
@ivangdavilaCapture, inspect, and compare screenshots of screens, windows, regions, web pages, simulators, and CI runs with the right tool, wait strategy, viewport, and...
Interview Simulator
@wscatsSimulates mock interviews for any role and experience level with tailored technical, behavioral, and case questions plus detailed feedback and scoring.