Spike Skill for Hermes Agent: Validate Ideas Before Building
Throwaway experiments to validate an idea before build.
Written by Neura Market from the official Hermes Agent documentation for Spike. Commands, paths, and version numbers are reproduced from the source unchanged.
Read the official documentationThe Spike skill is a structured way to run throwaway experiments that answer a specific feasibility question. You reach for it when you need to know whether something works before you commit to a full implementation. It is bundled with Hermes Agent and lives at skills/software-development/spike.
What it does
Spike turns a vague "I wonder if X works" into a repeatable loop: decompose the idea into testable questions, research the options, build the smallest thing that answers the question, and write a verdict. The output is disposable code and a clear yes/no/maybe that tells you what to do next.
Before you start
- Hermes Agent must be running with the Spike skill loaded. It is bundled by default, so no separate install is needed.
- The skill works on Linux, macOS, and Windows.
- You need a project directory where the
spikes/folder will be created. The skill assumes you are already in a repo or working directory. - For research steps, Hermes tools like
web_search,web_extract, andterminalmust be available. If you have Context7 MCP configured, the skill can usemcp_*_resolve-library-idandmcp_*_query-docsas well. - If the full GSD system is installed (via
npx get-shit-done-cc --hermes), the skill will prefergsd-spikefor persistent state and MANIFEST tracking. This standalone version is for users who do not have or do not want the full system.
Core method
Every spike follows this loop, regardless of scale:
decompose → research → build → verdict
↑__________________________________________↓
iterate on findings
1. Decompose
Break the user's idea into 2-5 independent feasibility questions. Each question is one spike. Present them as a table with Given/When/Then framing:
| # | Spike | Validates (Given/When/Then) | Risk |
|---|---|---|---|
| 001 | websocket-streaming | Given a WS connection, when LLM streams tokens, then client receives chunks < 100ms | High |
| 002a | pdf-parse-pdfjs | Given a multi-page PDF, when parsed with pdfjs, then structured text is extractable | Medium |
| 002b | pdf-parse-camelot | Given a multi-page PDF, when parsed with camelot, then structured text is extractable | Medium |
Spike types:
- standard, one approach answering one question
- comparison, same question, different approaches (shared number, letter suffix
a/b/c)
Good spike questions: specific feasibility with observable output. Bad spike questions: too broad, no observable output, or just "read the docs about X".
Order by risk. The spike most likely to kill the idea runs first. No point prototyping the easy parts if the hard part doesn't work.
Skip decomposition only if the user already knows exactly what they want to spike and says so. Then take their idea as a single spike.
2. Align (for multi-spike ideas)
Present the spike table. Ask: "Build all in this order, or adjust?" Let the user drop, reorder, or re-frame before you write any code.
3. Research (per spike, before building)
Spikes are not research-free, you research enough to pick the right approach, then you build. Per spike:
- Brief it. 2-3 sentences: what this spike is, why it matters, key risk.
- Surface competing approaches if there's real choice:
| Approach | Tool/Library | Pros | Cons | Status |
|---|---|---|---|---|
| ... | ... | ... | ... | maintained / abandoned / beta |
- Pick one. State why. If 2+ are credible, build quick variants within the spike.
- Skip research for pure logic with no external dependencies.
Use Hermes tools for the research step:
web_search("python websocket streaming libraries 2025"), find candidatesweb_extract(urls=["https://websockets.readthedocs.io/..."]), read the actual docs (returns markdown)terminal("pip show websockets | grep Version"), check what's installed in the project's venv
For libraries without docs pages, clone and read their README.md / examples/ via read_file. Context7 MCP (if the user has it configured) is also a good source, mcp_*_resolve-library-id then mcp_*_query-docs.
4. Build
One directory per spike. Keep it standalone.
spikes/
├── 001-websocket-streaming/
│ ├── README.md
│ └── main.py
├── 002a-pdf-parse-pdfjs/
│ ├── README.md
│ └── parse.js
└── 002b-pdf-parse-camelot/
├── README.md
└── parse.py
Bias toward something the user can interact with. Spikes fail when the only output is a log line that says "it works." The user wants to feel the spike working. Default choices, in order of preference:
- A runnable CLI that takes input and prints observable output
- A minimal HTML page that demonstrates the behavior
- A small web server with one endpoint
- A unit test that exercises the question with recognizable assertions
Depth over speed. Never declare "it works" after one happy-path run. Test edge cases. Follow surprising findings. The verdict is only trustworthy when the investigation was honest.
Avoid unless the spike specifically requires it: complex package management, build tools/bundlers, Docker, env files, config systems. Hardcode everything, it's a spike.
Building one spike, a typical tool sequence:
terminal("mkdir -p spikes/001-websocket-streaming")
write_file("spikes/001-websocket-streaming/README.md", "# 001: websocket-streaming\n\n...")
write_file("spikes/001-websocket-streaming/main.py", "...")
terminal("cd spikes/001-websocket-streaming && python3 main.py")
# Observe output, iterate.
Parallel comparison spikes (002a / 002b), delegate. When two approaches can run in parallel and both need real engineering (not 10-line prototypes), fan out with delegate_task:
delegate_task(tasks=[
{"goal": "Build 002a-pdf-parse-pdfjs: ...", "toolsets": ["terminal", "file", "web"]},
{"goal": "Build 002b-pdf-parse-camelot: ...", "toolsets": ["terminal", "file", "web"]},
])
Each subagent returns its own verdict; you write the head-to-head.
5. Verdict
Each spike's README.md closes with:
## Verdict: VALIDATED | PARTIAL | INVALIDATED
### What worked
- ...
### What didn't
- ...
### Surprises
- ...
### Recommendation for the real build
- ...
VALIDATED = the core question was answered yes, with evidence. PARTIAL = it works under constraints X, Y, Z, document them. INVALIDATED = doesn't work, for this reason. This is a successful spike.
Comparison spikes
When two approaches answer the same question (002a / 002b), build them back to back, then do a head-to-head comparison at the end:
## Head-to-head: pdfjs vs camelot
| Dimension | pdfjs (002a) | camelot (002b) |
|-----------|--------------|----------------|
| Extraction quality | 9/10 structured | 7/10 table-only |
| Setup complexity | npm install, 1 line | pip + ghostscript |
| Perf on 100-page PDF | 3s | 18s |
| Handles rotated text | no | yes |
**Winner:** pdfjs for our use case. Camelot if we need table-first extraction later.
Frontier mode (picking what to spike next)
If spikes already exist and the user says "what should I spike next?", walk the existing directories and look for:
- Integration risks, two validated spikes that touch the same resource but were tested independently
- Data handoffs, spike A's output was assumed compatible with spike B's input; never proven
- Gaps in the vision, capabilities assumed but unproven
- Alternative approaches, different angles for PARTIAL or INVALIDATED spikes
Propose 2-4 candidates as Given/When/Then. Let the user pick.
Output
- Create
spikes/(or.planning/spikes/if the user is using GSD conventions) in the repo root - One dir per spike:
NNN-descriptive-name/ README.mdper spike captures question, approach, results, verdict- Keep the code throwaway, a spike that takes 2 days to "clean up for production" was a bad spike
When not to use it
The source lists three clear exclusions:
- The answer is knowable from docs or reading code, just do research, don't build
- The work is production path, use the
planskill instead - The idea is already validated, jump straight to implementation
Limits and gotchas
- Spikes are disposable by design. Do not clean them up for production. If you find yourself spending two days refactoring a spike, you missed the point.
- The verdict is only as good as the depth of your investigation. One happy-path run is not enough.
- Avoid complex tooling unless the spike specifically requires it. Hardcode everything.
- The skill assumes Hermes tools are available. If
web_searchorweb_extractare not configured, research steps will fail.
Related skills
- sketch, for early ideation before a spike
- subagent-driven-development, for delegating work to subagents
- plan, for production-path work after validation
Attribution
Adapted from the GSD (Get Shit Done) project's /gsd-spike workflow, MIT © 2025 Lex Christopherson (gsd-build/get-shit-done). The full GSD system offers persistent spike state, MANIFEST tracking, and integration with a broader spec-driven development pipeline; install with npx get-shit-done-cc --hermes --global.