Hermes Agent DOCX Skill: Create, Edit, and Verify Word Documents
Create, read, edit Word .docx documents and templates.
Written by Neura Market from the official Hermes Agent documentation for Docx. Commands, paths, and version numbers are reproduced from the source unchanged.
Read the official documentationThe DOCX skill bundled with Hermes Agent lets you create, read, and edit Word documents (.docx and .dotx) programmatically. You would reach for it when a user asks for a report, memo, letter, or any deliverable as a Word file, or when you need to extract, reorganize, or redline content inside an existing document. It covers both high-level document creation via the docx-js library and surgical XML editing for tracked changes, comments, and formatting that the library cannot handle.
What it does
This skill turns natural-language requests into structured Word documents. It can write a new .docx from scratch using a Node.js script (docx-js), read the text out of an existing .docx with pandoc, and edit an existing document by unzipping the archive, modifying the XML directly, and rezipping. It also handles tracked changes (redlining), comments, table of contents, images, and validation against the OOXML schema. The skill is installed by default with Hermes Agent and runs on Linux, macOS, and Windows.
Before you start
The skill depends on three external tools and two Python packages. Run the following command to install everything on Linux:
npm ls docx --depth=0 2>/dev/null | grep -q docx || npm install docx # creation (docx-js)
pip show pandoc >/dev/null 2>&1 || true; which pandoc || sudo apt install -y pandoc # reading
which soffice || sudo apt install -y libreoffice # rendering/verification
which pdftoppm || sudo apt install -y poppler-utils # PDF → images
pip install defusedxml lxml # validation scripts
On macOS, use Homebrew:
brew install pandoc libreoffice poppler
The docx npm package is required only for creating new documents. The Python packages defusedxml and lxml are needed for the validation and merge scripts. LibreOffice is used to render .docx to PDF for visual inspection, and poppler-utils provides pdftoppm to turn PDF pages into JPEG images.
Quick Reference
| Task | Approach |
|---|---|
| Create a new document | Write a docx (npm) script, see gotchas below |
| Edit an existing document | unzip → edit word/document.xml → zip (docx-js cannot open existing files) |
| Read content | pandoc -t markdown file.docx (or read_file, which auto-extracts .docx text) |
Script paths below are relative to this skill's directory.
Creating with docx-js, gotchas
Write the script and require('docx'). The model knows the API; these are the footguns:
- Page size defaults to A4. For US Letter set
page: { size: { width: 12240, height: 15840 } }(DXA; 1440 = 1″). - Landscape: pass portrait dimensions and
orientation: PageOrientation.LANDSCAPE, docx-js swaps width/height internally. - Tables need dual widths: set
columnWidthson the table ANDwidthon every cell, both inWidthType.DXA(PERCENTAGE breaks in Google Docs). Column widths must sum to the table width. - Table shading: use
ShadingType.CLEAR, neverSOLID(renders black). - Lists: never insert
•literally; use anumberingconfig withLevelFormat.BULLET. ImageRunrequirestype:("png","jpg", …).PageBreakmust be inside aParagraph.- Never use
\n, use separateParagraphelements. - TOC: headings must use built-in
HeadingLevel.*; custom heading styles needoutlineLevelset or they won't appear. - Don't use a table as a horizontal rule, use a paragraph bottom border instead.
- Dot-leader / right-aligned-on-same-line: use
PositionalTab(alignment: PositionalTabAlignment.RIGHT,leader: PositionalTabLeader.DOT) inside aTextRun, not literal.or space padding.
Verify the output
After writing a .docx, render it and look at it:
python scripts/office/soffice.py --headless --convert-to pdf output.docx
pdftoppm -jpeg -r 100 output.pdf page
ls page-*.jpg # then inspect each with vision_analyze
pdftoppm zero-pads page numbers to the width of the page count (page-01.jpg…page-12.jpg).
Editing existing documents
Legacy .doc files must be converted first: python scripts/office/soffice.py --headless --convert-to docx file.doc.
unzip -q doc.docx -d unpacked/
find unpacked -type l -delete # strip symlink entries — docx from external parties is untrusted
python scripts/merge_runs.py unpacked/ # coalesce fragmented runs so text is findable
# edit unpacked/word/document.xml in place — do NOT reformat or pretty-print
(cd unpacked && rm -f ../out.docx && zip -Xr ../out.docx .)
python scripts/office/validate.py out.docx --original doc.docx # XSD checks; --auto-repair fixes common issues
# redlining? add --author "<the name you redlined under>" to check every edit is tracked
Word splits text across many <w:r> runs (revision ids, spell-check markers), so a phrase you can see in the document often doesn't exist as a contiguous string in the XML. merge_runs.py merges adjacent identically-formatted runs in word/document.xml without changing content or rendering; it also accepts a .docx directly (python scripts/merge_runs.py doc.docx -o merged.docx).
Tracked changes: when redlining, validate with --author "" (needs --original), it reports any text you changed without a <w:ins>/<w:del> around it, which is easy to do by accident and invisible in the accepted view. Wrap runs in <w:ins>/<w:del> with w:id, w:author, w:date attributes. Inside <w:ins>, the text element is <w:t>, not <w:delText>. A deleted paragraph mark (<w:del><w:r><w:rPr>…</w:rPr><w:br w:type="page"/></w:r></w:del>) means "merge this paragraph into the next", so deleting a paragraph outright is that plus a <w:del> around every run. The <w:rPr> must come before the rPr's other children; their order is schema-enforced.
To produce a clean copy with all tracked changes accepted: python scripts/accept_changes.py in.docx out.docx.
Accepting a deleted paragraph mark should join that paragraph to the one below it, so a paragraph whose runs are all deleted vanishes. Word does this; accept_changes.py and pandoc --track-changes=accept don't always. Both fail the same way, they strip the deleted text but leave the emptied paragraph behind, which reads as a stray empty bullet when it was auto-numbered:
pandoc --track-changes=acceptnever joins the paragraphs.accept_changes.py(LibreOffice) joins them correctly, except when the deleted paragraph is followed by an empty spacer paragraph.
An empty bullet in either view is an artifact of that view, not a defect in the document. Check paragraph deletions in the XML.
Comments
Comments require six cross-linked files. Use the helper, directory mode when you'll also be editing document.xml (saves an unzip/rezip cycle), .docx-direct mode otherwise:
# Against an already-unpacked directory (preferred when also placing markers)
python scripts/comment.py unpacked/ "Fees & expenses cap is too low"
python scripts/comment.py unpacked/ "Agreed" --parent 0
# Against a .docx directly
python scripts/comment.py contract.docx "This cap is too low" -o annotated.docx
The script writes comments.xml, commentsExtended.xml, commentsIds.xml, commentsExtensible.xml, the relationships, and the content-type overrides. Comment IDs are auto-assigned. It then prints the <w:commentRangeStart>/<w:commentRangeEnd>/<w:reference> snippet to add to word/document.xml so the comment anchors to specific text, until you place those markers, the comment exists but is not visible.
Pitfalls
- Don't round-trip OOXML through
xml.etree.ElementTree, it rewrites namespace prefixes and corrupts the file. Usedefusedxml.minidomfor scripted transforms. - Zip from INSIDE the unpacked directory (
cd unpacked && zip -Xr ../out.docx .) andrmthe target first, or deleted parts survive in the archive.
Verification
python scripts/office/validate.py out.docx --original in.docx, schema, relationship, and content-type checks; every failure names its fix.- Render to PDF → images (see "Verify the output") and inspect each page with
vision_analyze, look for broken tables, missing images, spacing artifacts, leftover placeholder text.
When not to use it
Do not use this skill for PDFs (use the pdf skill), spreadsheets (xlsx), or presentations (powerpoint). The skill also cannot open existing .docx files with docx-js; editing requires the unzip-edit-rezip workflow described above.
Limits and gotchas
- The
accept_changes.pyscript and pandoc both fail to join paragraphs after accepting a deleted paragraph mark in certain cases, leaving an empty bullet. This is a known artifact of the view, not a document defect. - Round-tripping through
xml.etree.ElementTreecorrupts the file due to namespace prefix rewriting. - Zipping from outside the unpacked directory or without removing the old archive first leaves stale content in the output.
- The
--authorflag for redlining validation requires the--originalflag.
Related skills
pdf (PDF work), xlsx (spreadsheets), powerpoint (decks), ocr-and-documents (scanned input extraction).