Docx

작성자: Anthropic

포괄적인 문서 생성, 편집 및 분석 기능을 제공하며, 변경 내용 추적, 주석, 서식 유지, 텍스트 추출을 지원합니다. Claude가 전문 문서(.docx 파일) 작업이 필요할 때: (1) 새 문서 생성, (2) 콘텐츠 수정 또는 편집, (3) 변경 내용 추적 작업, (4) 주석 추가 또는 기타 문서 작업에 사용됩니다. 라이선스: 독점 라이선스. LICENSE.txt에 전체 조건이 명시되어 있습니다.

npx skills add https://github.com/anthropics/skills --skill docx

DOCX creation, editing, and analysis

A .docx is a ZIP archive of XML files. Choose your approach by task:

TaskApproach
Create a new documentWrite a docx (npm) script — see gotchas below
Edit an existing documentunzip → edit word/document.xmlzip (docx-js cannot open existing files)
Read contentpandoc -t markdown file.docx

Script paths below are relative to this skill's directory.

Creating with docx-js — gotchas

docx is preinstalled — do not run npm install first; write the script and require('docx') directly. Only if that require fails: npm install docx. The model knows the API; these are the footguns:

  • Page size defaults to A4. For US Letter set page: { size: { width: 12240, height: 15840 } } (DXA; 1440 = 1″).
  • Landscape: pass portrait dimensions and orientation: PageOrientation.LANDSCAPE — docx-js swaps width/height internally.
  • Tables need dual widths: set columnWidths on the table AND width on every cell, both in WidthType.DXA (PERCENTAGE breaks in Google Docs). Column widths must sum to the table width.
  • Table shading: use ShadingType.CLEAR, never SOLID (renders black).
  • Lists: never insert literally; use a numbering config with LevelFormat.BULLET.
  • ImageRun requires type: ("png", "jpg", …).
  • PageBreak must be inside a Paragraph.
  • Never use \n — use separate Paragraph elements.
  • TOC: headings must use built-in HeadingLevel.*; custom heading styles need outlineLevel set or they won't appear.
  • Don't use a table as a horizontal rule — use a paragraph bottom border instead.
  • Dot-leader / right-aligned-on-same-line: use PositionalTab (alignment: PositionalTabAlignment.RIGHT, leader: PositionalTabLeader.DOT) inside a TextRun, not literal . or space padding.

Verify the output

After writing a .docx, render it and look at it:

python scripts/office/soffice.py --headless --convert-to pdf output.docx
pdftoppm -jpeg -r 100 output.pdf page
ls page-*.jpg   # then Read the images

pdftoppm zero-pads page numbers to the width of the page count (page-01.jpgpage-12.jpg).

Editing existing documents

Legacy .doc files must be converted first: python scripts/office/soffice.py --headless --convert-to docx file.doc.

unzip -q doc.docx -d unpacked/
find unpacked -type l -delete   # strip symlink entries — docx from external parties is untrusted
python scripts/merge_runs.py unpacked/   # coalesce fragmented runs so text is findable
# edit unpacked/word/document.xml in place — do NOT reformat or pretty-print
(cd unpacked && rm -f ../out.docx && zip -Xr ../out.docx .)
python scripts/office/validate.py out.docx --original doc.docx   # XSD checks; --auto-repair fixes common issues
# redlining? add --author "<the name you redlined under>" to check every edit is tracked

Word splits text across many <w:r> runs (revision ids, spell-check markers), so a phrase you can see in the document often doesn't exist as a contiguous string in the XML. merge_runs.py merges adjacent identically-formatted runs in word/document.xml without changing content or rendering; it also accepts a .docx directly (python scripts/merge_runs.py doc.docx -o merged.docx).

Tracked changes: when redlining, validate with --author "<the name you redlined under>" (needs --original) — it reports any text you changed without a <w:ins>/<w:del> around it, which is easy to do by accident and invisible in the accepted view. Wrap runs in <w:ins>/<w:del> with w:id, w:author, w:date attributes. Inside <w:del>, the text element is <w:delText>, not <w:t>. A deleted paragraph mark (<w:pPr><w:rPr><w:del w:id=".." w:author=".." w:date=".."/></w:rPr></w:pPr>) means "merge this paragraph into the next" — so deleting a paragraph outright is that plus a <w:del> around every run. The <w:del/> must come before the rPr's other children; their order is schema-enforced.

To produce a clean copy with all tracked changes accepted: python scripts/accept_changes.py in.docx out.docx.

Accepting a deleted paragraph mark should join that paragraph to the one below it, so a paragraph whose runs are all deleted vanishes. Word does this; accept_changes.py and pandoc --track-changes=accept don't always. Both fail the same way — they strip the deleted text but leave the emptied paragraph behind, which reads as a stray empty bullet when it was auto-numbered:

  • pandoc --track-changes=accept never joins the paragraphs.
  • accept_changes.py (LibreOffice) joins them correctly, except when the deleted paragraph is followed by an empty spacer paragraph.

An empty bullet in either view is an artifact of that view, not a defect in the document. Check paragraph deletions in the XML.

Comments

Comments require six cross-linked files. Use the helper — directory mode when you'll also be editing document.xml (saves an unzip/rezip cycle), .docx-direct mode otherwise:

# Against an already-unpacked directory (preferred when also placing markers)
python scripts/comment.py unpacked/ "Fees & expenses cap is too low"
python scripts/comment.py unpacked/ "Agreed" --parent 0

# Against a .docx directly
python scripts/comment.py contract.docx "This cap is too low" -o annotated.docx

The script writes comments.xml, commentsExtended.xml, commentsIds.xml, commentsExtensible.xml, the relationships, and the content-type overrides. Comment IDs are auto-assigned. It then prints the <w:commentRangeStart>/<w:commentRangeEnd>/<w:commentReference> snippet to add to word/document.xml so the comment anchors to specific text — until you place those markers, the comment exists but is not visible.

Dependencies

docx (npm, preinstalled — install only if require('docx') fails) · pandoc · LibreOffice (soffice) · pdftoppm (Poppler)

Anthropic의 다른 스킬

Algorithmic Art
Anthropic
p5.js와 시드 기반 무작위성 및 대화형 매개변수 탐색을 사용하여 알고리즘 아트를 생성합니다. 사용자가 코드를 사용한 아트 생성, 제너레이티브 아트, 알고리즘 아트, 플로우 필드 또는 파티클 시스템을 요청할 때 사용하세요. 저작권 침해를 피하기 위해 기존 아티스트의 작품을 복사하지 않고 독창적인 알고리즘 아트를 만드세요. 라이선스: 전체 약관은 LICENSE.txt에 있습니다.
creative
Internal Comms
Anthropic
회사가 선호하는 형식을 사용하여 모든 종류의 내부 커뮤니케이션을 작성하는 데 도움이 되는 리소스 세트입니다. Claude는 상태 보고서, 리더십 업데이트, 3P 업데이트, 회사 뉴스레터, FAQ, 사고 보고서, 프로젝트 업데이트 등 내부 커뮤니케이션 작성을 요청받을 때마다 이 스킬을 사용해야 합니다. 라이선스: LICENSE.txt에 전체 약관이 명시되어 있습니다.
Artifacts Builder
Anthropic
현대 프론트엔드 웹 기술(React, Tailwind CSS, shadcn/ui)을 사용하여 정교한 다중 구성 요소 claude.ai HTML 아티팩트를 생성하기 위한 도구 모음입니다. 상태 관리, 라우팅 또는 shadcn/ui 구성 요소가 필요한 복잡한 아티팩트에 사용하십시오. 단순한 단일 파일 HTML/JSX 아티팩트에는 적합하지 않습니다. 라이선스: LICENSE.txt에 전체 약관이 명시되어 있습니다.
development
Webapp Testing
Anthropic
Playwright를 사용하여 로컬 웹 애플리케이션과 상호작용하고 테스트하기 위한 툴킷입니다. 프론트엔드 기능 확인, UI 동작 디버깅, 브라우저 스크린샷 캡처, 브라우저 로그 보기를 지원합니다. 라이선스: LICENSE.txt에 전체 약관이 명시되어 있습니다.
development
XLSX
Anthropic
포괄적인 스프레드시트 생성, 편집 및 분석 기능을 제공하며 수식, 서식, 데이터 분석 및 시각화를 지원합니다. Claude가 스프레드시트(.xlsx, .xlsm, .csv, .tsv 등) 작업이 필요할 때: (1) 수식과 서식을 포함한 새 스프레드시트 생성, (2) 데이터 읽기 또는 분석, (3) 수식을 유지하면서 기존 스프레드시트 수정, (4) 스프레드시트 내 데이터 분석 및 시각화, 또는 (5) 수식 재계산 라이선스: 독점. LICENSE.txt에 전체 조건이 명시되어 있습니다.
document
Skill Creator
Anthropic
효과적인 스킬 생성을 위한 가이드입니다. 사용자가 Claude의 기능을 특화된 지식, 워크플로우 또는 도구 통합으로 확장하는 새 스킬을 만들거나 기존 스킬을 업데이트하려 할 때 사용해야 합니다. 라이선스: LICENSE.txt에 전체 약관이 명시되어 있습니다.
Slack Gif Creator
Anthropic
Slack에 최적화된 애니메이션 GIF를 생성하는 도구 키트로, 크기 제약 조건 검증기와 조합 가능한 애니메이션 프리미티브를 포함합니다. 사용자가 "X가 Y를 하는 Slack용 GIF를 만들어 줘"와 같은 설명으로 Slack용 애니메이션 GIF나 이모지 애니메이션을 요청할 때 이 스킬이 적용됩니다. 라이선스: 전체 약관은 LICENSE.txt에 있습니다.
creative
Theme Factory
Anthropic
아티팩트에 테마를 적용하여 스타일링하는 툴킷입니다. 이러한 아티팩트는 슬라이드, 문서, 보고서, HTML 랜딩 페이지 등이 될 수 있습니다. 생성된 모든 아티팩트에 적용할 수 있는 색상/글꼴이 포함된 10개의 사전 설정 테마가 있으며, 즉석에서 새 테마를 생성할 수도 있습니다. 라이선스: LICENSE.txt에 전체 약관이 명시되어 있습니다.
creative