doc

作者: firecrawl

Use when the task involves reading, creating, or editing `.docx` documents, especially when formatting or layout fidelity matters; prefer `python-docx` plus…

npx skills add https://github.com/firecrawl/openai-skills --skill doc

DOCX Skill

When to use

  • Read or review DOCX content where layout matters (tables, diagrams, pagination).
  • Create or edit DOCX files with professional formatting.
  • Validate visual layout before delivery.

Workflow

  1. Prefer visual review (layout, tables, diagrams).
    • If soffice and pdftoppm are available, convert DOCX -> PDF -> PNGs.
    • Or use scripts/render_docx.py (requires pdf2image and Poppler).
    • If these tools are missing, install them or ask the user to review rendered pages locally.
  2. Use python-docx for edits and structured creation (headings, styles, tables, lists).
  3. After each meaningful change, re-render and inspect the pages.
  4. If visual review is not possible, extract text with python-docx as a fallback and call out layout risk.
  5. Keep intermediate outputs organized and clean up after final approval.

Temp and output conventions

  • Use tmp/docs/ for intermediate files; delete when done.
  • Write final artifacts under output/doc/ when working in this repo.
  • Keep filenames stable and descriptive.

Dependencies (install if missing)

Prefer uv for dependency management.

Python packages:

uv pip install python-docx pdf2image

If uv is unavailable:

python3 -m pip install python-docx pdf2image

System tools (for rendering):

# macOS (Homebrew)
brew install libreoffice poppler

# Ubuntu/Debian
sudo apt-get install -y libreoffice poppler-utils

If installation isn't possible in this environment, tell the user which dependency is missing and how to install it locally.

Environment

No required environment variables.

Rendering commands

DOCX -> PDF:

soffice -env:UserInstallation=file:///tmp/lo_profile_$$ --headless --convert-to pdf --outdir $OUTDIR $INPUT_DOCX

PDF -> PNGs:

pdftoppm -png $OUTDIR/$BASENAME.pdf $OUTDIR/$BASENAME

Bundled helper:

python3 scripts/render_docx.py /path/to/file.docx --output_dir /tmp/docx_pages

Quality expectations

  • Deliver a client-ready document: consistent typography, spacing, margins, and clear hierarchy.
  • Avoid formatting defects: clipped/overlapping text, broken tables, unreadable characters, or default-template styling.
  • Charts, tables, and visuals must be legible in rendered pages with correct alignment.
  • Use ASCII hyphens only. Avoid U+2011 (non-breaking hyphen) and other Unicode dashes.
  • Citations and references must be human-readable; never leave tool tokens or placeholder strings.

Final checks

  • Re-render and inspect every page at 100% zoom before final delivery.
  • Fix any spacing, alignment, or pagination issues and repeat the render loop.
  • Confirm there are no leftovers (temp files, duplicate renders) unless the user asks to keep them.

来自 firecrawl 的更多技能

firecrawl-research-index
firecrawl
使用 Firecrawl Research 查找回答研究查询的论文,采用语义搜索、语义与结构扩展以及正文内验证。对于任何文献查找/论文检索任务——无论是单篇论文查找还是完整的多篇论文集合——始终使用此技能。
data-analysisresearchweb-scraping
oracle
firecrawl
使用oracle CLI的最佳实践(提示词与文件打包、引擎、会话及文件附件模式)。
pinecone
firecrawl
面向生产级AI应用的托管向量数据库。全托管、自动扩缩容,支持混合搜索(稠密+稀疏)、元数据过滤和命名空间。
wpds
firecrawl
在构建利用WordPress设计系统(WPDS)及其组件、令牌、模式等的用户界面时使用。
audiocraft-audio-generation
firecrawl
用于音频生成的PyTorch库,包括文本生成音乐(MusicGen)和文本生成声音(AudioGen)。当需要从文本生成音乐时使用…
skypilot-multi-cloud-orchestration
firecrawl
跨多云编排机器学习工作负载,自动优化成本。当您需要在多个云上运行训练或批处理作业时使用,利用…
firecrawl-seo-audit
firecrawl
使用 Firecrawl 对网站进行 SEO 审计。适用于用户要求进行 SEO 审计、元数据和标题审查、站点地图/网站结构分析、关键词机会、竞争对手 SERP 对比,或优先搜索优化建议的场景。
data-analysisresearchweb-scraping
gh-issues
firecrawl
获取GitHub议题,生成子代理来实施修复并开启拉取请求,然后监控并处理PR审查评论。用法:/gh-issues [owner/repo] [--label…]