report-to-google-doc

作者: openai

狭窄的转换技能。仅在用户明确要求将现有的本地或blob托管的HTML分析报告转换为Google文档、DOCX等格式时调用。

npx skills add https://github.com/openai/role-specific-plugins --skill report-to-google-doc

Report To Google Doc

Use this skill only when the user explicitly needs a shareable Google Drive document from an existing HTML analytics report. The source must be an HTML report: a local file, a downloaded blob-hosted report, or a report produced by $build-report HTML mode. This skill does not convert a live MCP app report directly.

The expected path is HTML -> DOCX -> Drive upload. It is acceptable for Drive to host the upload as a DOCX-backed viewer file rather than a native Google Docs MIME type. Do not use the old Google Docs batch-update request path.

Workflow

  1. Resolve the HTML report.

    Use an absolute local path. If the user provides a remote report, retrieve it first and pass the local HTML file to the helper. If the file is a sign-in page, redirect page, or tiny stub, stop and obtain the real report.

  2. Run the bundled helper.

    python3 <REPORT_TO_GOOGLE_DOC_SKILL_DIR>/scripts/report_to_google_doc_plan.py \
      /absolute/path/to/report.html \
      --out-dir /tmp/report_to_google_doc_plan
    

    Omit --render-workers on the normal path. Only pass a worker count after benchmarking the same report family locally. If dependencies are missing, use a local virtual environment with beautifulsoup4, pillow, and python-docx; cairosvg or headless Playwright are optional renderers.

  3. Inspect helper outputs.

    Required outputs:

    • skeleton.txt: source text with stable placeholders
    • manifest.json: parsed headings, tables, callouts, lists, styles, links, and rendered visual inventory
    • preflight_checks.json: source, width, DOCX, and rendered-image checks
    • report.docx: generated local Word document
    • docx_upload_plan.json: compact upload instructions
    • placeholder_queries.json: source mapping debug labels

    Do not upload until preflight_checks.json has status: "passed" with zero errors. Warnings must either be fixed or called out in the handoff.

  4. Upload the DOCX.

    mcp__codex_apps__google_drive._upload_file({
      "file_uri": "/tmp/report_to_google_doc_plan/report.docx",
      "file_name": "Report Name.docx",
      "mime_type": "application/vnd.openxmlformats-officedocument.wordprocessingml.document"
    })
    

    Treat the returned Drive URL as the deliverable. Do not attempt to force native Google Docs conversion, and do not fall back to _batch_update_document.

  5. Validate the uploaded result.

    Confirm the uploaded file is readable and non-empty. Compare uploaded text against the source inventory: title, section headings, executive summary or answer callout, caveats, recommendations, source notes, and source links. Inspect the local report.docx structure when available: heading counts, lists, tables, hyperlink relationships, and image relationships should match manifest.json.

  6. Hand off the link.

    Return the Drive/Docs URL and, when useful, the local DOCX path or source HTML path. Connector success alone is not enough; the handoff is complete only after the uploaded file and local DOCX structure have been checked against the source report. Keep routine check and preflight details in support artifacts. Do not list internal checks in the user-facing handoff unless a check failed, was unavailable, or produced a user-relevant caveat.

Standards

  • Preserve every section, headline claim, metric card, metric definition, source note, chart takeaway, recommendation, caveat, and link from the HTML report.
  • Preserve semantic formatting: headings, paragraphs, inline bold/emphasis, inline code, positive/negative colors, links, lists, callouts, metric-card grids, tables, notes, captions, and charts.
  • Use DOCX-native structures wherever practical: headings, paragraphs, tables, bullets/numbered lists, links, inline images, paragraph shading, table cell shading, and text styles.
  • Keep the report text column readable. Tables, charts, rendered table grids, screenshots, and visual blocks must not exceed the DOCX page text width.
  • Preserve charts as inline images rendered from the source visual or its same-data SVG fallback, aligned to the same left edge as text and tables. A chart with missing bars, missing legend swatches, or all-black/all-white marks fails validation even if the DOCX contains an image object.
  • Preserve multi-column report blocks. A two-column grid made only of titled mini-tables can remain a two-column rendered image capped to the text width; mixed two-column blocks with narrative text, pills, callouts, or non-table panels should preserve that content natively instead of dropping it.
  • Do not expose customer-level details or sensitive links that were intentionally omitted from a sanitized report.
  • Do not write Google Docs batch-update artifacts such as seed_requests.json, remote_write_plan.json, or all_requests*.json.

Repairs

Use these fixes when validation exposes a conversion issue:

ProblemFix
Heading is regular weightFix the DOCX writer heading style or explicit run bolding.
Blank line after title or headingDelete spacer paragraphs; skeletons should emit \n, not \n\n, after headings.
Body paragraphs have too much spaceDelete literal blank paragraphs and set modest style spacing.
Table/image too wideSet table columns or image width to the DOCX text column width.
Chart has labels but missing barsInline SVG class styles or use a headless screenshot before inserting.
Two-column table group is flattenedTreat the grid as one layout block with mini-table titles preserved.
Executive summary, metric cards, or section is missingFix parser inventory before upload.
Chart overlaps tableInsert the image on its own paragraph after the table.
Chart duplicatedRemove the extra image; keep exactly one source-order copy.
Inline bold/code/color missingFix manifest range splitting in the DOCX writer.
Source links show raw URLsReplace with native linked labels or native bullets using HTML link text.

Local Checks

python3 -m py_compile <REPORT_TO_GOOGLE_DOC_SKILL_DIR>/scripts/report_to_google_doc_plan.py
python3 <REPORT_TO_GOOGLE_DOC_SKILL_DIR>/scripts/report_to_google_doc_plan.py \
  /absolute/path/to/report.html \
  --out-dir /tmp/report_to_google_doc_plan_smoke
jq '.status, .summary' /tmp/report_to_google_doc_plan_smoke/preflight_checks.json
test -f /tmp/report_to_google_doc_plan_smoke/report.docx
test -f /tmp/report_to_google_doc_plan_smoke/docx_upload_plan.json
git diff --check -- plugins/data-analytics/skills/build-report/report-to-google-doc

来自 openai 的更多技能

user-context
openai
加载或管理数据分析插件的持久化源路由偏好、引导逻辑、设置进度及语义层注册表。
official
notion-research-documentation
openai
研究Notion内容,并将其综合成带有引用的结构化简报、报告或对比。通过定向查询搜索并获取Notion页面,然后按主题组织发现,附带内联来源引用和参考文献部分。根据范围和用户目标,从四种输出格式(快速简报、研究摘要、对比、综合报告)中选择。使用内置模板创建和更新Notion页面;直接链接来源,并在新信息到达时跟踪变更...
official
rcsb-pdb-skill
openai
提交紧凑的RCSB PDB请求以获取核心元数据、Search API查询和FASTA下载。当用户需要简洁的RCSB摘要时使用;保存原始JSON或…
official
pdf
openai
PDF的读取、创建与验证,支持可视化渲染与程序化生成。使用Poppler(pdftoppm)将PDF页面渲染为PNG,以便在交付前直观检查布局、间距与排版;通过reportlab程序化生成PDF,确保格式可靠;利用pdfplumber或pypdf提取文本与元数据。执行质量标准:无文本裁剪、元素重叠、表格损坏或渲染伪影;仅使用ASCII连字符,引用内容需可读。使用...
official
test-coverage-improver
openai
改进OpenAI Agents JS mon
official
playwright
openai
基于终端驱动的浏览器自动化,支持元素快照与交互式UI工作流。通过playwright-cli包装脚本运行(需npx),支持无头模式与有头模式进行可视化调试。核心工作流:打开页面、获取快照以稳定元素引用、使用引用进行交互、在导航或DOM变更后重新快照。包含表单填写、点击、输入、多标签页管理、截图/PDF捕获及用于流程调试的追踪记录。元素引用(如e3、e15)...
official
ukb-topmed-phewas-skill
openai
通过接受rsID、GRCh37或GRCh38输入并解析为所需的GRCh38查询,获取单个变体的紧凑型UKB-TOPMed PheWAS摘要。当需要…时使用。
official
code-review-context
openai
模型可见上下文
official