nvidia-kaggle-skill

oleh nvidia

Gunakan untuk mengambil ringkasan kompetisi Kaggle, penulisan, riset diskusi/kernel, pengiriman, dan unggahan dataset. Bukan untuk kode ML yang tidak terkait.

npx skills add https://github.com/nvidia/nvidia-kaggle --skill nvidia-kaggle-skill

NVIDIA Kaggle Skill

Purpose

Use this skill for Kaggle competition work: context gathering, writeups, discussions, kernels, local reproduction, submission, and dataset upload.

Do not use it for unrelated ML training, generic notebook editing, general data analysis, or non-Kaggle dataset management unless the user explicitly ties the task to Kaggle.

Inputs

InputRequiredDescription
Kaggle slug, URL, writeup URL, kernel ref, or local folderDepends on taskPrimary target for the requested Kaggle action.
KAGGLE_API_TOKENRequired for API/CLI-backed workflowsKGAT token string for Kaggle API, CLI, and SDK calls.
Disk spaceRequired for kernel setupMust fit input datasets, competition data, models, and extracted archives.

Prerequisites

  • Run commands from this skill directory unless a referenced workflow says otherwise.
  • Install only the runtime packages needed for the requested workflow.
  • Set KAGGLE_API_TOKEN before API, CLI, kernel, discussion, dataset, or submission workflows.
  • Confirm local disk space before downloading competition data, kernel inputs, or extracted archives.
  • Require explicit user confirmation before sensitive, externally visible actions: competition submissions (each can consume a daily slot), dataset uploads, and creating a public dataset. Treat KAGGLE_API_TOKEN as a secret — never print, log, or echo it.

Runtime Dependencies

Install only the packages needed for the requested task into the current environment, then run scripts with python.

Kaggle API, CLI, kernels, discussions, datasets, competition pages, and writeups:

if command -v uv >/dev/null 2>&1; then
  uv pip install httpx kaggle kagglesdk nbformat pydantic python-dotenv rich
else
  python -m pip install httpx kaggle kagglesdk nbformat pydantic python-dotenv rich
fi

For API/CLI tasks, verify credentials before calling Kaggle:

: "${KAGGLE_API_TOKEN:?ERROR: KAGGLE_API_TOKEN environment variable is not set}"

Workflows

Use this workflow catalog to choose the right path. Run the direct script commands for quick tasks. For workflows that point to another markdown file, read that file only when the request needs that workflow.

Prefer the runtime's run_script helper when it exists, for example run_script("scripts/fetch_competition_info.py", args=["titanic"]). Otherwise run the equivalent python ./scripts/<script>.py ... command from this skill directory.

Competition Details

Use this when the user asks to retrieve or summarize a Kaggle competition overview, rules, evaluation, timeline, or dataset description.

Fetch overview:

python ./scripts/fetch_competition_info.py <competition-slug-or-url>

Fetch dataset description:

python ./scripts/fetch_dataset_info.py <competition-slug-or-url>

The scripts accept a bare competition slug or https://www.kaggle.com/competitions/<slug> URL and extract the slug automatically. Convert output to markdown when the user asks for saved documentation, using {slug}_competition_overview.md and {slug}_dataset_description.md in the current working directory.

Research Brief

Use this when the user asks you to research a competition and write a strategy brief in natural terms (e.g. "research this competition and brief me, with links and a few charts"). You chain the skill's individual research workflows yourself, write your own analysis/plotting code, and produce the brief. Read ./research-brief.md for the principles that keep the brief accurate, informative, and useful to a reader — how to cite real sources as links, and how to make plots honest and legible (every plotted number traces to what you gathered). These principles live in the skill so the user does not have to spell them out.

Writeups

Use this when the user asks to fetch one writeup, fetch top-k writeups, discover leaderboard writeup links, or summarize solution posts. Read ./writeups.md.

Discussions

Use this when the user asks for Kaggle competition discussions, community insights, questions, tips, or a specific discussion thread.

python ./scripts/discussion_ingest.py <competition_id> [--max-pages N] [--sort-by hotness|votes|comments|created|updated] [--page-size N] [--nofetch-comments]
python ./scripts/discussion_query.py <competition_id> [--search TERM] [--min-votes N] [--author NAME] [--limit N] [--as-json]
python ./scripts/discussion_read.py <discussion_id> [--competition-id ID]
python ./scripts/discussion_db_info.py [competition_id]

Storage:

PathContents
data/discussions.dbSQLite cache for discussions, comments, and competition metadata

Always run ingest before query/read if the database is empty. Keep retries bounded if Kaggle rate limits or API shapes change.

Kernels

Use this when the user asks to ingest, query, or read kernels; research top public kernels; fetch kernel scores; or analyze kernel lineage. Read ./kernels.md.

Kernel Setup

Use this when the user asks to download and reproduce a Kaggle notebook locally with its inputs. Read ./kernel-setup.md.

Submission

Use this when the user asks to push, poll, or submit a Kaggle kernel to a competition. Read ./submission.md.

Upload Dataset

Use this when the user wants to create or update a Kaggle dataset from local files.

python ./scripts/upload_dataset.py <path-to-data-folder> [--title "My Dataset"] [--public] [--version-notes "notes"] [--dir-mode zip|tar|skip] [--collaborator user:reader]

Defaults:

  • Datasets are private unless the user explicitly asks for --public.
  • If --title is omitted, derive it from the folder name.
  • If an existing dataset-metadata.json has description, keywords, subtitle, or license fields, preserve them.
  • If the dataset already exists, do not overwrite silently; ask for or use --version-notes.

Outputs

  • Competition detail scripts print cleaned text that can be saved as markdown.
  • Discussion scripts write/read data/discussions.db and print tables, JSON, or rendered threads.
  • Dataset upload writes dataset-metadata.json in the data folder and prints the Kaggle dataset URL.
  • Referenced workflows may write markdown reports, notebook caches, local kernel workspaces, or submission logs as described in their markdown files.

Troubleshooting

Use this table for common failure modes across Kaggle workflows. Workflow files may add only narrow entries that are not covered here.

SymptomCauseAction
KAGGLE_API_TOKEN missing or invalidAPI/CLI-backed workflow started without valid Kaggle credentials.Stop before Kaggle API/CLI calls, set KAGGLE_API_TOKEN, and rerun the exact command.
Empty discussion or kernel query resultsThe local cache has not been populated for that competition.Run the matching ingest script first, then query again.
Private, restricted, or unavailable Kaggle contentThe active account lacks access, rules were not accepted, or the content was removed.Report the URL/ref and ask the user for access context before retrying.
Kaggle API, SDK, rate-limit, or page-structure failureKaggle returned partial data, changed an API/layout, or limited requests.Preserve the failing command and output, keep retries bounded, and label unavailable evidence.
Disk space or archive extraction failureCompetition data, kernel inputs, models, or extracted archives exceed local capacity or extraction failed.Stop, report the partial workspace state, and ask before deleting files or retrying.
Submission retry or uncertain submission statusA successful submit can spend a competition submission slot.Read existing logs and require explicit user intent before rerunning a submission workflow.

Runtime Compatibility

This skill works with any agent runtime that follows the Agent Skills convention. Codex uses the repository checkout or plugin installation, Claude Code uses marketplace plugin installation, and Claude Agent SDK can load the same project-scoped plugin settings. Scripts are self-contained under this skill's scripts/ directory.

Lebih banyak skill dari nvidia

compileiq-debug
nvidia
Gunakan ketika ada yang salah: Search() menggantung, semua evaluasi mengembalikan INVALID_SCORE, skor tidak kunjung membaik, setiap konfigurasi mengembalikan angka yang sama, error ptxas…
create-github-pr
nvidia
Buat pull request GitHub menggunakan gh CLI. Gunakan saat pengguna ingin membuat PR baru, mengirimkan kode untuk ditinjau, atau membuka pull request. Kata kunci pemicu -…
nemoclaw-maintainer-cross-issue-sweep
nvidia
Memindai isu terbuka lainnya untuk menemukan isu yang mungkin juga diperbaiki atau secara tidak sengaja dirusak oleh suatu PR tertentu. Menghasilkan peluang perbaikan yang berdekatan dan risiko kontradiksi dengan file:baris…
fhir-basics
nvidia
Mengajarkan agen cara kerja API FHIR R4, sumber daya apa saja yang tersedia, cara melakukan kueri dengan parameter pencarian, dan cara mengurai semua format respons dengan benar…
compileiq-validate-result
nvidia
Gunakan SETELAH Pencarian selesai dan SEBELUM mengklaim percepatan atau mengirim ACF. Muat CSV dump_results, ekstrak kandidat top-K (tujuan tunggal)…
changelog-audit
nvidia
Audit Warp CHANGELOG.md sebelum rilis: pulihkan entri yang hilang, urutkan berdasarkan dampak pengguna, perbaiki bahasa entri, bungkus baris, dan (mode cabang rilis) naikkan bandingkan…
maintain-dynamic-plugins
nvidia
Mempertahankan pemuat plugin dinamis NeMo Relay, manifes, SDK asli Rust, protokol pekerja gRPC, SDK pekerja Python, dokumen, pengujian, dan cakupan alur kerja rilis
dgx-diagnose
nvidia
Diagnosis masalah umum DGX Station GB300 — crash CUDA, penargetan GPU yang salah, bug kontainer vLLM/SGLang, masalah status MIG, kesalahan NVLink/Fabric Manager,…