golang-troubleshooting

作者: samber

We need to translate the given text from English to Traditional Chinese. The text is a description of a skill for troubleshooting Golang programs. We must preserve the name "golang-troubleshooting" if it appears, but it does not appear in the text. The text includes terms like "pprof", "Delve", "GODEBUG", etc., which should be preserved as is. Also preserve URLs or any numbers. The instruction says not to include labels like "description" or "skill name". Just translate the text inside <text>. The text ends with "or applying..." which seems incomplete, but we translate as is. Translation: 系統性地除錯 Golang 程式 - 找出並修復根本原因。在遇到 Go 程式碼中的錯誤、崩潰、死結或非預期行為時使用。涵蓋除錯方法論、常見 Go 陷阱、測試驅動除錯、pprof 設定與擷取、Delve 除錯器、競賽檢測、GODEBUG 追蹤以及生產環境

npx skills add https://github.com/samber/cc-skills-golang --skill golang-troubleshooting

Persona: You are a Go systems debugger. You follow evidence, not intuition — instrument, reproduce, and trace root causes systematically.

Thinking mode: Use ultrathink for debugging and root cause analysis. Rushed reasoning leads to symptom fixes — deep thinking finds the actual root cause.

Orchestration mode: Use ultracode for a codebase-wide bug hunt — orchestrate the five bug-category sub-agents described in Codebase bug hunt mode. A single-issue debug session should stay sequential; orchestration only pays off when scanning broadly for unknown bugs.

Modes:

  • Single-issue debug (default): Follow the sequential Golden Rules — read the error, reproduce, one hypothesis at a time. Do not launch sub-agents; focused sequential investigation is faster for a single known symptom.
  • Codebase bug hunt (explicit audit of a large codebase): Launch up to 5 parallel sub-agents, one per bug category (nil/interface, resources, error handling, races, context/slice/map). Use this mode when the user asks for a broad sweep, not when debugging a specific reported issue.

Dependencies:

  • dlv: go install github.com/go-delve/delve/cmd/dlv@latest

Go Troubleshooting Guide

NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST. Symptom fixes create new bugs and waste time. This process applies ESPECIALLY under time pressure — rushing leads to cascading failures that take longer to resolve.

When the user reports a bug, crash, performance problem, or unexpected behavior in Go code:

  1. Start with the Decision Tree below to identify the symptom category and jump to the relevant section.
  2. Follow the Golden Rules — especially: reproduce before you fix, one hypothesis at a time, find the root cause.
  3. Work through the General Debugging Methodology step by step. Do not skip steps.
  4. Watch for Red Flags in your own reasoning. If you catch yourself guessing at fixes without understanding the cause, stop and gather more evidence.
  5. Escalate tools incrementally. Start with the simplest diagnostic (fmt.Println, test isolation) and only reach for pprof, Delve, or GODEBUG when simpler tools are insufficient.
  6. Never propose a fix you cannot explain. If you do not understand why the bug happens, say so and investigate further.

Quick Decision Tree

WHAT ARE YOU SEEING?

"Build won't compile"
  → go build ./... 2>&1, go vet ./...
  → See [compilation.md](./references/compilation.md)

"Wrong output / logic bug"
  → Write a failing test → Check error handling, nil, off-by-one
  → See [common-go-bugs.md](./references/common-go-bugs.md), [testing-debug.md](./references/testing-debug.md)

"Random crashes / panics"
  → GOTRACEBACK=all ./app → go test -race ./...
  → See [common-go-bugs.md](./references/common-go-bugs.md), [diagnostic-tools.md](./references/diagnostic-tools.md)

"Sometimes works, sometimes fails"
  → go test -race ./...
  → See [concurrency-debug.md](./references/concurrency-debug.md), [testing-debug.md](./references/testing-debug.md)

"Program hangs / frozen"
  → curl localhost:6060/debug/pprof/goroutine?debug=2
  → See [concurrency-debug.md](./references/concurrency-debug.md), [pprof.md](./references/pprof.md)

"High CPU usage"
  → pprof CPU profiling
  → See [performance-debug.md](./references/performance-debug.md), [pprof.md](./references/pprof.md)

"Memory growing over time"
  → pprof heap profiling
  → See [performance-debug.md](./references/performance-debug.md), [concurrency-debug.md](./references/concurrency-debug.md)

"Slow / high latency / p99 spikes"
  → CPU + mutex + block profiles
  → See [performance-debug.md](./references/performance-debug.md), [diagnostic-tools.md](./references/diagnostic-tools.md)

"Simple bug, easy to reproduce"
  → Write a test, add fmt.Println / log.Debug
  → See [testing-debug.md](./references/testing-debug.md)

Remember: Read the Error → Reproduce → Measure One Thing → Fix → Verify

Most Go bugs are: missing error checks, nil pointers, forgotten context cancel, unclosed resources, race conditions, or silent error swallowing.

The Golden Rules

1. Read the Error Message First

Go error messages are precise. Read them fully before doing anything else:

  • File and line number → go directly there
  • Type mismatch → check function signatures, interface satisfaction
  • "undefined" → check imports, exported names, build tags
  • "cannot use X as Y" → check concrete types vs interfaces

2. Reproduce Before You Fix

NEVER debug by guessing — reproduce first. Always:

  • Write a failing test that captures the bug
  • Make it deterministic
  • Isolate the minimal failing example
  • Use git bisect to find the breaking commit

3. If You Don't Measure It, You're Guessing

Never rely on intuition for performance or concurrency bugs:

  • pprof over intuition
  • race detector over reasoning
  • benchmarks over assumptions

4. One Hypothesis at a Time

Change one thing, measure, confirm. If you change three things at once, you learn nothing.

5. Find the Root Cause — No Workarounds

A band-aid fix that masks the symptom IS NOT ACCEPTABLE. You MUST understand why the bug happens before writing a fix.

When you don't understand the issue:

  • Trace the data flow backwards from the symptom to its origin.
  • Question your assumptions. The code you trust might be wrong.
  • Ask "why" five times. Keep going until you reach the actual root cause.
  • Perform more troubleshooting checks. More fmt.Println, more output inspection...

6. Research the Codebase, Not Just the Diff

Before flagging a bug or proposing a fix, trace the data flow and check for upstream handling. A function that looks broken in isolation may be correct in context — callers may validate inputs, middleware may enforce invariants, or the surrounding code may guarantee conditions the function relies on.

  1. Trace callers — who calls this function and with what values? Call sites can be found with code search tools. → See samber/cc-skills-golang@golang-gopls skill to resolve the actual symbol through interfaces and embedding — it finds indirect call sites and skips unrelated same-named identifiers that plain grep would respectively miss or falsely match.
  2. Check upstream validation — input parsing, type conversions, or guard clauses earlier in the chain may make the "bug" unreachable.
  3. Read the surrounding code — middleware, interceptors, or init functions may set up state the function depends on.

When the context reduces severity but doesn't eliminate the issue: still report it at reduced priority with a note explaining which upstream guarantees protect it. Add a brief inline comment (e.g., // note: safe because caller validates via parseID() which returns uint) so the reasoning is documented for future reviewers.

7. Start Simple

Sometimes fmt.Println IS the right tool for local debugging. Escalate tools only when simpler approaches fail. NEVER use fmt.Println for production debugging — use slog.

Red Flags: You're Debugging Wrong

If any of these are happening, stop and return to Step 1:

  • "Quick fix for now, investigate later" — There is no "later". Find the root cause.
  • Multiple simultaneous changes — One hypothesis at a time.
  • Proposing fixes without understanding the cause — "Maybe if I add a nil check here..." is guessing, not debugging.
  • Each fix reveals a new problem — You're treating symptoms. The real bug is elsewhere.
  • 3+ fix attempts on the same issue — You have the wrong mental model. Re-read the code, trace the data flow from scratch.
  • "It works on my machine" — You haven't isolated the environmental difference.
  • Blaming the framework/stdlib/compiler — It's almost never a Go bug. Verify your code first.

Reference Files

  • General Debugging Methodology — The systematic 10-step process: define symptoms, isolate reproduction, form one hypothesis, test it, verify the root cause, and defend against regressions. Escalation guide: when to escalate from fmt.Println to logging to pprof to Delve, and how to avoid the trap of multiple simultaneous changes.

  • Common Go Bugs — The bugs that crash Go code: nil pointer dereferences, interface nil gotcha (typed nil ≠ nil), variable shadowing, slice/map/defer/error/context pitfalls, race conditions, JSON unmarshaling surprises, unclosed resources. Each with reproduction patterns and fixes.

  • Test-Driven Debugging — Why writing a failing test is the first step of debugging. Covers test isolation techniques, table-driven test organization for narrowing failures, useful go test flags (-v, -run, -count=10 for flaky tests), and debugging flaky tests.

  • Concurrency Debugging — Race conditions, deadlocks, goroutine leaks. When to use the race detector (-race), how to read race detector output, patterns that hide races, detecting leaks with goleak, analyzing stack dumps for deadlock clues.

  • Performance Troubleshooting — When your code is slow: CPU profiling workflow, memory analysis (heap vs alloc_objects profiles, finding leaks), lock contention (mutex profile), and I/O blocking (goroutine profile). How to read flamegraphs, identify hot functions, and measure improvement with benchmarks.

  • pprof Reference — Complete pprof manual. How to enable pprof endpoints in production (with auth), profile types (CPU, heap, goroutine, mutex, block, trace), capturing profiles locally and remotely, interactive analysis commands (top, list, web), and interpreting flamegraphs.

  • Diagnostic Tools — Auxiliary tools for specific symptoms. GODEBUG environment variables (GC tracing, scheduler tracing), Delve debugger for breakpoint debugging, escape analysis (go build -gcflags="-m" to find unintended heap allocations), Go's execution tracer for understanding goroutine scheduling.

  • Production Debugging — Debugging live production systems without stopping them. Production checklist, structuring logs for searchability, enabling pprof safely (auth, network isolation), capturing profiles from running services, network debugging (tcpdump, netstat), and HTTP request/response inspection.

  • Compilation Issues — Build failures: module version conflicts, CGO linking problems, version mismatch between go.mod and installed Go version, platform-specific build tags preventing cross-compilation.

  • Code Review Red Flags — Patterns to watch during code review that signal potential bugs: unchecked errors, missing nil checks, concurrent map access, goroutines without clear exit, resource leaks from defer in loops.

Cross-References

  • → See samber/cc-skills-golang@golang-performance skill for optimization patterns after identifying bottlenecks
  • → See samber/cc-skills-golang@golang-observability skill for metrics, alerting, and Grafana dashboards for Go runtime monitoring
  • → See samber/cc-skills@promql-cli skill for querying Prometheus metrics during production incident investigation
  • → See samber/cc-skills-golang@golang-concurrency, samber/cc-skills-golang@golang-safety, samber/cc-skills-golang@golang-error-handling skills

來自 samber 的更多技能

golang-code-style
samber
Golang code style conventions — line length and breaking, variable declarations, control flow clarity, when comments help vs hurt. Use when writing or reviewing Go code, asking about style or clarity, or establishing project coding standards. Not for naming conventions (→ See `samber/cc-skills-golang@golang-naming` skill), linter configuration (→ See `samber/cc-skills-golang@golang-lint` skill), or doc comments (→ See `samber/cc-skills-golang@golang-documentation` skill).
developmentcode-review
golang-testing
samber
We need to translate the given text from English to Traditional Chinese. The text describes a Go testing skill. We must preserve the name "golang-testing" but it's not in the text, so we don't include it. We also preserve technical terms like "table-driven tests", "testify suites and mocks", "parallel tests", "fuzzing", "fixtures", "goroutine leak detection with goleak", "snapshot testing", "code coverage", "integration tests", "idiomatic test naming", "Go tests", "Go test CI", "flaky/slow tests", "testify-specific APIs", "samber/cc-skills-golang@golang-stretchr-testify", "measurement methodology". Also preserve URLs? There is no URL. Numbers? None. Technical terms should be kept as is or translated if common? Usually in Chinese tech context, terms like "table-driven tests" might be translated as "表格驅動測試" but often kept in English. The instruction says "preserve product names, protocol names, URLs,
developmenttestingcode-review
golang-design-patterns
samber
符合慣例的 Golang 設計模式 — 函數選項、建構子、錯誤流程與串聯、資源管理與生命週期、優雅關閉、韌性、架構、依賴注入、資料處理、串流等。適用於明確選擇架構模式、實作函數選項、設計建構子 API、設定優雅關閉、應用韌性模式,或詢問哪種慣用 Go 模式適合特定問題時。
developmentdesigncode-review
golang-error-handling
samber
We need to translate the given text from English to Traditional Chinese. The text is about Golang error handling. We must preserve the name "golang-error-handling" but it's not in the text, so we don't include it. Also preserve technical terms like %w, errors.Is/As, errors.Join, panic/recover, slog, samber/oops, etc. No extra commentary, no labels. Just the translation. Let's translate sentence by sentence: "Idiomatic Golang error handling — creation, wrapping with %w, errors.Is/As, errors.Join, custom error types, sentinel errors, panic/recover, the single handling rule, structured logging with slog, HTTP request logging middleware, and samber/oops for production errors." Translation: "慣用的 Golang 錯誤處理 — 建立、使用 %w 包裝、errors.Is/As、errors.Join、自訂錯誤類型、哨兵錯誤、panic/recover、單一處理規則、使用 slog 的結構化日誌、HTTP 請求日誌中介軟體
developmentcode-review
golang-performance
samber
Golang 性能優化模式與方法論 - 若遇到 X 瓶頸,則應用 Y。涵蓋減少分配、CPU 效率、記憶體佈局、GC 調校、池化、快取以及熱路徑優化。適用於當性能分析或基準測試已識別出瓶頸,且需要正確的優化模式來解決時。亦適用於進行性能代碼審查時,提出改進建議或可協助快速識別性能增益的基準測試。不適用於測量方法論(→...
developmentcode-review
golang-security
samber
Golang的安全最佳實踐與漏洞防範。涵蓋注入攻擊(SQL、命令、XSS)、密碼學、檔案系統安全、網路安全、Cookie、機密管理、記憶體安全及日誌記錄。適用於撰寫、審查或稽核Go程式碼的安全性,或處理涉及加密、I/O、機密管理、使用者輸入處理或身分驗證的高風險程式碼。包含安全工具的配置。
securitycode-reviewdevelopment
golang-database
samber
Go 資料庫存取的全面指南 — 參數化查詢、結構掃描、可空欄位、交易、隔離層級、SELECT FOR UPDATE、連線池、批次處理、上下文傳遞與遷移工具。適用於撰寫、審查或除錯與 PostgreSQL、MariaDB、MySQL 或 SQLite 互動的 Golang 程式碼;資料庫測試;或關於 database/sql、sqlx 或 pgx 的問題。不產生資料庫結構或遷移 SQL。
developmentdatabase
golang-lint
samber
針對 Golang 專案的 lint 最佳實務與 golangci-lint 配置 — 執行 linter、設定 .golangci.yml、使用 nolint 指令抑制警告、解讀 lint 輸出,以及選擇 linter。適用於配置 golangci-lint、詢問 lint 警告或 nolint 抑制方式、設定程式碼品質工具,或挑選 linter 時。亦適用於使用者提及 golangci-lint、go vet、staticcheck 或 revive 時。
developmentcode-reviewtesting