attribution

작성자: coreyhaines31

사용자가 어떤 마케팅이 실제로 전환과 매출을 유도하는지 파악하려 하거나, 기여 모델을 선택하거나 해석하려 하거나, 여러 도구 간의 상충되는 수치를 조정하려 할 때 사용합니다. 또한 사용자가 "기여", "기여 모델", "퍼스트터치 대 라스트터치", "멀티터치", "어떤 채널이 매출을 유도하는지", "실제 CAC는 얼마인지", "대시보드가 서로 다르게 나와요", "Google/Meta는 X라고 하는데 GA는 Y라고 해요", "미디어 믹스 모델", "MMM", "증분성", "지오 리프트", "홀드아웃 테스트", "어떻게 했는지..."라고 언급할 때도 사용합니다.

npx skills add https://github.com/coreyhaines31/marketingskills --skill attribution

Attribution

You help users answer the hardest question in marketing: which of my efforts actually caused this conversion and this revenue? Attribution is where marketers lose the most money — to channels that look good in one dashboard and terrible in another, to "direct" and "branded search" that hide the real source, and to models that quietly encode an opinion as if it were fact.

This skill has two pillars. Know which one the user needs before you dive in:

  • (A) Interpretation — choosing an attribution model, picking a measurement approach, and reconciling the conflicting numbers your tools report. This applies to everyone, even with zero engineering.
  • (B) Own your attribution (first-party) — instrumenting and stitching attribution yourself when you control the site/app. This is the build track. Use it when the user says "I want to track this myself" or is hitting a conversion that lives on a domain they don't own.

Most requests start with (A). Reach for (B) only when they control the surface and want to build.

Product context: check for .agents/product-marketing.md and read it if present — business type, sales cycle, and primary conversion drive almost every recommendation here.

Boundaries — what this skill does NOT own

State these up front so you don't rebuild neighboring skills:

  • General event tracking, tracking plans, UTM setup, GA4/GTManalytics. Attribution assumes tracking exists. The line: analytics = "what events and how to fire them"; attribution = "how touches join to conversions and survive to revenue."
  • Ad-platform pixels, CAPI, server-side conversion trackingads (references/conversion-tracking.md). Attribution consumes platform-reported numbers and corrects for their bias; it doesn't set up the pixels.
  • Pipeline stages, lead lifecycle, CRM revenue dashboardsrevops. Attribution feeds pipeline data; it doesn't define stages.
  • Showing up in / measuring AI searchai-seo. Attribution names AI traffic as a blind spot only.

Pillar A — Interpretation

1. What attribution can and can't tell you

Set expectations before touching a number:

  • Attribution is directional, not truth. It's a model of causality built from incomplete data (cookies expire, sessions fragment, offline touches vanish, people research on one device and buy on another). Treat it as a strong hint, never a verdict.
  • Every model is an opinion. "First-touch" says the first ad gets all the credit; "last-touch" says the closing click does. Both are wrong in opposite directions. Choosing a model is choosing whose story to believe — say so out loud.
  • The attribution gap is normal. The sum of channel-reported conversions almost always exceeds real conversions, because every platform claims credit for the same sale. Your job is to shrink and explain the gap, not to make the numbers tie out perfectly. They won't.

When a user demands one true number, reframe: "We can get you a defensible, consistent number and a read on which channels are trending up. A single objective truth doesn't exist — here's why, and here's what we use to make decisions anyway."

2. Attribution models

The six standard models and when each one lies:

ModelCredit ruleBest forHow it lies
First-touch100% to the first known touchTop-of-funnel / demand-gen valuation; short cyclesIgnores everything that closed the deal; over-credits awareness channels
Last-touch100% to the last touch before conversionDirect-response, quick e-commOver-credits bottom-funnel + branded search/direct; ignores what created demand
Last non-direct100% to last touch, skipping "direct"A cheap fix for direct pollutionStill single-touch; just moves the blind spot
LinearEqual credit to every touchLong, multi-touch journeys where every step mattersTreats a throwaway visit like a demo; flatters high-frequency channels
Time-decayMore credit to touches nearer conversionLonger cycles where recency mattersUnder-credits the top of funnel; still an assumption, not a measurement
Position-based (U-shaped)40% first, 40% last, 20% middleB2B with clear "created" + "closed" momentsThe 40/40/20 split is arbitrary; middle touches get shortchanged
Data-driven (algorithmic/Shapley)Credit from modeled marginal contributionHigh-volume accounts with enough conversionsA black box; needs volume; can't see offline/dark touches it was never fed

Rules of thumb:

  • Never report a single model in isolation for a long sales cycle. Show first-touch and last-touch side by side — the truth lives between them, and the gap between them is the insight.
  • Data-driven attribution needs volume (Google Ads historically gated it behind ~3,000 ad interactions and ~300 conversions in 30 days; it has since relaxed the minimums and made DDA the default, but low volume still makes it noise dressed as science). Use position-based instead when you're thin.
  • The model matters far less than being consistent and pairing it with an out-of-model sanity check (Pillar A §4, self-reported).

For the model math, worked examples of one journey scored six ways, and Shapley explained plainly, see references/attribution-models.md.

3. The three measurement paradigms

Models split credit within your tracked data. Paradigms are how you get at causality — increasingly rigorous, increasingly expensive:

ParadigmWhat it isAnswersNeedsWatch out
MTA (multi-touch attribution)Stitch user-level touches, apply a model"Which touchpoints appear on converting journeys?"Clean cross-device user-level trackingCookie loss + privacy have gutted user-level data; it silently under-measures
MMM (media/marketing mix modeling)Top-down regression of spend vs. outcomes over time"What's each channel's aggregate contribution, including offline/brand?"2–3 yrs of weekly data, spend variationCorrelational; slow to react; needs real budget swings to learn
Incrementality (geo holdout, PSA, ghost ads, on/off)Controlled experiment: exposed vs. withheld"Did this channel cause lift I wouldn't have gotten anyway?"Ability to withhold; enough volume for significanceThe gold standard, but you can only test a few things at a time

How to choose: small budget / short cycle → good UTM + last-non-direct + a self-reported survey beats a fancy model. Mid budget, several channels → MTA for day-to-day + periodic incrementality tests on your biggest line items. Large budget, offline + brand spend → MMM for the portfolio + incrementality to validate MMM's coefficients. Incrementality is the tiebreaker whenever two channels both claim the same conversions.

Decision table by budget × sales cycle × channel count, and how to read a geo-holdout / PSA test (not a stats tutorial), in references/measurement-paradigms.md.

4. Self-reported attribution

The most underused signal, and often the most honest for long cycles and dark social. A post-conversion "How did you hear about us?" survey catches what tracking structurally cannot: podcasts, word of mouth, Slack communities, a founder's tweet, "a friend told me."

  • When it beats tracking: long consideration cycles, high word-of-mouth, brand/community-led, or heavy dark-social (see §5). If a big slice of your journeys are "direct," you have a self-reported-shaped hole.
  • Ask at the moment of conversion (signup, first purchase, demo request) — highest recall, before memory fades.
  • Wording: open-ended ("How did you first hear about us?") captures dark social; a short pick-list is easier to quantify but pre-biases the answer. Best practice: pick-list of your known channels plus a free-text "other/tell us more."
  • Treat it as a triangulation input, not gospel — recall is fuzzy and people credit the memorable touch, not the first. It's the out-of-model check that keeps your tracked models honest.
  • On the build side, this is a form field written to your CRM/analytics as a person property — see Pillar B and references/first-party-tracking.md.

5. Reconciling conflicting sources

The request behind most attribution work: "Google says 50, Meta says 40, GA says 60, my CRM says 35 — who's right?" Nobody is. Here's the framework.

Why each source systematically lies:

SourceBiased towardBecause
Ad platforms (Google/Meta/LinkedIn)Over-counts itselfClaims view-through + click conversions in its own window; every platform counts the same sale; motivated to look good
GA / web analyticsLast non-direct clickLoses cross-device, loses cookie-blocked users, dumps the unknown into direct
CRMWhatever the rep typed / the form capturedHuman entry, lead-source overwrites, offline deals with no digital trail
Self-reported surveyThe memorable touchRecall bias; under-counts boring-but-real touches like retargeting

How to triangulate:

  1. Pick one source of truth for the conversion count — usually your CRM or backend (the system where money is real). Everything else explains where those came from, they don't get to redefine how many.
  2. Never sum across platforms. If Google and Meta both claim a conversion, you have one conversion with two claimants, not two conversions. De-dupe against the source-of-truth total.
  3. Read directional agreement, not absolute match. If every source says paid search is up and organic is down this quarter, that trend is trustworthy even though no two numbers match.
  4. Use self-reported as the tiebreaker when platforms fight over the same conversions, and incrementality when the stakes justify a test.
  5. Expect and budget for the gap. Report "platforms claim N; we can verify M; the delta is over-claiming + view-through + untracked — here's our best allocation."

The output is an honest allocation with confidence levels, not a false reconciliation to the decimal.

6. The blind spots

Where conversions hide, making real channels look weak:

  • Direct — the junk drawer. Bookmarks and typed URLs, yes, but also stripped referrers, app-to-web, dark social, and any touch your tracking dropped. A large direct share is a measurement problem, not a channel.
  • Branded search — people who discovered you elsewhere and Googled your name. Last-touch hands the credit to paid/organic branded search; the real driver was whatever made them search. Segment branded vs. non-branded or you'll defund the top of funnel.
  • Dark social — sharing that carries no referrer: DMs, Slack/Discord, podcasts, newsletters, screenshots. Structurally invisible to tracking; self-reported is the only way to see it (§4).
  • AI traffic — assistants and AI search increasingly influence buyers, then send them via branded search or direct, so the AI touch is invisible in analytics. Name it and hand deeper work to ai-seo.

The through-line: when "direct" and "branded search" dominate, your top of funnel is working and your attribution is hiding it. Say that explicitly — it's the single most common misread in marketing.

7. Business-type fork

Defaults differ sharply. Summary here; full playbooks in references/by-business-type.md.

  • B2B SaaS (long cycle, sales-assisted): journeys span weeks–months and multiple people, so single-touch models mislead badly. Anchor on the CRM as source of truth, use first-touch + position-based side by side, lean hard on self-reported at demo/signup, and treat pipeline/revenue attribution (→ revops) as the real scoreboard. Offline touches (events, sales convos) make MTA weakest and self-reported strongest here.
  • Ecommerce / DTC (short cycle, self-serve): fast journeys, high volume, spend concentrated in paid social + search. Anchor on platform ROAS but distrust it (iOS/CAPI inflation), validate with MMM once spend is material and incrementality/geo-holdouts on your biggest channels, and use a post-purchase survey to catch what pixels miss. Last-touch is defensible for quick-turn SKUs; MMM+incrementality is how you allocate the real budget.

Pillar B — Own your attribution (first-party)

Use this when the user controls the site/app and wants to instrument attribution themselves — especially for a conversion that happens on a domain they don't own (a SavvyCal/Calendly/Cal.com booking, a Stripe Checkout page). This pillar is grounded in real production builds; the full runbook with code patterns is in references/first-party-tracking.md. The essentials:

The identity graph

First-party attribution is one idea: join anonymous browsing to the eventual conversion.

  1. A visitor arrives anonymously; your analytics tool assigns an anonymous distinct_id and stamps first-touch properties ($initial_referrer, $initial_utm_*) on their events.
  2. At conversion (signup, booking, purchase) you call identify() with a stable id (email or user UUID). This merges the anonymous history into a known person — first-touch now survives all the way to the conversion.
  3. Every conversion event can now be broken down by first-touch channel. That's the whole game.

Closing the identify() gap

The most common first-party failure: nothing ever calls identify(), so conversions never join to browsing history and every customer looks like they appeared from nowhere. (Framing adapted from Tessa Kriesel's PostHog approach.) The fix is to call identify at each real conversion. Audit first — many SaaS apps already identify at signup; don't rebuild what works. Find the specific un-instrumented conversions and close only those.

Stitching conversions on a third-party domain

The one case that needs real machinery: a conversion that completes on a domain you don't control (a booking tool, a hosted checkout). You can't run your analytics there, so:

  1. At click time, a capture-phase link decorator appends the visitor's anonymous distinct_id to the outbound URL via the tool's metadata passthrough (e.g. ?metadata[ph_distinct_id]=<id>). One document-level listener covers every CTA — no per-link edits.
  2. The third-party tool stores that metadata and returns it in its webhook.
  3. Your webhook handler fires an identity merge ($identify with the booking email as distinct_id and the smuggled anonymous id as $anon_distinct_id) plus a conversion event — joining the booking back onto the marketing journey.

Guardrails (do not skip)

  • Anonymity guard — fail closed. Only ever smuggle the anonymous id. After identify(), the current id becomes the user's email/UUID; leaking that into a third-party URL or merging on it corrupts profiles (person A's email folds into whoever books). Reject ids that look like PII (contain @), cap length, and when identity is ambiguous, send nothing. If the app identifies by UUID, test distinct_id === device_id rather than an @ check.
  • First-touch data quality. Redirects overwrite the true first touch. Exclude OAuth/checkout referrers (accounts.google.com, checkout.stripe.com, login.*), your own subdomains (self-referrals), and dev hosts (localhost) from referrer classification. This is usually a settings change, not code, and it's the highest-trust-per-effort fix.
  • Cross-subdomain stitching. Marketing site → app on a subdomain must share one analytics project + a cross-subdomain cookie, or the journey breaks at the handoff. Expect near-zero numbers until the stitch is verified in prod — don't panic at empty data; use a campaign-window heuristic fallback and backfill the pre-stitch cohort in the meantime (details in the reference).

Reporting and the last mile

The first payoff is one insight: your conversion event broken down by first-touch channel ($initial_utm_source / $initial_referring_domain), and — joined to revenue — channel → conversion → revenue. Confirm first-touch vs. last-touch config in the tool (many default to last-touch; first-party attribution wants $initial_*).

But first-touch alone can't run the multi-touch models from §2. Store the full ordered touch path (not just $initial_*) and the build track feeds the interpretation track — you can score your own journeys position-based / linear / time-decay instead of only reading about them.

The last mile — get it into the CRM (production refinement from Tessa Kriesel). A breakdown in an analytics tool is a report; sales and lifecycle act on attribution written onto the record. Sync a source field with confidence and basis (journey-linked vs self-reported vs campaign-window fallback) plus a Paid-vs-Organic read off the medium, rolled up to the account (not just the contact — one B2B org is several people with mixed work/personal emails). How pipeline/lifecycle then use it is revops' job.

The pattern is tool-agnostic: identify + merge exists in PostHog, Segment, Amplitude, and via user-id in GA4; the third-party stitch works with any tool that has a metadata passthrough + webhook. PostHog + SavvyCal are the worked example in references/first-party-tracking.md.


Output format

Deliver an attribution readout, not a data dump:

# Attribution Readout — [date]

## The question
[What decision this informs — e.g. "where should next quarter's budget go?"]

## Source of truth
[Which system defines the conversion count, and why]

## What each source says
| Channel | Platform-reported | GA | CRM | Self-reported | Our read |
|---------|------------------|----|----|--------------|----------|
[De-duped against source of truth; not summed]

## Model comparison (for long cycles)
[First-touch vs last-touch side by side; the gap is the insight]

## Confidence & gaps
[The attribution gap, the blind spots, what we can't see]

## Recommendation
[Allocation call with confidence levels; the tiebreaker test worth running]

Tool Integrations

For implementation, see the tools registry. Key tools:

ToolBest ForMCPGuide
PostHogFirst-party attribution, identify/merge, funnels-posthog.md
GA4Web analytics, model comparison, user-id stitchingga4.md
DubShort-link + click attributiondub-co.md
SegmentCDP — route identify/track to every destination-segment.md
HubSpotCRM lead-source + self-reported fieldshubspot.md
SalesforceCRM as revenue source of truth-salesforce.md
SupermetricsPull platform numbers into one place to reconcilesupermetrics.md
RB2BDe-anonymize B2B website visitors-rb2b.md

Related Skills

  • analytics — event tracking, tracking plans, UTMs, GA4/GTM setup. Do this before attribution.
  • ads — ad-platform pixels, CAPI, server-side conversion tracking (references/conversion-tracking.md).
  • revops — pipeline stages, lead lifecycle, CRM revenue reporting. Attribution feeds it.
  • ai-seo — the AI-search attribution blind spot in depth.
  • ab-testing — controlled experiments; the incrementality mindset applied to on-site changes.

coreyhaines31의 다른 스킬

seo-audit
coreyhaines31
사용자가 자신의 사이트에 대한 SEO 감사, 검토 또는 진단을 원할 때 사용합니다. 또한 사용자가 "SEO 감사", "기술적 SEO", "왜 순위가 오르지 않나요", "SEO 문제", "온페이지 SEO", "메타 태그 검토", "SEO 건강 점검", "트래픽이 떨어졌어요", "순위를 잃었어요", "구글에 나타나지 않아요", "사이트가 순위에 오르지 않아요", "구글 업데이트가 저를 강타했어요", "페이지 속도", "코어 웹 바이탈", "크롤 오류", 또는 "색인 문제"를 언급할 때도 사용합니다. 사용자가 "내 SEO가 안 좋아요" 또는 "도와주세요..."와 같이 모호하게 말하는 경우에도 사용합니다.
marketingresearchdata-analysis
copywriting
coreyhaines31
사용자가 홈페이지, 랜딩 페이지, 가격 페이지, 기능 페이지, 소개 페이지, 제품 페이지 등 모든 페이지에 대한 마케팅 카피를 작성, 재작성 또는 개선하려 할 때 사용합니다. 또한 사용자가 "카피 작성", "이 카피 개선", "이 페이지 재작성", "마케팅 카피", "헤드라인 도움", "CTA 카피", "가치 제안", "태그라인", "서브헤드라인", "히어로 섹션 카피", "어브 더 폴드", "이 카피가 약해요", "더 설득력 있게 만들어 주세요", "내 제품 설명을 도와주세요"라고 말할 때도 사용합니다.
marketingcreativecommunication
marketing-psychology
coreyhaines31
사용자가 심리학 원리, 정신 모델, 또는 행동 과학을 마케팅에 적용하고자 할 때 사용합니다. 또한 사용자가 '심리학', '정신 모델', '인지 편향', '설득', '행동 과학', '사람들이 구매하는 이유', '의사 결정', '소비자 행동', '앵커링', '사회적 증거', '희소성', '손실 회피', '프레이밍', 또는 '넛지'를 언급할 때도 사용합니다. 마케팅 맥락에서 사람들이 어떻게 생각하고 결정을 내리는지 이해하거나 활용하려는 경우에 사용하세요. 적용을 위해...
marketingresearch
content-strategy
coreyhaines31
사용자가 콘텐츠 전략을 계획하거나, 어떤 콘텐츠를 만들지 결정하거나, 다룰 주제를 파악하려 할 때 사용합니다. 또한 사용자가 "콘텐츠 전략", "무엇에 대해 써야 할까", "콘텐츠 아이디어", "블로그 전략", "주제 클러스터", "콘텐츠 기획", "편집 일정", "콘텐츠 마케팅", "콘텐츠 로드맵", "어떤 콘텐츠를 만들어야 할까", "블로그 주제", "콘텐츠 핵심 주제", 또는 "무엇을 써야 할지 모르겠어"라고 언급할 때도 사용합니다. 누군가 어떤 콘텐츠를 만들지 결정하는 데 도움이 필요할 때마다 사용하세요.
marketingresearchcreative
ai-seo
coreyhaines31
사용자가 AI 검색 엔진을 위한 콘텐츠 최적화, LLM에 인용되기, 또는 AI 생성 답변에 표시되기를 원할 때 사용합니다. 또한 사용자가 'AI SEO', 'AEO', 'GEO', 'LLMO', 'answer engine optimization', 'generative engine optimization', 'LLM optimization', 'AI Overviews', 'optimize for ChatGPT', 'optimize for Perplexity', 'AI citations', 'AI visibility', 'zero-click search', 'how do I show up in AI answers', 'LLM mentions', 또는 'optimize for Claude/Gemini'를 언급할 때도 사용합니다. 누군가가... 할 때마다 사용하세요.
marketingresearch
programmatic-seo
coreyhaines31
사용자가 템플릿과 데이터를 사용하여 SEO 기반 페이지를 대량으로 생성하려 할 때 사용합니다. 또한 사용자가 "programmatic SEO", "템플릿 페이지", "대량 페이지", "디렉토리 페이지", "지역 페이지", "[키워드] + [도시] 페이지", "비교 페이지", "통합 페이지", "SEO를 위한 다수 페이지 생성", "pSEO", "100개 페이지 생성", "데이터 기반 페이지", "템플릿 랜딩 페이지"를 언급할 때도 사용합니다. 다양한 키워드나 위치를 대상으로 유사한 페이지를 많이 만들고자 할 때 이 기능을 사용하세요. 예를 들어...
marketingdata-analysisweb-scraping
marketing-ideas
coreyhaines31
사용자가 SaaS 또는 소프트웨어 제품에 대한 마케팅 아이디어, 영감 또는 전략이 필요할 때. 또한 사용자가 '마케팅 아이디어', '성장 아이디어', '마케팅 방법', '마케팅 전략', '마케팅 전술', '홍보 방법', '성장 아이디어', '또 무엇을 시도할 수 있을까', '이걸 어떻게 마케팅해야 할지 모르겠어', '마케팅 브레인스토밍', '어떤 마케팅을 해야 할까'라고 물을 때 사용합니다. 누군가 막혀 있거나 성장을 위한 영감을 찾고 있을 때 출발점으로 사용하세요. 구체적인...
marketing
copy-editing
coreyhaines31
사용자가 기존 마케팅 카피를 편집, 검토, 개선하거나 오래된 콘텐츠를 업데이트하려 할 때 사용합니다. 또한 사용자가 '이 카피를 편집해 줘', '내 카피를 검토해 줘', '카피 피드백', '교정', '이걸 다듬어 줘', '더 좋게 만들어 줘', '카피 정리', '이걸 간결하게 해 줘', '이게 어색하게 읽혀', '이 텍스트를 정리해 줘', '너무 장황해', '메시지를 날카롭게 해 줘', '이 콘텐츠를 새롭게 해 줘', '이 페이지를 업데이트해 줘', '이 콘텐츠는 오래됐어', 또는 '콘텐츠 감사'라고 말할 때도 사용합니다. 사용자가 이미 카피를 가지고 있고
documentcommunicationmarketing