← the whole session plugin/skills/feature-measurement-workflow/SKILL.md

Canonical workflow for major feature creation with built-in measurement.

Canonical Major Feature Creation Workflow

This is the end-to-end workflow for creating, shipping, and measuring major features in any product. Every substantive feature follows this pipeline. Measurement is not an afterthought — it's step 2.

This workflow was built running a coaching product. Wherever a step below names that product's specific tool, metric, or policy, it's labelled as an example — swap in whatever you actually use. The decision tree in step 2 and the evaluation matrix in step 10 have no product baked in; they work for any feature in any product as written.


The Full Pipeline

┌─────────────────────────────────────────────────────────────┐
│  1. IDEATION                                                │
│     Intuitive, passion-driven:                              │
│     "Can I 2x the value users get?"                         │
│     "Can I 2x engagement / conversion / time-to-ROI?"       │
│                                                             │
│  2. LEVERAGE + MEASUREMENT DECISION  ← NEW                  │
│     Identify highest leverage within what's possible         │
│     Run the Substantiveness Test (see below)                 │
│     If substantive: pick primary KPI + secondary KPI         │
│     Write one-line prediction                                │
│                                                             │
│  3. PLANNING                                                │
│     Scope, design, identify affected surfaces                │
│     Name the feature for tracking (see naming conventions)   │
│                                                             │
│  4. ALIGNMENT CHECK: How This Actually Works                │
│     Consult institutional memory for architectural context   │
│     Verify assumptions about existing code                   │
│                                                             │
│  5. EXECUTION                                               │
│     Build the feature                                       │
│     Add one feature-usage event during development           │
│     (one event: feature_{name}_used, fired on first use)     │
│                                                             │
│  6. TDD TESTING                                             │
│     Automated tests covering the new behavior                │
│                                                             │
│  7. UX TESTING FOR PERFECTION                               │
│     UX assignment tool — human verification                  │
│                                                             │
│  8. ALIGNMENT CHECK: Intent DB                              │
│     Verify against documented UX intentions                  │
│                                                             │
│  9. DEPLOY TO PRODUCTION                                    │
│     Ship to ALL users (no holdout unless pricing/paywall)    │
│     Release marker auto-logged at deploy time                │
│                                                             │
│ 10. 4-WEEK REVIEW  ← NEW                                    │
│     Evaluate adoption + KPI impact                           │
│     Advisory output: discovery problem? product problem?     │
│     Log result — calibrates future intuition                 │
└─────────────────────────────────────────────────────────────┘

Steps 4 and 8 (the alignment checks): if you have an institutional-memory search tool set up (see /alignment-harness:harness-setup), use it to check what already exists and what past sessions decided before you build. Otherwise, do it by hand: grep the repo's own docs and notes, read git log for the area you're touching, and search your own past Claude Code sessions (~/.claude/projects/*/*.jsonl) for prior decisions on this feature. Say plainly when the search turns up nothing — don't build on a guess dressed up as a finding.


Step 2 Deep Dive: The Substantiveness Test

Not every change needs measurement. Run this decision tree to determine if a feature is substantive enough:

Q1: Does it change a core user flow?
    (coaching, onboarding, paywall, engagement loop, navigation)
    → YES: substantive. Measure it.

Q2: Does it add a new surface area users interact with?
    (new page, new feature, new modal, new capability)
    → YES: substantive. Measure it.

Q3: Could it plausibly affect engagement, conversion, or churn?
    (even indirectly — e.g. faster load → more sessions)
    → YES: substantive. Measure it.

Q4: Is it cosmetic, admin-only, bug fix, or internal tooling?
    → NOT substantive. Ship without measurement.

If Q1-Q3 are all NO and Q4 is NO:
    → Use judgment. When in doubt, add measurement — it's 30 seconds.

When you mark it substantive:

  1. Pick a primary KPI — the single metric you believe this will most impact
  2. Pick a secondary KPI — a supporting signal from a different angle
  3. Write a one-line prediction: "I expect [primary KPI] to [increase/decrease] by ~[X%] within 4 weeks"

This prediction is not binding. It's calibration data for your future intuition.


Step 3 Deep Dive: Feature Naming

Feature names appear in your analytics tool's dashboards (Statsig, PostHog, or whatever you use) among potentially 100+ other features. They MUST be instantly identifiable by a person scanning a list.

See full naming conventions: if you have a /how-to-measure-feature-impact skill installed, it has the deeper naming reference; otherwise the quick rules below are the whole of it.

Quick rules:

  • Pattern: {domain}_{surface}_{specific_action}
  • UX-translated: describe what the USER experiences, not what the CODE does
  • Hyper-specific: another human must distinguish this from 99 other features at a glance

Step 5 Deep Dive: Adding the Feature-Usage Event

During development, add ONE event per feature, in whatever analytics tool your product already uses (Statsig, PostHog, Amplitude, Segment, a homegrown events table — the pattern is the same regardless of tool):

// Generic shape — adapt to your own event-tracking utility and naming scheme
export const APP_EVENTS = {
  // ... existing
  FEATURE_MY_NEW_THING_SHOWN: 'feature_my_new_thing_shown',
};

// Fire on first interaction per session:
trackAppEvent(APP_EVENTS.FEATURE_MY_NEW_THING_SHOWN, {
  sessionId: currentSessionId,
});

Example (opt-in reference, not a default): one project's events file is web-app/src/utilities/unifiedEventTracking.js, and a real event from it looks like FEATURE_COACHING_CHAT_SOURCE_CITATIONS_SHOWN: 'feature_coaching_chat_source_citations_shown'.

If nothing in the project tracks events yet, that's the gap to close before this step can complete — pick any lightweight tool (or a simple logged table) rather than skipping instrumentation entirely.

For detailed instrumentation including backend events, missing metrics, and edge cases: if you have a /how-to-measure-feature-impact skill installed, it covers that in depth; otherwise, the pattern above (one event, fired once per session on first use, named feature_{name}_used) is the whole of what's required to close this step.


Step 9 Deep Dive: Deploy Batch Markers

The Reality: Features Ship in Batches

Features don't ship individually to production. They accumulate on staging, then merge to production as a batch. This means:

  • The unit of measurement is the deploy batch, not the individual feature
  • You can't isolate "which feature caused the KPI change" within a batch — you measure the batch's collective impact
  • Individual feature adoption is still tracked per-feature (via your feature-usage events), but KPI impact is measured at the batch level

Automatic Deploy Markers

Every merge of staging → production should automatically fire a marker event in whatever analytics tool you use. This removes manual labor and creates clean before/after comparison points. (One approach fires this as a Statsig event — any tool with a timeline and custom-event support works the same way.)

The deploy marker captures:

Field Value
deployId Auto-generated unique ID for this deploy
deployTimestamp ISO timestamp of production merge
features Array of substantive features in this batch (each with their name, primary KPI, prediction)
totalFeatureCount How many substantive features in this batch
reviewDueDate deployTimestamp + 4 weeks
status pending_review → reviewed → archived

The marker event:

  • Fire a production_deploy event at deploy time with metadata about the batch
  • This creates a clean timestamp in your analytics tool's timeline for all metrics
  • Enables automatic before/after contrast on any metric without manual date selection

If you have no analytics tool configured at all yet, the fallback is a plain log: append one line (timestamp, batch contents, features and their predictions) to a markdown file in the folder printed by alignment-harness records deploys. It's less automatic than a timeline marker, but it still gives you a fixed point to compare metrics against later.

Per-Feature Predictions Within a Batch

Each substantive feature still gets its own prediction (written during step 2), but they're grouped under the deploy batch:

Deploy Batch: 2026-02-15T14:30:00Z
├── feature_coaching_chat_source_citations_shown
│   Primary: coaching_replies_per_logged_in_user (+15%)
│   Secondary: words_shared_per_session (+10%)
├── feature_onboarding_page_goal_selection_carousel
│   Primary: visitor_to_trial_rate (+8%)
│   Secondary: funnel_stage_velocity (+5%)
└── feature_settings_sidebar_memory_tier_upgrade_prompt
    Primary: trial_to_paid_rate (+3%)
    Secondary: distinct_features_per_logged_in_user (+5%)

At the 4-week review, you evaluate the batch's KPI impact collectively, then cross-reference individual feature adoption to reason about which features likely drove the change.

What This Means for the Review

  • Batch-level question: "Did this deploy improve our KPIs?"
  • Feature-level question: "Which features in this batch got adopted?"
  • Synthesis: If batch KPIs improved and Feature A had 45% adoption but Feature B had 3% adoption, Feature A likely drove the change
  • If batch KPIs declined: All features in the batch are suspects. Check adoption of each to narrow the investigation.

This marker is what powers the 4-week review. Without it, there's nothing to evaluate against.


Step 10 Deep Dive: The 4-Week Review

At 4 weeks post-deploy, every marked feature gets reviewed. This is NOT a pass/fail — it's a calibration exercise.

The Evaluation Matrix

Adoption Primary KPI What it means Action
High (>30% of logged-in users touched it) Improved Success. Feature works and users found it. Log win. Note prediction accuracy.
High Flat Product gap. Users found it but it didn't move the needle. Investigate: is the feature solving a real problem?
High Declined Possible harm. Users found it and something got worse. Urgent: investigate if feature caused regression or if external factor.
Low (<30%) Improved Indirect benefit. Small group using it but it's powerful for them. Consider: is this a power-user feature? That's fine.
Low Flat Discovery problem OR irrelevant feature. Run discovery advisory (see atomic skill).
Low Declined Investigate correlation vs causation. Low usage + decline is likely coincidental. Check for external factors. Don't blame the feature without evidence.

IMPORTANT: The 30% threshold is a reasoning aid, NOT a cutoff. A feature used by 5% of users that 10x'd their engagement is a massive win. Always reason about the numbers in context.

Discovery Advisory

When adoption is low and you suspect a discovery problem, the agent should:

  1. Identify the feature's entry points — how does a user currently reach this feature?
  2. Assess discoverability friction — is it buried? Behind a menu? Only visible in certain states?
  3. Generate discovery proposals with considerations:
    • Could it be surfaced in onboarding?
    • Could a changelog/spotlight announcement drive awareness?
    • Could a contextual tooltip introduce it at the right moment?
    • Could the feature be moved to a more visible surface?
    • Is the entry point named in a way users would look for?
  4. Present to the person as a prioritized list — each proposal with expected adoption lift and implementation effort

See full evaluation and advisory protocol: how-to-measure-feature-impact


The Metrics: Pick Your Own, Normalized to a Real User Base

CRITICAL: Whatever metrics you pick, express them as a ratio per active/logged-in user, not a raw count. Raw counts are meaningless because traffic fluctuates with marketing spend. Gating to a real, identifiable user (not an anonymous visitor) filters bots and normalizes for acquisition volume.

If your project doesn't have named KPIs yet, don't invent plausible-sounding ones. Look first at whatever's already inspectable — an existing analytics/events file, a dashboard config, a metrics.md — and if nothing exists, ask directly: "what are the two or three numbers that tell you this product is working?" Step 2 (the substantiveness test and prediction) still runs without a named KPI; it just can't tie the prediction to one until you have it, and it should say that plainly rather than guessing.

Every product needs its own set. Roughly three shapes are worth having, whatever you call them:

  • Engagement — is the core loop of the product getting used, and how deeply?
  • Funnel velocity — are people moving from first contact to paying/committed faster or slower?
  • Negative signal, tracked separately — churn, complaints, refunds. A feature can raise engagement AND raise churn at the same time (e.g. cognitive overload); if you fold both into one number you can't see that happening.

Example (opt-in reference, not a default to inherit): for a coaching app, the engagement metric is "coaching replies per logged-in user" ("this is the heartbeat — if this is up, the product is working"), plus "words shared per session" for depth and "distinct features used per logged-in user" for breadth. The funnel metrics are visitor→trial and trial→paid conversion rate. The negative-signal metric is churn rate, tracked independently of the engagement numbers above. None of these names are meant to be copied — they're what a heartbeat metric looks like for one specific product, so you can see the shape of a good one.


When to Use A/B Testing Instead

Default: Ship to all. Measure with time-series. This is a real, statable position, and it's a product-risk call you make for your own project — not a law of measurement. Set your own threshold for what's too risky to ship to everyone.

One example threshold, showing where to draw that line — use a real A/B test (an experiment with a holdout group) when:

  • Pricing / paywall changes — directly affects revenue, easy to get wrong
  • Core interaction changes to the product's central loop — if the primary experience itself changes (not additions, but modifications to existing behavior)
  • You genuinely don't know if it helps or hurts — and the risk of shipping a harmful change to 100% is too high

For A/B test setup: if you have a /how-to-create-manage-and-run-statsig-experiments skill installed (or an equivalent for whatever experimentation tool you use), follow that; otherwise any tool that can hold back a random slice of users and compare its metrics against everyone else works the same way.


Skill When to use
/how-to-measure-feature-impact (if installed) Atomic: naming, flags, evaluation, discovery advisory, edge cases
/how-to-create-a-new-statsig-kpi-and-confirm-it-works (if installed, or your tool's equivalent) When the metric you need doesn't exist yet
/how-to-create-manage-and-run-statsig-experiments (if installed, or your tool's equivalent) When you DO need a real A/B test (pricing, core interaction changes)
/leverage-predictor Predicting highest leverage move before ideation
/intent-db Alignment check against documented UX intentions
/ux-assignment-reasoning-protocol (if you don't have a dedicated UX-assignment generator installed) Generating UX verification assignments