Back to prompts

Prompt 034

Product outcome measurement

An outcome-review prompt for comparing mature, authorized post-release evidence with the original baseline, target behavior, customer result, and delivery cost.

Ready-to-use prompt

Copy the assignment.

# Product outcome measurement

## Goal

Determine whether a released product increment is associated with the intended user or business outcome, for whom, over what mature observation window, and with what guardrail effects.

Usage is not outcome. Adoption can explain exposure to the product change; it does not prove the change caused the result. A cohort that has not had time to experience the job is not a failed or retained cohort.

This is read-only analysis of authorized evidence already available in the workspace. Do not add instrumentation, connect to production or third-party systems, query live accounts, change data, deploy, message users, conduct new research, or collect new evidence.

Complete the analysis autonomously. Do not stop to ask clarifying questions. When the evidence is immature or cannot support the registered question, choose `TOO EARLY`. Never manufacture a baseline, comparison group, customer result, or causal claim to avoid that decision.

## Inputs

Locate and read:

* all applicable `AGENTS.md` files
* the approved product contract, target job, intended outcome, success threshold, and guardrails
* `docs/product/rollout-observation.md`, `docs/product/rollout-observation.yaml`, and `docs/product/rollout-observation-changelog.md`
* the exact release identity and verified rollout evidence referenced by that observation
* the rollout plan’s pre-registered measures, cohorts, windows, gates, and deviations
* the adoption-enablement definitions, intervention plan, and observed execution records
* authorized product, performance, reliability, support, billing, cost, and outcome extracts already present in the workspace
* event dictionaries, metric definitions, identity-resolution rules, data-quality audits, and known instrumentation changes
* cohort eligibility, exposure, account, segment, acquisition, and product-cadence records
* relevant incidents, migrations, pricing changes, campaigns, seasonality, and concurrent product releases
* current ICP, onboarding, retention, product feedback, and allowed proof in `docs/gtm/`
* consent, contract, data-processing, purpose-limitation, retention, deletion, residency, and research restrictions
* any existing `docs/product/outcome-measurement.md`, `docs/product/outcome-measurement.yaml`, and `docs/product/outcome-measurement-changelog.md`

Analyze only sources whose intended use is authorized for this product decision. The presence of customer or user data in a repository does not establish permission to reuse it. If authority, provenance, or identity resolution is unclear, exclude the source and record the resulting limitation.

The rollout-observation decision gates this analysis. `OBSERVED RELEASE` may proceed only within its verified exposure boundary. `PARTIAL OBSERVATION` may analyze an explicitly verified subset while inheriting every unknown, but choose `TOO EARLY` when that subset cannot answer the registered question. `NO VERIFIED RELEASE` forces `TOO EARLY`. `REVISIT ROLLOUT` stops outcome analysis and returns the discrepancy to the authorized rollout owner.

## Evidence standard

Label every material claim with exactly one shared product-evidence label:

* `OBSERVED`: directly supported by an identified, authorized source
* `DERIVED`: calculated or logically produced from identified `OBSERVED` inputs, with the method shown
* `ASSUMED`: an explicit premise not established by the available observations
* `UNKNOWN`: absent, inaccessible, immature, conflicting, unauthorized for this use, or not safely inferable

Every quantitative result must state:

* metric definition and unit
* numerator and denominator
* raw counts as well as rates
* source files and extraction dates
* eligibility, exposure, and maturity rules
* observation window and product cadence
* exclusions and missingness
* calculation method
* uncertainty and known bias

Do not replace missing observations with targets, industry benchmarks, anecdotes, or model-generated estimates. Keep plan values and actual observations separate.

## Step 1: Pass the data-authority gate

Create a source ledger before calculating outcomes. For each source record:

* owner and provenance
* collection purpose and authorized reuse
* consent or other lawful basis where applicable
* contractual and data-processing limitations
* fields actually needed for the decision
* access, retention, deletion, and residency requirements
* whether it contains personal, confidential, protected, or sensitive data
* whether aggregation or de-identification is sufficient
* known quality and identity-linkage limitations
* include or exclude decision with reason

Minimize data. Do not reproduce raw message content, names, emails, account secrets, or sensitive attributes in the artifacts. Suppress small cells that could identify a person. Do not join sources merely because identifiers can be matched; the join must also be authorized for the stated purpose.

If a legally, contractually, or ethically required approval is missing, exclude that evidence and refer the question to the authorized human owner. This prompt does not grant permission.

## Step 2: Freeze the question and evidence window

Restate, without improving it after seeing results:

* release and product-contract IDs
* intended eligible population and exclusions
* primary user or business outcome
* expected direction and registered threshold
* adoption prerequisite or mechanism
* guardrails and stop conditions
* baseline and comparison design, if pre-registered
* natural product cadence
* minimum exposure and follow-up windows
* analysis cutoff date
* concurrent changes known before analysis

If a threshold, segment, or measure is introduced after results are visible, label it exploratory. Do not use an exploratory cut to declare the registered outcome supported.

## Step 3: Audit cohort maturity and data quality

Reconstruct the eligible population and report the flow in raw counts:

```text
eligible → actually exposed → attempted → completed → valued → mature for outcome
```

Identify:

* planned versus actual exposure dates
* users or accounts never eligible or never exposed
* internal, test, bot, duplicate, refunded, or corrupted records
* partial, staggered, interrupted, or contaminated exposure
* late entrants and insufficient follow-up
* missing or changed event semantics
* identity mismatches across sources
* attrition and survivorship bias
* support-assisted versus unassisted paths
* delayed, episodic, seasonal, or censored outcomes

Do not treat an immature observation as zero. Exclude it from a mature denominator and report it separately. If the mature cohort cannot answer the question, choose `TOO EARLY` even when early directional signals look favorable.

## Step 4: Establish the comparison honestly

Use the strongest comparison the authorized evidence actually supports:

1. randomized or otherwise controlled exposure, if it genuinely occurred and remained intact
2. a pre-registered concurrent comparison with comparable eligibility
3. the same cohort’s valid pre-release baseline
4. a descriptive post-release cohort with no counterfactual

Document assignment, balance, crossover, contamination, concurrent interventions, missingness, and selection into adoption. Never call adopters versus non-adopters a causal comparison without addressing why people adopted.

If no valid baseline or comparison exists, report levels and changes descriptively. Do not invent a counterfactual from an industry average or an earlier incomparable customer set.

## Step 5: Calculate outcomes and guardrails

Calculate the registered primary outcome first. Then calculate only the secondary and guardrail measures needed to understand it.

Include, where authorized and relevant:

* outcome level and change, with raw counts
* exposure, attempt, completion, value, and recurrence as diagnostics
* time to first value and time to outcome
* reliability, error, latency, accessibility, and data-integrity effects
* support burden, manual intervention, and founder or operator labor
* infrastructure or delivery cost
* retention or expansion only when the cohort and cadence are mature
* adverse or uneven effects across pre-specified, decision-relevant cohorts

For segmented results, protect privacy, report cohort size, and distinguish a registered segment from exploratory slicing. Do not search many cuts for a flattering result. Do not infer sensitive traits or use protected attributes as performance explanations without a lawful, necessary, reviewed purpose.

Show missing values as `null` with an explanation. Do not average away severe incidents or aggregate a benefit that depends on one exceptional account.

## Step 6: State the causal limit

Classify the strongest warranted interpretation in plain language:

* descriptive: the outcome was observed after exposure
* associational: exposure or adoption and outcome moved together
* contribution-supported: timing, mechanism, comparison, and competing explanations make contribution credible, but not isolated
* causal: the design and execution support attribution under stated assumptions

List plausible competing explanations: selection, seasonality, maturation, regression to the mean, support attention, pricing or campaign changes, another release, measurement drift, or changed customer mix.

Do not use statistical significance as a substitute for practical importance, or practical importance as permission to claim causation. When uncertainty intervals or statistical tests are appropriate, show the method and assumptions; when sample size does not support them, report counts and limitations instead.

## Step 7: Compare against the registered decision rule

Create an outcome table with registered threshold, observed result, evidence label, maturity, uncertainty, guardrail status, and interpretation.

Choose exactly one decision:

* `OUTCOME SUPPORTED`: the mature authorized evidence meets the registered outcome rule, guardrails remain acceptable, and the causal wording does not exceed the design
* `MIXED`: meaningful benefit and weakness coexist across outcomes, guardrails, cohorts, or evidence sources; the tradeoff must be resolved
* `NOT SUPPORTED`: a mature, sufficiently valid observation did not meet the registered outcome rule, or material harm outweighs the intended result
* `TOO EARLY`: cohort maturity, exposure, sample, data quality, authorization, or observation time cannot yet answer the registered question

`OUTCOME SUPPORTED` does not automatically authorize expansion. `NOT SUPPORTED` does not identify the cause. `TOO EARLY` must include the earliest evidence-based reevaluation date or condition, not a fabricated calendar promise.

## Step 8: Desk-test the analysis

Verify:

1. Only authorized existing evidence was analyzed.
2. Planned exposure and interventions were not treated as observed.
3. Eligibility, exposure, adoption, and maturity denominators are distinct.
4. Immature records were not counted as success or failure.
5. The registered outcome was analyzed before exploratory cuts.
6. Every calculation is reproducible from cited sources without exposing personal data.
7. Causal language matches the comparison design.
8. Negative, inconvenient, and guardrail evidence remains visible.
9. No production query, user contact, data collection, or product action occurred.

## Decision

Lead with:

“As of [analysis cutoff], release [identifier] has [N] eligible, [N] exposed, and [N] outcome-mature units. For [registered outcome], the observed result is [result or unknown] against [threshold], with interpretation [descriptive / associational / contribution-supported / causal] and decision [OUTCOME SUPPORTED / MIXED / NOT SUPPORTED / TOO EARLY].”

Add: “This analysis used only authorized evidence already present in the workspace and did not contact users or change production.”

## Required artifacts

### 1. `docs/product/outcome-measurement.md`

Lead statement, data-authority ledger, frozen question, cohort flow, maturity and quality audit, comparison design, outcome and guardrail tables, causal limits, competing explanations, decision, desk test, and handoff.

### 2. `docs/product/outcome-measurement.yaml`

Include `version`, `status`, `as_of`, `decision`, `release_id`, `product_contract_id`, `registered_question`, `analysis_cutoff`, `data_authority`, `cohort_flow`, `maturity`, `comparison`, `primary_outcome`, `secondary_outcomes`, `guardrails`, `adoption_diagnostics`, `cost_and_labor`, `causal_strength`, `confounders`, `privacy_controls`, `assumptions`, `unknowns`, and `sources`.

Unknown values are `null` with an explanation. Keep raw personal-level data out of YAML.

### 3. `docs/product/outcome-measurement-changelog.md`

Append only. Each observation records timestamp, analysis cutoff, version, cohort definition, source snapshots, formula changes, maturity changes, decision, and reason. Never overwrite an earlier result or silently redefine its denominator.

## Closed-loop handoff

For `OUTCOME SUPPORTED`, pass the complete evidence package directly to Product cycle decision; do not manufacture friction work. For `MIXED` or `NOT SUPPORTED`, pass it to Product friction diagnosis so competing causes are tested before a product change is prioritized. For `TOO EARLY`, stop the series, preserve the current definitions, and return only when the named maturity or data condition is met; do not collect new data without separate authorization.

Pass validated current-product proof, observed adoption constraints, support burden, and explicit claim limits to Full-cycle GTM. GTM may update positioning, onboarding, retention, and product feedback only within the strength of this evidence. It must not turn association into a guarantee or publish customer-level outcomes without permission.

The friction diagnosis and product cycle decision return a bounded choice to the next discovery, prioritization, delivery, rollout, and measurement cycle. This prompt measures; it does not decide or execute the next product change.

## Boundaries

* Analyze only authorized evidence already available in the workspace.
* Do not connect to production, query live systems, add tracking, collect data, or contact users.
* Do not invent users, exposure, telemetry, baselines, comparisons, results, costs, retention, or customer outcomes.
* Do not reuse customer, support, research, or personal data outside its authorized purpose.
* Do not expose personal data, secrets, message content, or identifying small cohorts.
* Do not count immature cohorts as retained, successful, failed, or zero.
* Do not change registered thresholds after seeing the result.
* Do not claim causation beyond the design.
* Preserve uncertainty and human decision authority.

## Done when

* Data authority, cohort eligibility, actual exposure, maturity, and comparison strength are explicit.
* Every material claim uses `OBSERVED`, `DERIVED`, `ASSUMED`, or `UNKNOWN`.
* Every quantitative result has counts, formula, window, source, exclusions, and limitations.
* The registered outcome and all material guardrails are reported, including negative evidence.
* The Markdown, YAML, and append-only changelog agree.
* The decision uses one allowed value and follows the supported-outcome, diagnosis, or too-early branch before the next product cycle and Full-cycle GTM handoff.

Expected result

Decision-ready evidence, not manufactured certainty.

The finished work separates observed evidence, derived judgment, assumptions, and the next commitment-bearing test.

Use this when

Use this after the relevant rollout cohort has reached the product contract's observation window and authorized behavior, support, and outcome evidence can be reviewed.

What it produces

  • An outcome review at docs/product/outcome-measurement.md
  • A machine-readable outcome model at docs/product/outcome-measurement.yaml
  • An append-only record at docs/product/outcome-measurement-changelog.md
  • Baseline comparisons, cohort maturity, confidence limits, quality, and cost evidence
  • An OUTCOME SUPPORTED, MIXED, NOT SUPPORTED, or TOO EARLY decision

Guardrails

  • Does not invent users, events, baselines, outcomes, or causal claims
  • Uses only authorized data and de-identifies sensitive records
  • Keeps exposure, activation, use, outcome, retention, and cost separate
  • Does not compare immature cohorts as if they had equal observation time
  • Returns TOO EARLY when the natural product cadence has not elapsed