15.2 Statistics refresher
15.2.1 Variables and measurement
Distinguish categorical from numeric variables and nominal from ordered categories. A numeric storage type does not guarantee a quantitative meaning. Member identifiers, postal codes, and category codes should not be averaged simply because they contain digits.
Identify the observational unit, population, sampling or collection process, and time period. These decisions determine which summaries and comparisons are meaningful.
15.2.2 Centre and spread
The mean uses every value and is sensitive to extremes. The median identifies the middle ordered value and is resistant to extreme observations. Standard deviation summarizes typical distance from the mean under its usual interpretation, while the interquartile range describes the middle half of the distribution.
Choose summaries that fit the distribution and decision. Median visit duration may be more representative when a small number of long events create strong skew. Total facility hours may still matter for staffing even when the median is preferred for describing a typical visit.
15.2.3 Counts, proportions, and rates
A count answers “how many?” A proportion answers “what share of the defined total?” A rate relates events to a population, exposure, or time period.
Always show the denominator. A 20 percent cancellation rate could mean 2 of 10 registrations or 2,000 of 10,000. The proportions are equal, but their stability and operational importance differ.
15.2.4 Conditional comparison
Overall patterns can differ from subgroup patterns. Compare retention within acquisition channel, location, membership type, or joining period when those variables are relevant. Do not create so many subgroups that the analysis becomes unstable or disclosive.
15.2.5 Confidence intervals and tests
A confidence interval expresses uncertainty under a statistical procedure and its assumptions. A hypothesis test evaluates compatibility with a stated null model. Neither determines whether the result matters to the organization.
Report effect size, uncertainty, assumptions, and practical consequence. Avoid using “significant” without clarifying whether it means statistical or practical significance.
The OpenIntro-based statistics reference provides deeper review (Diez et al. 2015).
Worked Example: Distinguishing Statistical and Practical Importance
NVRW estimates that one acquisition channel has a twelve-month retention rate 1.2 percentage points higher than another. The confidence interval excludes zero because the dataset is large. However, the higher-retention channel costs substantially more per acquired member. The decision requires effect size, CAC, contribution margin, and uncertainty, not the p-value alone.