6.5 Chapter Review
6.5.1 Chapter summary
A data matrix organizes cases and variables. Samples provide statistics that estimate population parameters, but sampling methods determine representativeness. Observational studies establish association, while randomized experiments can support causal conclusions. Mean and standard deviation suit symmetric data; median and IQR resist skewness and outliers. Potential outliers should be investigated in context.
6.5.2 Common mistakes
- Treating an identifier as a measured variable.
- Generalizing from a convenience sample without qualification.
- Claiming causation from an observational association.
- Using the population variance formula for a sample.
- Removing a potential outlier without investigation.
- Reporting mean and standard deviation for strongly skewed data without explanation.
6.5.3 Exercises
- Classify flow rate and pipe material.
Check Your Work
Flow rate is numerical continuous; pipe material is categorical. - Explain why sampling only one neighbourhood may be biased.
One Possible Answer
The neighbourhood may differ from others in pipe age, pressure zone, or water residence time. - Find mean and median of \(3,4,4,5,19\).
Check Your Work
Mean 7; median 4. - Find sample standard deviation of \(2,4,6\).
Check Your Work
2. - Find fences when \(Q_1=20\) and \(Q_3=32\).
Check Your Work
\(IQR=12\); fences 2 and 50. - Choose summaries for strongly right-skewed overflow volumes.
Check Your Work
Median and IQR.