6.5 Chapter Review

6.5.1 Chapter summary

A data matrix organizes cases and variables. Samples provide statistics that estimate population parameters, but sampling methods determine representativeness. Observational studies establish association, while randomized experiments can support causal conclusions. Mean and standard deviation suit symmetric data; median and IQR resist skewness and outliers. Potential outliers should be investigated in context.

6.5.2 Common mistakes

  • Treating an identifier as a measured variable.
  • Generalizing from a convenience sample without qualification.
  • Claiming causation from an observational association.
  • Using the population variance formula for a sample.
  • Removing a potential outlier without investigation.
  • Reporting mean and standard deviation for strongly skewed data without explanation.

6.5.3 Exercises

  1. Classify flow rate and pipe material.
    Check Your Work Flow rate is numerical continuous; pipe material is categorical.
  2. Explain why sampling only one neighbourhood may be biased.
    One Possible Answer The neighbourhood may differ from others in pipe age, pressure zone, or water residence time.
  3. Find mean and median of \(3,4,4,5,19\).
    Check Your Work Mean 7; median 4.
  4. Find sample standard deviation of \(2,4,6\).
    Check Your Work 2.
  5. Find fences when \(Q_1=20\) and \(Q_3=32\).
    Check Your Work \(IQR=12\); fences 2 and 50.
  6. Choose summaries for strongly right-skewed overflow volumes.
    Check Your Work Median and IQR.

6.5.4 Questions for discussion

  1. What practical conditions can make a water-quality sample unrepresentative?
  2. When would a randomized water-treatment experiment be unethical or impractical?
  3. Why should an apparent outlier be investigated before it is removed?