Chapter 9 Researching Secondary Data

Secondary data can supply context, denominators, benchmarks, and explanatory variables that a client dataset does not contain. This chapter presents a structured process for identifying information needs, locating credible sources, evaluating compatibility, preserving provenance, and combining evidence without implying more comparability than the sources allow.

Learning outcomes

After completing this chapter, you should be able to:

  • identify secondary evidence that improves a project;
  • translate an information need into a structured search;
  • evaluate authority, relevance, comparability, and reuse conditions;
  • document sources and transformations through a source log; and
  • combine external and organizational data without hiding differences in definition.

Key terms

  • Primary data: Data collected directly for the project or purpose currently being investigated.

  • Secondary data: Data previously collected by another person or organization for a different or broader purpose and reused in the current analysis.

  • Provenance: The documented origin, collection context, ownership, version, and transformation history of data.

  • Comparability: The degree to which values from different sources can be meaningfully compared because their populations, definitions, periods, units, and methods align.

  • Benchmark: A reference value or standard used to evaluate performance, position, or change.

  • Metadata: Information that describes a dataset’s content, structure, definitions, collection methods, quality, ownership, and use.

  • Citation: A formal acknowledgement that identifies the source of data, ideas, methods, or other material.

  • Data lineage: The traceable path from an original source through transformations, joins, calculations, and outputs.