HSCI 341 · Lesson 2

Surveillance and Sampling

Fundamental Epidemiological Concepts and Approaches

Learning objectives for this lesson:

  • Distinguish passive, active, sentinel, and syndromic surveillance, identify Canadian examples of each, and trace a notifiable disease report from the clinic to the Public Health Agency of Canada
  • Navigate the major federal and BC surveillance products (CNDSS, FluWatch, CCDSS, BCCDC dashboards, CVSD) and the registry and vital-statistics infrastructure beneath them, and interrogate any source on its timeliness, completeness, representativeness, sensitivity, and predictive value positive
  • Apply the classic CDC 10-step outbreak investigation framework and the Canadian FIORP analogue to a real Canadian outbreak, identifying the surveillance signals that triggered the response
  • Explain why an investigation must sample when the population at risk cannot be enumerated, and connect the case-control side of the outbreak workflow to formal sampling design
  • Distinguish between a census and a sample and between descriptive and analytic studies, and describe the target, source, and study population hierarchy and the sampling frame
  • Explain sampling error, Type I and Type II errors, and statistical power, and use the central limit theorem to justify inference from a sample to a population
  • Compare non-probability sampling methods (judgement, convenience, purposive) with probability sampling methods (simple random, systematic, stratified, cluster, multistage, targeted) and choose among them for a given question
  • Account for complex sampling designs in analysis through weights, stratification, clustering, and the design effect
  • Compute required sample sizes for common descriptive and analytic objectives, including adjustments for clustering, attrition, and finite populations

This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Dohoo, I. R., Martin, S. W., & Stryhn, H. (2012). Methods in Epidemiologic Research. VER Inc.

Reference

Glossary: Key Terms, People & Concepts

📚 Reference page, available throughout the lesson

This glossary collects the key concepts, people, and ideas you will meet in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.

Surveillance Concepts
Public Health Surveillance Ongoing systematic collection, analysis, and dissemination of health-related data for public health action. Distinguished from research by its operational mandate and rapid feedback loop to action.
Passive Surveillance Surveillance system in which clinicians, labs, or facilities report cases on their own initiative according to legal or professional requirements. Cheap and broad but vulnerable to under-reporting.
Active Surveillance Surveillance staff regularly contact reporters or examine records to actively elicit case reports. More complete and timely than passive systems but resource-intensive; often used during outbreaks.
Sentinel Surveillance High-quality data collected from a selected subset of providers, sites, or populations chosen to represent broader trends. Trades coverage for timeliness and depth.
Syndromic Surveillance Use of pre-diagnostic data (chief complaints, ED visits, school absences, pharmacy sales) to detect outbreaks earlier than confirmed-case reporting allows.
Notifiable (Reportable) Disease A condition that, by law, providers and laboratories must report to public-health authorities (e.g., measles, TB, COVID-19). Lists are maintained provincially in Canada.
Case Definition Standardised criteria (clinical, epidemiologic, laboratory) used to classify suspect, probable, and confirmed cases. Tightening or loosening the definition trades sensitivity against specificity.
Cluster, Outbreak, Epidemic, Pandemic Cluster: aggregation of cases in a place or time. Outbreak: cases above expected (often local). Epidemic: widespread excess (often used interchangeably with outbreak). Pandemic: epidemic across multiple countries or continents.
Index Case (Primary Case) The first case identified in an outbreak investigation (index) or the first case in a transmission chain (primary). They are not always the same person.
Super-Spreader / Super-Spreading Event An individual or event that generates substantially more secondary infections than the population average, typically driven by host, agent, environmental, and behavioural factors converging.
R₀ (Basic Reproduction Number) The expected number of secondary cases produced by a typical infected individual in a fully susceptible population. R₀ > 1 implies epidemic potential (Wikipedia overview).
Rt (Effective Reproduction Number) The average number of secondary cases per infected individual at a specific time, accounting for current immunity and interventions. Rt < 1 indicates a shrinking epidemic.
Attack Rate The proportion of an at-risk population that develops disease during an outbreak (cumulative incidence). Secondary attack rate refers to spread among contacts of a primary case.
Incubation Period Time from exposure to onset of symptoms. Knowing the typical incubation distribution allows back-calculation from onset to likely exposure window.
Generation Time / Serial Interval Generation time: interval between infection of a primary case and infection of a secondary case. Serial interval: interval between symptom onsets in successive cases. Used to estimate Rt.
Herd (Population) Immunity Indirect protection against infection that occurs when a sufficient proportion of the population is immune, reducing the chance that susceptible people contact infectious ones.
Outbreak Investigation Methods
Ten Steps of Outbreak Investigation (CDC) A standardised sequence: prepare, confirm outbreak, verify diagnosis, define and find cases, descriptive epi, hypothesise, test hypotheses, refine hypotheses, implement control, communicate. Steps overlap in practice.
Descriptive Epidemiology (Person, Place, Time) Characterisation of cases by who is affected, where they live or were exposed, and when illness began. The first analytic step in any outbreak investigation.
Epidemic (Epi) Curve A histogram of case onsets over time. Shape suggests transmission pattern: point-source (sharp peak), continuous-source (plateau), propagated/person-to-person (successive peaks at intervals of one incubation period).
Line List A structured table with one row per case capturing demographic, clinical, exposure, and outcome variables. The operational data backbone of outbreak investigation.
Contact Tracing Identification, notification, and follow-up of people exposed to a confirmed case to enable testing, isolation, prophylaxis, or vaccination. Foundational for STI, TB, Ebola, COVID-19 control.
Hypothesis-Generating Interview An open-ended structured interview with a small number of cases (typically 5–10) to elicit possible exposures, such as foods, places, and activities, before a focused case-control or cohort analysis is launched.
Retrospective Cohort (Outbreak) When the at-risk population is well defined (a wedding, a cruise, a daycare), exposed and unexposed cohorts can be reconstructed and attack rates compared directly, the design of choice in defined-population outbreaks.
Case-Control (Outbreak) When the at-risk population is unbounded (community), confirmed cases are compared with appropriate controls on possible exposures to identify the source.
Environmental Sampling & Molecular Subtyping Laboratory analysis of suspect food, water, surfaces, or air, often combined with whole-genome sequencing of patient isolates, to confirm the outbreak source and link cases.
PHAC & BCCDC Public Health Agency of Canada (national surveillance and outbreak coordination) and the British Columbia Centre for Disease Control (provincial counterpart). Both publish public dashboards used in this lesson.
Key People
Alexander Langmuir (1910–1993) Founder of the CDC Epidemic Intelligence Service (EIS) in 1951; established modern shoe-leather outbreak investigation as a core public-health discipline (Langmuir biography).
William Farr (1807–1883) First compiler of statistical abstracts at the General Register Office in London; established systematic mortality surveillance and the foundation of routine vital statistics (Farr biography).
John Snow (1813–1858) London physician whose 1854 investigation of the Broad Street cholera outbreak combined a spot map, water-supply comparison, and shoe-leather case interviews, an enduring template for outbreak epidemiology (Snow biography).
Key Concepts & Ideas
Target Population The full population to which study findings are intended to apply, defined by the research question.
Study (Source) Population The accessible subset of the target population from which a sample can actually be drawn (also called the source population).
Sampling Frame The actual list, register, or operational mechanism used to enumerate and reach members of the study population. Coverage error arises when the frame omits or double-counts target members.
Sampling Unit The element selected at each stage of the sampling process; this could be an individual, household, school, or cluster, depending on design.
Census Data collection from every member of a population. Eliminates sampling error but is usually impractical and may suffer from large non-response.
Probability Sampling Sampling design in which every unit in the frame has a known, non-zero probability of selection. Required for valid statistical inference using survey weights.
Non-Probability Sampling Sampling without known selection probabilities, for example convenience, purposive, snowball, or quota. Useful for hard-to-reach groups but inference to a population is limited.
Sampling Error Random variation in estimates that arises because we observe a sample rather than the whole population. Quantified by the standard error and decreases as sample size increases.
Non-Sampling Error Errors arising from sources other than random sampling, such as coverage, non-response, measurement, and processing. Often more consequential than sampling error.
Type I Error (α) Rejecting a true null hypothesis: a false positive. Conventionally controlled at 0.05.
Type II Error (β) Failing to reject a false null hypothesis: a false negative. Power = 1 − β, conventionally targeted at 0.80 or higher.
Statistical Power The probability that a study correctly detects an effect of a given size when one exists. Depends on sample size, effect size, variability, and α.
Design Effect (DEFF) Ratio of the variance under the actual complex design to the variance under simple random sampling of the same size. DEFF > 1 indicates the effective sample size is smaller than n.
Oversampling Deliberately sampling certain subgroups at higher rates than their population share to improve subgroup precision. Requires weighting in analysis.
Non-Response Bias Systematic error that occurs when individuals who do not respond differ in important ways from those who do.
Sampling Designs & Methods
Simple Random Sample (SRS) A probability sample in which every possible sample of size n from the frame has an equal chance of selection. The reference design against which others are compared.
Systematic Sampling Selecting every kth unit from the frame after a random start. Approximates SRS unless periodicity in the list aligns with the sampling interval.
Stratified Sampling Dividing the population into mutually exclusive strata and sampling independently within each. Improves precision when strata differ in the outcome and ensures representation of subgroups.
Cluster Sampling Sampling intact groups (clusters) such as schools or villages instead of individuals. Reduces field cost but inflates variance because units within a cluster are correlated.
Multistage Sampling A design that selects clusters in successive stages (e.g., regions → neighbourhoods → households → people). Common in national surveys.
Sampling Weight The reciprocal of the selection probability (often adjusted for non-response and post-stratification). Each respondent represents 1/weight people in the population.
Convenience Sampling Recruiting whoever is most accessible. Fast and cheap but vulnerable to selection bias; common in pilot studies.
Purposive (Judgement) Sampling Researcher selects units believed to be informative or representative based on theoretical criteria. Common in qualitative work.
Snowball Sampling Existing participants refer additional participants. Useful for hidden or stigmatised populations; respondent-driven sampling (RDS) is a probability-based variant.
Quota Sampling Non-probability sampling that fills predefined quotas (e.g., by age and sex) using convenience recruitment within each cell.
No matching entries. Try a different search term.
Section 1 of 9

Foundations: What Surveillance Is and the Four System Types

⏱ Estimated reading time: 25 minutes
Lesson 2 · Part 1 · HSCI 341

The Phone Call That Drives the Course

A duty epidemiologist gets a call about four children with bloody diarrhoea from one school.

Section 1 of 5

Foundations: What Surveillance Is and the Four System Types

The definition, the action loop, and the four system types with Canadian examples.

Langmuir 1963

What surveillance is, and is not

The ongoing, systematic collection, analysis, interpretation, and dissemination of health-related data for public-health action.Langmuir, 1963; refined by Thacker & Berkelman, 1988

The key distinction: surveillance that does not feed back into decisions, alerts, or programs is bookkeeping, not surveillance.

The loop

The surveillance action loop

Data Sources Data Systems Analysis & Interpretation Action closes the loop
Four types

Passive, active, sentinel, syndromic

Passive

Clinicians and labs report on their own initiative. Cheap and broad; vulnerable to non-random under-reporting.

Active

Public-health staff elicit reports. Higher quality; resource-intensive and narrow.

Sentinel

A designated network reports systematically. Trades coverage for depth and data quality.

Syndromic

Pre-diagnostic signals flag clusters early. High sensitivity; low specificity.

Canadian examples

Matching types to real systems

Passive

CNDSS: 50 notifiable conditions, weekly provincial aggregates to PHAC.

Active

FluWatch active: 100-clinician search component; COVID-19 contact tracing.

Sentinel

FluWatch sentinel: ~150 GPs reporting weekly; Canadian Paediatric Surveillance Program.

Syndromic

BC ED monitoring; PHAC wastewater dashboards (SARS-CoV-2, influenza, mpox).

Carry forward

What to take into the next section

  • Surveillance closes the loop: collection without feedback to action is bookkeeping.
  • The four types trade coverage, cost, depth, and timeliness; modern products are hybrids.
  • Passive under-reporting is non-random: severe and urban cases are over-represented.

Introduction and Overview

You are the duty epidemiologist at a regional health authority on a Tuesday afternoon. Three minutes ago, your phone buzzed: a paediatrician at a community clinic just called to say she has seen four children from the same school presenting with bloody diarrhoea over two days. She is requesting stool cultures and wants to know whether you have seen anything similar elsewhere. By the time you put the phone down, you will need to know, quickly, whether four cases is unusual for this organism in this catchment, whether a notifiable disease report has already been filed, who else needs to be looped in, and what data sources you can pull within the next hour. That sequence of questions is what this lesson is about. The infrastructure that lets you answer them is called public-health surveillance, and the structured response that follows is called outbreak investigation.

Learning Objectives

  • State Langmuir's working definition of surveillance and explain what it means to "close the loop."
  • List the five core purposes of surveillance and the four conventional system types (passive, active, sentinel, syndromic).
  • Distinguish strengths, weaknesses, and Canadian examples of each system type.
  • Trace the notifiable-disease reporting flow from clinician to MOH to province to PHAC, and explain why under-reporting in passive systems is non-random.

2.1 What Public-Health Surveillance Is, and What It Is For

The classic working definition, attributed to Langmuir (1963) and refined by Thacker & Berkelman (1988), is the ongoing, systematic collection, analysis, interpretation, and dissemination of health-related data for the planning, implementation, and evaluation of public-health practice. Each of those verbs is doing real work. Ongoing distinguishes surveillance from a one-time study. Systematic rules out anecdote. Analysis and interpretation rule out a passive data warehouse. Dissemination for action rules out research that is never returned to the people who can act on it. A famous shorthand attributed to Langmuir (EIS founder) is that surveillance only counts if it “closes the loop”: information must come back out as decisions, alerts, or programs, otherwise the system is just bookkeeping.

The five purposes of surveillance

When you read a published surveillance report, the authors are usually doing one or more of these five things: (1) detect outbreaks, clusters, and unusual events early; (2) characterize who is getting sick, where, and why (descriptive epi by person, place, and time); (3) monitor trends in incidence, prevalence, and risk factors over time; (4) evaluate the effect of interventions and programs; and (5) plan resource allocation and policy. A single dataset can serve more than one purpose, but the design tradeoffs are different for each.

2.2 The Surveillance Action Loop

It helps to picture surveillance as a closed loop with four moving parts: data sources (clinical reports, lab results, vital records, administrative data), data systems (the registries, dashboards, and notifiable-disease platforms that ingest and store data), analysis and interpretation (the epidemiologists who turn case counts into rates, anomalies, and narratives), and action (the case finding, control measures, public communications, and policy decisions that flow from the analysis). Each handoff between these stages is also a place where the system can leak: a clinician who never reports a case, a province that does not share data with the federal level, an analyst who does not see a signal in time, or a recommendation that never reaches a decision-maker.

For the rest of this section we focus on the data-system layer because it determines what kinds of signals you can detect at all. The four conventional system types differ on a single axis: who is doing the work of finding cases.

2.3 The Four Surveillance System Types

Most public-health surveillance systems can be sorted into one of four types; the syndromic category is the newest, codified by the CDC framework of Buehler and colleagues (2004). They are not mutually exclusive, since a single disease can be tracked by several at once, but they have different strengths, costs, and biases.

TypeWho initiates the reportStrengthsLimitationsCanadian example
Passive Clinicians and labs report when they encounter a notifiable disease. Cheap, broad coverage, mandated by law, runs continuously. Under-reporting (often substantial), variable timeliness, completeness depends on clinician burden. The Canadian Notifiable Disease Surveillance System (CNDSS), aggregated case counts for ~50 reportable conditions submitted by provinces to PHAC.
Active Public-health staff actively contact providers, labs, or households to find cases. Higher case ascertainment, better data quality, useful in outbreak investigations. Resource-intensive, narrow scope, hard to sustain over time. The 100-clinician active-search component of FluWatch, and contact tracing during COVID-19 case investigations.
Sentinel A small, designated network of providers reports systematically, trading breadth for depth. High data quality, manageable cost, can collect richer data than passive systems. Not population-representative; trends but not absolute counts. FluWatch sentinel practitioners (general practitioners reporting weekly influenza-like-illness rates) and the Canadian Paediatric Surveillance Program (CPSP).
Syndromic Real-time signals (chief-complaint codes, EMS calls, OTC drug sales, school absences) flag clusters before lab confirmation. Fast, and can detect events before a definitive diagnosis. Useful for emerging or rare events. Low specificity (lots of false alarms); validation is hard; needs analytic infrastructure. BC's Acute and Communicable Disease Prevention ED chief-complaint monitoring; PHAC's pandemic-era wastewater surveillance dashboards.

A fifth category, laboratory-based surveillance, is sometimes broken out separately because it sits beside (rather than under) the clinician's desk: provincial public-health labs aggregate isolates from clinical labs, perform serotyping or whole-genome sequencing, and feed the results into both passive and active systems. PulseNet Canada is the best-known example, the network that enabled the 2008 Maple Leaf listeriosis outbreak to be linked across provinces by genome (Gilmour et al., 2010).

Why “passive” under-reporting is not random

A common student misconception is that under-reporting in passive surveillance just makes case counts smaller. It does, but it also biases who is counted. Cases that present to the health system, get tested, and produce a positive lab result are systematically over-represented. The result is that severe cases, urban cases, and cases in well-insured populations are more visible in the data than mild, rural, or uninsured ones. When you read a CNDSS rate, the denominator is the catchment population, but the numerator is selected.

2.4 The Notifiable Disease Reporting Flow in Canada

For passive surveillance to function, every link in the chain has to do its part. The Canadian flow looks like this:

  1. Clinician or laboratory identifies a case of a notifiable disease (e.g., a positive shiga-toxin-producing E. coli culture or a clinical diagnosis of measles). Each province publishes its own list of conditions reportable under public-health legislation.
  2. Local Medical Officer of Health (MOH) or regional health authority receives the report (typically within 24 hours for urgent conditions, longer windows for routine ones). The MOH may immediately initiate case investigation, contact tracing, or public-health control measures.
  3. Provincial / territorial public-health authority (e.g., BCCDC, Public Health Ontario, the Quebec INSPQ) aggregates reports from the regional level, performs initial analysis, and shares anonymized aggregate data with the federal level.
  4. Public Health Agency of Canada (PHAC) publishes nationally aggregated counts in CNDSS and feeds disease-specific products like FluWatch and the CCDSS. PHAC also reports onward under the WHO International Health Regulations (2005) for events of international concern; the post-SARS rationale for that regime is laid out by Heymann & Rodier (2004).

Two features of this flow deserve emphasis. First, public-health legislation is provincial in Canada, so the list of notifiable conditions and reporting timelines differ across provinces. A condition might be urgent-reportable in one province and routine in another. Second, the federal level receives aggregated, de-identified data only; PHAC cannot pull individual records. This federal-provincial division is one reason that the COVID-19 pandemic exposed gaps in real-time data sharing, because the legal architecture was not designed for the latency that a respiratory pandemic demands.

Reflection

You are designing a surveillance system to track post-secondary student mental-health outcomes in BC. Which of the four system types (or which combination) would you choose, and why? What kinds of cases would your design miss, and what would you do about that gap?

Model answerA defensible design uses a mixed surveillance system: (a) a passive system reporting from campus counselling centres (mandatory case reports of presenting concerns) for population denominator and trend, (b) an active sentinel system in a stratified sample of post-secondary institutions where standardised screening (PHQ-9, GAD-7) is administered to every student visit, and (c) periodic syndromic surveillance via EHR data linkage and anonymised counselling-line usage. Cases missed by passive: students who never present (the majority), international students who use community providers, and underserved campuses without counselling. Plug the gap with (i) a probability-sample survey component (annual ABLE survey or CCHS-style oversample on post-secondary students), and (ii) outreach to community providers and Indigenous campus services with mandatory de-identified reporting. Privacy and stigma considerations should drive opt-in / opt-out architecture and Indigenous data sovereignty for relevant subsamples.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A provincial laboratory operates a network of 80 family-medicine clinics across Canada that submit weekly counts of patients presenting with influenza-like illness, alongside the results of throat-swab subtyping. This is best described as:

A small, designated network of providers reporting systematically is the defining feature of sentinel surveillance. The system trades the breadth of passive surveillance for higher data quality and richer per-case information.

Question 2: Which of the following is the strongest reason that passive surveillance under-counts cases?

Passive surveillance can only count cases that enter the health system and get diagnosed. Mild, asymptomatic, and untested cases are systematically invisible, the dominant source of under-ascertainment for most conditions.

Question 3: In Canada, who is legally responsible for maintaining the list of notifiable diseases that clinicians must report?

Public-health legislation is a provincial responsibility under the Constitution Act. Each province and territory maintains its own list of notifiable conditions and reporting timelines. PHAC aggregates the data nationally but does not set the lists.
Section 2 of 9

Canadian Surveillance Products and Data Infrastructure

⏱ Estimated reading time: 25 minutes
Section 2 of 5

Canadian Surveillance Products and Data Infrastructure

Federal and BC systems, long-running data infrastructure, and five dimensions of data quality.

Federal layer

Four PHAC headline products

CNDSS

50+ notifiable conditions; weekly provincial aggregates; best for long-run trends and inter-provincial comparisons.

FluWatch

Hybrid sentinel + lab + outbreak counts; weekly respiratory-virus picture.

CCDSS

Case definitions on admin records; prevalence and incidence of chronic conditions (diabetes, hypertension, dementia).

CVSD

Statistics Canada mortality spine; life expectancy and excess-mortality analyses.

Provincial layer: BC

BCCDC dashboards worth knowing

  • Respiratory pathogens: weekly influenza, RSV, SARS-CoV-2 with regional breakdowns.
  • Enteric pathogens: Salmonella, Campylobacter, STEC, Listeria; updated weekly.
  • Toxic drug supply mortality: BCCDC and BC Coroners Service co-production.
  • BC Cancer Registry: incidence by health-service-delivery area.
  • STI quarterly: gonorrhea, chlamydia, syphilis trends.

Back-end platforms: IRIS (communicable disease) and Panorama (outbreak management and immunization).

Data infrastructure

The long-running layer everything reads from

Vital statistics

Continuous since late 19th century; denominator for life expectancy and mortality surveillance.

Cancer registries

Near-complete case ascertainment; active follow-up; BC Cancer Registry is the provincial source.

Health-admin data

DAD, NACRS, MSP billing. Powers CCDSS and most admin-data studies.

CCHS + wastewater

Survey self-report; wastewater independent of testing behaviour.

Data quality

Five dimensions to interrogate every source

Timeliness
Completeness
Representativeness
Sensitivity
Predictive value positive

Speed vs completeness

Wastewater: days. CNDSS: weeks to months. STIs: large under-count. Cancer registries: near-complete.

Sensitivity vs specificity

Syndromic: catches events early, many false alarms. Passive: misses mild cases, few false alarms.

Carry forward

What to take into the next section

  • Federal and provincial products are complementary: national trends versus case-level resolution.
  • The five quality dimensions determine what a data source can and cannot answer.
  • Always ask which dimension is binding for your question before selecting a source.

Introduction and Overview

An earlier section sketched four types of surveillance system. This section turns to the actual products a Canadian epidemiologist works with. They sit in three layers: federal, provincial/territorial, and the long-running data infrastructure that underpins both. Knowing which product to reach for, and what its limitations are, is most of the practical skill of the job.

Learning Objectives

  • Identify the major PHAC surveillance products (CNDSS, FluWatch, CCDSS, CVSD) and what each is best used for.
  • Describe the provincial layer in BC: BCCDC dashboards, IRIS, and Panorama.
  • Recognize the long-running data infrastructure (vital stats, cancer registries, DAD/NACRS, CCHS, wastewater) that surveillance products read from.
  • Apply the five dimensions of surveillance data quality (timeliness, completeness, representativeness, sensitivity, and predictive value positive) to interrogate any data source.

2.5 The Federal Layer

PHAC operates several headline products. They are aggregated and curated, the agency cannot dispense individual-level data, and they each have a different cadence, scope, and intended audience.

CNDSS: Canadian Notifiable Disease Surveillance System

The flagship passive system. Provinces submit weekly aggregated counts for ~50 nationally notifiable conditions (the list is harmonized but not identical to provincial lists). CNDSS feeds the agency's annual Notifiable Diseases Online reports and the underlying open-data tables on Open.canada.ca. Useful for long-run trends and inter-provincial comparisons; less useful for real-time outbreak detection because of the multi-week lag from clinic to PHAC.

FluWatch: weekly respiratory virus surveillance

A hybrid system: a sentinel network of ~150 family-medicine clinicians reports influenza-like-illness rates each week, lab partners across the country submit subtyped influenza and (post-2020) SARS-CoV-2 results, and provincial outbreak counts feed the national picture. The weekly FluWatch report is what most public-health communicators reach for during respiratory season.

CCDSS: Canadian Chronic Disease Surveillance System

A federal-provincial collaboration that uses validated case definitions on top of health-administrative records (physician billing claims, hospital discharge abstracts) to estimate prevalence and incidence of chronic conditions like diabetes, hypertension, asthma, dementia, and ischaemic heart disease. CCDSS is the largest single source of population-level chronic-disease data in Canada and powers the agency's chronic-disease infobase.

CVSD: Canadian Vital Statistics Death Database

Not strictly a surveillance product but the spine of mortality-based surveillance. Statistics Canada compiles every death registered in Canada, with cause coded to ICD-10. CVSD enables life-expectancy estimates, cause-specific mortality trends, and excess-mortality analyses (most prominently used during COVID-19 to estimate pandemic burden).

Specialty surveillance systems

PHAC also runs targeted systems for HIV (including the Canadian Perinatal Surveillance System branch), tuberculosis (CTBRS), antimicrobial resistance (CIPARS, CARSS), opioid-related harms, and others, alongside Canada's contribution to global digital surveillance, the Global Public Health Intelligence Network (GPHIN), described by Mykhalovskiy & Weir (2006) and complemented internationally by HealthMap (Brownstein, Freifeld, & Madoff, 2009). Each lives at canada.ca/en/public-health and is worth knowing exists.

2.6 The Provincial Layer (with BC examples)

The federal products are valuable for the country-level view, but most case-level work happens provincially. In British Columbia, the BC Centre for Disease Control (BCCDC) is the analytic and operational arm of provincial public health, and most of its products are publicly available.

A short tour of BCCDC dashboards worth bookmarking

  • BCCDC respiratory pathogens dashboards: weekly influenza, RSV, and SARS-CoV-2 surveillance with regional breakdowns.
  • BCCDC enteric pathogens dashboards: Salmonella, Campylobacter, STEC, Listeria; updated weekly.
  • BC Vital Statistics overdose deaths: the unrestricted-toxic-drug-supply mortality reports that the BC Coroners Service and BCCDC co-produce.
  • BC Cancer Registry and surveillance: cancer-incidence dashboards by health-service-delivery area.
  • BC Sexually Transmitted Infection Quarterly: gonorrhea, chlamydia, syphilis, congenital syphilis, and infectious syphilis trends.

Behind these dashboards sit two case-management platforms: IRIS (Integrated Reporting Information System, BCCDC's communicable-disease platform) and Panorama (a multi-province public-health platform used in BC for outbreak management and immunization records).

2.7 The Long-Running Data Infrastructure

Most surveillance products do not generate their own data; they read from infrastructure that exists for other reasons.

  • Vital statistics (births, deaths, marriages, divorces) are the oldest population health data in Canada, with continuous coverage since the late 19th century in some provinces. They are the denominator for life expectancy and the numerator for mortality surveillance.
  • Cancer registries: the Canadian Cancer Registry aggregates provincial cancer registries and is one of the few systems with active follow-up and complete case ascertainment. The BC Cancer Registry is the provincial source.
  • Health-administrative data: the Discharge Abstract Database (DAD, hospital admissions), the National Ambulatory Care Reporting System (NACRS, ED visits), and provincial Medical Services Plan billing claims (MSP in BC). CCDSS and many academic studies are built on top of these.
  • Population health surveys: the Canadian Community Health Survey (CCHS) is the workhorse cross-sectional survey for self-reported health behaviours and outcomes; the Canadian Health Measures Survey (CHMS) layers in physical measurement.
  • Wastewater-based surveillance: a rapidly maturing infrastructure since 2020. PHAC and many provinces now sample wastewater for pathogens (SARS-CoV-2, influenza, mpox) as a non-clinical signal that is independent of who seeks testing.

2.8 Data Quality: Five Dimensions to Question Every Source

Every surveillance product makes tradeoffs across these five dimensions. When you read a public-health report, the report's author has implicitly resolved each one:

  1. Timeliness: how long from event to data product? Wastewater can be days. CNDSS can be months.
  2. Completeness: what fraction of true events are captured? STIs are notoriously incomplete; cancer registries are nearly complete.
  3. Representativeness: do the captured cases reflect the affected population? Sentinel networks usually do not.
  4. Sensitivity: will the system detect a true outbreak when it occurs? Syndromic surveillance is highly sensitive, passive surveillance often is not.
  5. Predictive value positive: when the system flags an event, is it real? Inversely related to sensitivity, especially for syndromic systems.
🔎 Hands-on: Explore a real dashboard

Pick one of the BCCDC dashboards listed in 2.6 (or a comparable PHAC product) and spend ten minutes with it. Locate (a) the most recent week's case count, (b) the underlying case definition, and (c) any explicit data-quality caveats the dashboard publishes. Note the date of the most recent update versus today's date: how big is the lag?

BCCDC disease & system statistics · Health Infobase Canada

Reflection

You are asked by a journalist for the “most accurate” current count of chlamydia cases in BC. CNDSS, the BCCDC STI quarterly, and a recent academic estimate using CCHS self-report all give different numbers. How would you explain the discrepancy without making any of the three systems sound discredited?

Model answerThe three numbers differ because they answer different questions, not because any system is broken. CNDSS counts nationally reported confirmed lab-positive cases, an undercount, because asymptomatic infections never test, and a lag, because reporting goes through public-health pipelines. The BCCDC STI quarterly is closer to real-time at provincial scale but still depends on testing. The CCHS self-report estimate captures population prevalence of self-reported diagnoses, subject to recall and disclosure biases, but it includes people who tested years ago and never re-presented. The honest explanation: each system measures a different definition of "chlamydia cases" (lab-confirmed-this-period vs. ever-diagnosed-by-self-report); the discrepancy is informative about the size of the iceberg below the lab-confirmed tip; for any policy use, name which definition is appropriate before picking a number.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: Which Canadian surveillance product is built on top of physician billing claims and hospital discharge abstracts to estimate the prevalence of conditions like diabetes and hypertension?

The Canadian Chronic Disease Surveillance System (CCDSS) applies validated case definitions to administrative health data (physician claims and hospital abstracts) to produce population-level prevalence and incidence estimates for chronic conditions.

Question 2: A surveillance system has high sensitivity but low predictive value positive. What does this mean in practice?

High sensitivity means most true events trigger an alert; low predictive value positive means many alerts turn out not to be real events. This combination is typical of syndromic surveillance, which is designed to err on the side of catching events early.

Question 3: Which feature distinguishes wastewater-based surveillance from clinical case-based surveillance?

Wastewater surveillance samples the population at the sewershed level regardless of testing behaviour, which makes it complementary to clinical surveillance, especially when test access is uneven or self-tests are not reported.
Section 3 of 9

Defining and Investigating Outbreaks

⏱ Estimated reading time: 30 minutes
Section 3 of 5

Defining and Investigating Outbreaks

Definitions, the CDC 10-step framework, and the Canadian Foodborne Illness Outbreak Response Protocol.

Vocabulary

Cluster, outbreak, epidemic, pandemic

Cluster

Aggregation in time or space. May or may not exceed the expected baseline.

Outbreak

Cluster judged to exceed the baseline by a public-health authority. Partly statistical, partly operational.

Epidemic

Large-scale outbreak, often multi-jurisdictional. Often used interchangeably with outbreak.

Pandemic

Worldwide geographic spread. A geographic claim, not a severity claim.

CDC 10-step framework, part 1

Steps 1 through 6

  • Step 1: Prepare for fieldwork (authority, supplies, liaisons)
  • Step 2: Establish the outbreak exists vs. baseline
  • Step 3: Verify the diagnosis (review charts, talk to clinicians)
  • Step 4: Construct a working case definition
  • Step 5: Find cases; build the line list
  • Step 6: Descriptive epi covering person, place, time and the epi curve

The epidemic curve shape (point source, continuous, or propagated) constrains your hypotheses about the exposure window.

The epi curve and steps 7–10

Reading the curve, then testing hypotheses

Point source Continuous source Propagated Time (days from exposure)
Canadian FIORP

Three federal partners, one protocol

PHAC

Epidemiology and surveillance lead. Operates PulseNet Canada for WGS-based cluster detection.

CFIA

Food-safety investigations, traceback from grocery purchases, and product recalls.

Health Canada

Risk assessment of contaminated products; health-impact guidance.

FIORP defines escalation triggers and the Outbreak Investigation Coordinating Committee for multi-jurisdictional events.

Carry forward

What to take into the next section

  • The CDC 10-step framework overlaps in practice: control begins before the analytic study is done.
  • FIORP divides labour across PHAC, CFIA, and Health Canada by function, not hierarchy.
  • Speed versus accuracy is not just operational: it is an ethical question about who bears the cost of error.

Introduction and Overview

Surveillance is meant to detect signals; outbreak investigation is what you do when a signal turns into a problem. This section gives you the operational vocabulary: what counts as an outbreak, the standard 10-step framework that organizes the response, and the specifically Canadian protocol used for foodborne investigations.

Learning Objectives

  • Distinguish cluster, outbreak, epidemic, and pandemic, and explain why each is partly a statistical and partly an operational judgement.
  • Walk through the CDC 10-step outbreak investigation framework and recognize which steps typically run in parallel.
  • Describe the Canadian Foodborne Illness Outbreak Response Protocol (FIORP) and the roles of PHAC, CFIA, and Health Canada.
  • Articulate the speed-vs-accuracy tension in outbreak response and why equity considerations belong in surveillance design.

2.9 What Counts as an Outbreak?

The textbook definition is the occurrence of more cases of a disease than expected in a given population, place, or time. Each italicized phrase is doing work. “More than expected” presupposes a baseline: the local 5-year average for influenza in the same week, or the seasonal threshold from a regression model. “Population, place, time” says outbreaks are local: 30 cases of Campylobacter across Canada in a week is not unusual; 30 cases at one wedding is.

Three working terms you need to use precisely

  • Cluster: an aggregation of cases in time and/or space that may or may not be statistically unusual. Clusters are flagged by surveillance and triaged for further investigation.
  • Outbreak: a cluster judged to exceed the expected baseline. The threshold is decided by the public-health authority, not by the data alone.
  • Epidemic: a large-scale outbreak, often across multiple jurisdictions; in international usage often synonymous with outbreak.
  • Pandemic: an epidemic with worldwide geographic spread. The WHO declares pandemics; severity is a separate dimension and is not part of the definition.

A common student error is to treat “pandemic” as “severe outbreak”: it is a geographic claim, not a severity claim.

The decision to call something an outbreak is partly statistical and partly operational. Statistically, you can compare current counts to a baseline distribution and apply a threshold (e.g., the upper 95% confidence limit of the 5-year mean). Operationally, an outbreak declaration mobilizes resources, triggers a coordinated response, and may require public communication, and authorities are reasonably cautious about both over- and under-calling.

2.10 The CDC 10-Step Outbreak Investigation Framework

The reference framework most North American epidemiologists learn is the CDC's 10-step process, articulated in the modern era by Reingold (1998) and elaborated in the CDC Field Epidemiology Manual. The steps look orderly on paper but in practice you often work several at once and revisit earlier ones as new information arrives.

▸ INTERACTIVE STORY: LEGIONNAIRES' 1976 Open full screen ↗

Walk through the most famous post-war outbreak investigation, scene by scene. Next ▶ advances.

An 8-scene retelling of the 1976 American Legion convention outbreak in Philadelphia, illustrating the CDC's 10-step framework: detection, case definition, descriptive epi, hypothesis generation, analytic study, environmental investigation, agent identification (a brand-new bacterium), and control measures.

Step 1: Prepare for fieldwork

Before you arrive, you confirm authority and roles, brief on the suspected etiology, gather supplies (case-report forms, lab kits, PPE), and identify local liaisons (MOH, environmental health, lab). Preparation is the step new investigators most often skip and most often regret.

Step 2: Establish the existence of an outbreak

Compare current counts to a baseline. If the baseline is unstable (small denominators, seasonal variation), state how you constructed it. Ruling out artefact, whether a new lab test, a reporting policy change, or a clinician on a reporting kick, is part of this step.

Step 3: Verify the diagnosis

Talk to clinicians, review charts, check that lab results are correctly attributed. A pseudo-outbreak driven by a contaminated lab reagent or a misclassification is not unheard of.

Step 4: Construct a working case definition

A case definition has three parts: clinical criteria (symptoms, signs, lab tests), person/place/time restrictions (e.g., attendees of the August 12 potluck), and a level of certainty (suspect, probable, confirmed). You will revise it as the investigation evolves, which is normal.

Step 5: Find cases systematically and record information

Active case finding (chart review, asking clinicians, contacting attendees of the implicated event) yields a line list, with one row per case carrying demographic, clinical, exposure, and outcome variables. The line list is the working dataset for everything that follows.

Step 6: Perform descriptive epidemiology

The classic person, place, time triad. The flagship visualization is the epidemic curve (epi curve), a histogram of case counts by date of symptom onset. Its shape (point-source vs propagated) constrains your hypotheses about the exposure window.

Three epidemic-curve shapes: point source, continuous common source, and propagatedPoint sourcesingle sharp peakone brief exposureContinuous sourcebroad plateauongoing exposurePropagatedsuccessive wavesperson-to-person spread
Idealized epidemic-curve shapes. A point source produces a single sharp peak about one incubation period wide; a continuous common source produces a broad plateau; a propagated outbreak produces successive waves spaced roughly one incubation period apart as cases infect others.

Step 7: Develop hypotheses

From the descriptive epi (and from open-ended interviews of cases) you generate plausible exposures: a specific food, a specific water source, a specific event, a specific procedure. Good hypotheses are testable with the data you have or can collect.

Step 8: Evaluate hypotheses with an analytic study

The two workhorse designs in outbreak settings are the retrospective cohort (when you can enumerate everyone who attended an event, e.g., the wedding-guest list) and the case-control study (when you cannot). The 2×2 tables and risk ratios introduced here are formalized later in this course (measures of disease frequency in a later lesson; measures of association in a later lesson); the outbreak workflow gives you an early concrete reason to want them.

Step 9: Refine hypotheses and carry out additional studies

The first analytic pass often points to several plausible exposures. Environmental sampling, traceback investigations (where did the implicated food come from?), and lab characterization (genome-typing of isolates) sharpen the inference.

Step 10: Implement control and prevention; communicate findings

Control measures (recalling a product, closing a venue, prophylaxis, isolation) often happen before Step 8; the precautionary principle does not require you to wait for a p-value. Communication runs throughout: with the public, with affected communities, with policymakers, and through the final outbreak report.

2.11 The Canadian FIORP and Its Multi-Jurisdictional Structure

For foodborne outbreaks specifically, Canada operates under the Foodborne Illness Outbreak Response Protocol (FIORP), summarised by Vik & Hexemer (2014), which formalizes the roles of the federal partners and provincial/territorial public-health authorities. Three federal partners share the load:

  • PHAC: epidemiology and surveillance lead, including PulseNet Canada (the lab network that does whole-genome sequencing of bacterial isolates).
  • Canadian Food Inspection Agency (CFIA): food-safety investigations, traceback, and product recalls.
  • Health Canada: risk assessment of contaminated products and health-impact guidance.

FIORP defines escalation triggers (when a multi-jurisdictional outbreak exists), establishes an Outbreak Investigation Coordinating Committee (OICC) for inter-provincial events, and lays out communication protocols. The 2008 Maple Leaf listeriosis outbreak (the deli-meat-associated Listeria monocytogenes outbreak that killed 22 Canadians) is the case that drove the modern revisions to FIORP and to the supporting infrastructure of PulseNet.

2.12 Real-Time vs Retrospective: A Standing Tension

An outbreak investigation is run under two competing pressures. Speed, since every day of delay can mean more illness, pushes you toward early hypotheses and precautionary control measures. Accuracy, since falsely accusing a food product or a venue has real costs, pulls in the other direction. Experienced investigators learn to act on confident-enough evidence, communicate uncertainty honestly, and revise control measures as data evolve. The skill is partly statistical and partly ethical: who bears the cost of being wrong in either direction is rarely symmetric.

Equity in surveillance and outbreak response

Surveillance systems do not see all populations equally. Data quality, case ascertainment, and willingness to be tested all vary by social position; outbreak investigators are increasingly expected to ask whose communities are over- or under-represented in the line list and how to adjust the response accordingly. The COVID-19 pandemic made this question impossible to ignore in Canada: differential burdens by neighbourhood income, racialized status, and Indigenous identity were visible in surveillance data once collected, and absent when not.

Reflection

Imagine a hospital flags 11 cases of Clostridioides difficile on one ward over a single week. The ward typically sees 1–2 cases per week. As the consulting epidemiologist, walk through the first three CDC steps you would take in the next 24 hours, and identify which of those steps you would do in parallel rather than serially.

Model answerCDC steps 1–3 (in compressed form): (1) confirm the diagnosis and verify the cluster: pull charts to confirm C. difficile by lab criteria, exclude duplicates, and check baseline ward incidence to verify this is genuinely above-threshold; (2) establish a case definition: for example, lab-positive C. diff in a ward-N patient between dates X–Y, with a defined window post-admission to distinguish hospital-acquired from community-onset; (3) begin case-finding and line-listing: review admissions for the period, search for missed cases on adjacent wards, and start a line list capturing dates of admission, antibiotic exposure, procedures, and shared rooms. Steps 1 and 2 happen serially (you cannot define cases without confirming the cluster), but step 3 (case-finding) and immediate infection-control measures (contact precautions, hand-hygiene audits, environmental cleaning) should run in parallel so that risk does not accrue while the line-list is being built.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: A working case definition for an outbreak typically includes:

A working case definition has three parts: clinical criteria, restrictions on person/place/time, and a level of certainty (suspect, probable, confirmed). It is expected to be revised as the investigation matures.

Question 2: Under the Canadian Foodborne Illness Outbreak Response Protocol (FIORP), which federal agency leads the food-safety investigation and product traceback?

Under FIORP, CFIA leads food-safety investigation, traceback, and product recall. PHAC leads epidemiology and surveillance, while Health Canada leads risk assessment.

Question 3: An epidemic curve with a steep rise, a single peak, and a tail roughly equal to the duration of one incubation period is most consistent with which exposure pattern?

A point-source outbreak (everyone exposed in a single brief window, like a wedding meal) produces an epi curve whose shape mirrors the incubation-period distribution: steep rise, single peak, tail roughly one incubation period long.
Section 4 of 9

Case Study and Hands-on Outbreak Investigation

⏱ Estimated reading time: 30 minutes · R activity: 30–45 minutes
Section 4 of 5

Case Study and Hands-on Outbreak Investigation

The 2017–18 romaine lettuce E. coli O157:H7 outbreak, and a line-list analysis in R.

The outbreak

2017–18 E. coli O157:H7, romaine lettuce

Scale

42 confirmed cases · 5 eastern provinces.
25 US cases · 15 states.
CFR: about 2% (1 death among 42 cases).

Three teaching points

1. Genomic surveillance made the cluster visible.
2. Case-case design traded validity for speed.
3. Precautionary advisory issued without a confirmed source.

Steps 2–3: The genomic signal

How PulseNet Canada made the cluster visible

ON QC NB NS NL Before WGS: scattered background noise

WGS, whole-genome sequencing, revealed a shared phylogenetic neighbourhood. Same cases, newly visible as a cluster.

Steps 6–10

Descriptive epi, case-case analysis, precautionary advisory

Epi curve shape: several waves, consistent with a continuous common-source outbreak from an ongoing supply-chain contamination.

Case-case analysis: outbreak isolates vs. historical sporadic STEC cases pointed to romaine lettuce.

Advisory: issued Dec 2017 before a source was confirmed. Outbreak closed Jan 2018.

Outcome

No specific grower or facility was definitively identified. Precautionary action was still the right call.

R activity preview

What the line-list analysis asks you to do

Attack Rate (closed cohort)
\[ \color{#0B7B6B}{\text{AR}} = \frac{\color{#C2410C}{\text{cases among exposed}}}{\color{#6D28D9}{\text{total exposed}}} \]
AR attack ratecases among exposed ill people who ate the foodtotal exposed everyone who ate the food
Risk Ratio
\[ \color{#0B7B6B}{RR} = \frac{\color{#C2410C}{AR_{\text{exposed}}}}{\color{#6D28D9}{AR_{\text{unexposed}}}} \]
RR risk ratioARexposed attack rate among those who ate the foodARunexposed attack rate among those who did not

If two foods both show high RR and many guests ate both, stratify: compute the RR for Food A within strata of Food B. The food whose RR holds across both strata is the source.

Carry forward

What this section connected

  • WGS made the cluster visible; without PulseNet Canada, it would have been background noise.
  • Case-case design trades validity for speed when community controls are hard to recruit quickly.
  • Precautionary action is defensible when waiting for certainty means more cases.
  • The R activity is a preview: later lessons give these methods their formal foundations.

Introduction and Overview

The two preceding sections gave you the institutional and methodological vocabulary. This section makes them concrete: a real Canadian outbreak walked through against the 10-step framework, and a hands-on R activity in which you investigate a simulated outbreak yourself. You will leave this section having previewed the methods that the rest of this course formalizes: the attack rate (the proportion of an at-risk group who fall ill), the 2×2 table, and the risk and rate ratios, all computed on data you have generated and interrogated yourself.

Learning Objectives

  • Walk the 2017–18 romaine lettuce E. coli O157:H7 outbreak against the CDC 10-step framework and identify the role of whole-genome sequencing.
  • Explain when a case-case design is preferred to a traditional case-control design in a fast-moving foodborne investigation.
  • Compute attack rates and risk ratios across multiple suspect foods from a line list, and interpret an epi curve.
  • Identify the validity threats (recall bias, selection bias, confounding by other shared foods) that limit early outbreak inference and articulate what a defensible control action looks like under uncertainty.

2.13 Case Study: The 2017–18 Romaine Lettuce E. coli O157:H7 Outbreak

Between November 2017 and February 2018, PHAC and US CDC investigated a multi-jurisdictional outbreak of E. coli O157:H7 infections that ultimately involved 42 confirmed cases across 5 Canadian provinces (and 25 cases across 15 US states). All five affected Canadian provinces were in eastern Canada (Ontario, Quebec, New Brunswick, Nova Scotia, and Newfoundland and Labrador); western provinces were not involved. One death was reported among the 42 Canadian cases, a case fatality rate of about 2% (roughly 2 in every 100 diagnosed cases died). The outbreak is a useful teaching case because it shows the full FIORP machinery in motion and makes the role of whole-genome sequencing visible.

Step 2–3: Establishing the outbreak and verifying the diagnosis

PulseNet Canada flagged a cluster of E. coli O157:H7 isolates with matching pulsed-field gel electrophoresis (PFGE) patterns, later confirmed by whole-genome sequencing (WGS) to share a tight phylogenetic neighbourhood. The genomic signal is what made this an investigable outbreak rather than a scattered set of unrelated cases; without WGS, the same cases would have been distributed across the routine STEC surveillance baseline.

Steps 4–6: Case definition and descriptive epidemiology

Confirmed cases were Canadian residents with WGS-matched E. coli O157:H7 isolates and symptom onset between mid-November 2017 and the closing date. Probable cases were household contacts of a confirmed case with compatible symptoms. The descriptive analysis (epi curve, geographic distribution, age and sex distribution) showed a polymorphic temporal pattern with several waves, typical of a continuous common-source outbreak driven by an ongoing contaminated supply chain, not a single point exposure.

Steps 7–8: Hypothesis development and the case-case analysis

Standard hypothesis-generating interviews were conducted with confirmed cases using PHAC's Hypothesis Generating Questionnaire (HGQ) for STEC, which lists hundreds of food and exposure variables. Cases reported leafy greens consumption substantially more often than expected based on national consumption patterns. A case-case analysis comparing outbreak cases to historical sporadic STEC cases pointed strongly to romaine lettuce. CFIA conducted parallel traceback investigations from grocery purchases reported by cases.

Steps 9–10: Refinement, control, and communication

Joint PHAC–CFIA–US-CDC coordination produced a public advisory in late December 2017 urging Canadians in the affected provinces to avoid romaine lettuce. The Canadian outbreak was declared over in mid-January 2018. Importantly, despite intensive traceback, no specific grower or facility could be definitively implicated, a sobering reminder that even good investigations sometimes end without the closure of a single confirmed source.

Three teaching points are worth pulling out of this case. First, the outbreak would not have been detected at all without genomic surveillance; the case counts in any one province in any one week looked like background noise. Second, the case-case study design is a pragmatic alternative to a traditional case-control study when controls are hard to recruit during a fast-moving foodborne investigation; you trade some validity for considerable speed. Third, control measures (the public advisory) were issued before the source was definitively confirmed; precaution is a defensible public-health stance when downside asymmetry favours acting early.

2.14 Why Surveillance Comes Early, and What You Will Build On It

Surveillance comes early in this course on purpose: every later lesson depends on having data to design against, sample from, measure with, and reason about. The outbreak workflow you will run below previews methods you have not yet formally met: sampling (a later lesson), questionnaire design for hypothesis-generating interviews (a later lesson), measures of disease frequency such as attack rates and incidence (a later lesson), measures of association including risk ratios and odds ratios (a later lesson), and the validity and confounding vocabulary that arrives in later lessons. Modern surveillance is also being reshaped by big-data and machine-learning approaches reviewed by Mooney & Pejaver (2018). Treat the 2×2 tables and risk ratios you compute in the R activity as a working preview; the formal derivations and validity diagnostics come in the lessons that follow. The point of doing it now is to ground the rest of the course in a concrete operational setting that needs all of it.

R Activity: Investigating a foodborne outbreak from a line list

The companion R script r-activities/HSCI_341_Lesson_2_Surveillance_and_Outbreak_Investigation.R walks through a canonical foodborne-outbreak workflow on a simulated 200-attendee community potluck (line list phaa_outbreak.csv) plus 60 days of daily new-case counts in the surrounding town (phaa_outbreak_curve.csv). The same datasets are revisited later in this course once measures of association and regression tools have been formally introduced.

# PART A -- read the line list and describe the outbreak
ll <- read.csv("phaa_outbreak.csv", stringsAsFactors = FALSE,
               na.strings = c("", "NA"))
mean(ll$ill, na.rm = TRUE)                  # overall attack rate

# Epidemic curve: histogram of symptom-onset days
hist(ll$onset_day, breaks = seq(0.5, 12.5, 1),
     col = "tomato", main = "Days since the potluck",
     xlab = "Day of symptom onset")

# PART B -- 2x2 table for one suspect food (chicken)
tab <- table(Chicken = ll$chicken, Ill = ll$ill)
addmargins(tab)

a <- tab["1","1"];  b <- tab["1","0"]
c <- tab["0","1"];  d <- tab["0","0"]

rr <- (a/(a+b)) / (c/(c+d))                  # relative risk
se <- sqrt(1/a - 1/(a+b) + 1/c - 1/(c+d))
lo <- exp(log(rr) - 1.96*se);  hi <- exp(log(rr) + 1.96*se)
round(c(RR = rr, lo = lo, hi = hi), 2)

chisq.test(tab)                              # or fisher.test(tab) when small

# PART C -- surveillance time series with a 7-day moving average
curve <- read.csv("phaa_outbreak_curve.csv")
curve$ma7 <- stats::filter(curve$new_cases, rep(1/7, 7), sides = 2)
plot(curve$day, curve$new_cases, type = "h", col = "#0B7B6B",
     xlab = "Day", ylab = "New cases", main = "Surveillance time series")
lines(curve$day, curve$ma7, col = "#CC0033", lwd = 2)

# PART D -- attack rates across all suspect foods, ranked
foods <- c("chicken", "potato_salad", "raw_milk_cheese", "green_salad", "punch")
sapply(foods, function(f) {
  ex <- ll[[f]] == 1
  ar_e <- mean(ll$ill[ex], na.rm = TRUE)
  ar_u <- mean(ll$ill[!ex], na.rm = TRUE)
  round(c(ar_exposed = ar_e, ar_unexposed = ar_u, rr = ar_e/ar_u), 2)
})
What you should be able to do after this activity: compute attack rates and risk ratios across multiple suspect foods, draw and interpret an epi curve, identify the food with the strongest association, and articulate the limitations (small numbers, confounding by other shared foods, recall bias) before recommending a control action.

R Reflect on what you just ran

Use the questions below to interpret the actual numbers, table, and plot you produced. Answer in the matching numbered boxes in your R script.

1. What overall attack rate did mean(ll$ill, na.rm = TRUE) return, and how do you interpret that single number for the potluck?

Model answerThe overall attack rate sits around 0.40 (40%): 40 of every 100 attendees became ill at the potluck. As a single number it tells you the outbreak is substantial (much larger than background gastroenteritis rates of a few per cent per week) and clusters in time, suggesting a common-source point exposure. Attack rate is the right epi metric for an outbreak because the cohort is closed (everyone there is at risk), the time window is short, and the outcome is detected by active follow-up.

2. From your ranked sapply() output, which food had the largest RR, and was its 95% CI clearly above 1? In one sentence, name the prime suspect.

Model answerThe food with the largest RR is typically the chicken or potato salad in this simulation; assume chicken with RR ≈ 4.0–4.5 and a 95% CI that clearly excludes 1 (e.g., 2.1–7.5). The prime suspect is that food: a fourfold higher attack rate among consumers vs. non-consumers, with a CI that does not span the null. The univariable RR is the first-pass diagnostic; multi-food adjustment (next reflection) refines it.

3. Looking at the surveillance plot, on roughly which day does the 7-day moving average peak? What problem does the moving average solve that the raw daily bars do not?

Model answerThe 7-day moving average peaks around day 8–10 of the surveillance window, smoothing the day-to-day jitter that the raw bars show. Daily counts are noisy because of reporting lags, weekend reporting drop-offs, and small denominators; the moving average filters that high-frequency variation while preserving the underlying trend, making the rising and falling phases of the outbreak visible. It is the standard tool for distinguishing signal from reporting noise in surveillance time series.
Saved.

Reflection

Your ranked attack-rate table from the R activity puts two foods at the top: chicken (RR = 4.1, attack rate 62% vs 15%) and potato salad (RR = 3.4, attack rate 55% vs 16%). Many guests ate both. What additional analytic step would you take to disentangle the two? (Think about stratifying on one food while looking at the other, a technique that this course will formalize as confounding control in a later lesson.) Briefly describe what evidence would convince you that one is the source rather than the other.

Model answerThe right next step is cross-tabulated (stratified) analysis: compute the RR for chicken among potato-salad eaters, the RR for chicken among non-eaters of potato salad, and the symmetric two strata for potato salad. If one food's RR remains elevated within both strata of the other while the other food's RR drops to near 1 when stratified, the persistent food is the source and the other was riding along (confounded by co-consumption). If both ratios remain elevated, joint contamination is plausible. Algebraically this is the Mantel-Haenszel logic that this course develops in a later lesson. For a quick sanity check, look at attack rates within the four joint cells (chicken+salad, chicken only, salad only, neither): the joint cell should not have the highest rate by much if the true source is one food.

Minimum 20 characters required.

✓ Reflection saved
Knowledge Check: this section

Question 1: In the 2017–18 romaine lettuce E. coli O157:H7 outbreak, what was the role of whole-genome sequencing?

WGS through PulseNet Canada showed that isolates from cases across multiple provinces shared a tight phylogenetic neighbourhood, turning what looked like background-rate STEC into a recognizable cluster. WGS does not pinpoint a grower (traceback does) or compute attack rates (epidemiologic study does).

Question 2: Why might investigators use a case-case design (comparing outbreak cases to sporadic background cases of the same organism) rather than a traditional case-control design during a fast-moving foodborne outbreak?

Case-case studies trade some validity for considerable operational speed: you compare outbreak cases to historical sporadic cases of the same organism, avoiding the recruitment burden of healthy controls. Recall bias and confounding are not eliminated.

Question 3: An R-activity attack-rate table shows that 60% of attendees who ate chicken got ill, vs 15% of those who did not. The crude risk ratio is approximately:

Risk ratio = (exposed attack rate) / (unexposed attack rate) = 0.60 / 0.15 = 4.0. A risk ratio of 4 means people who ate the food were four times as likely to fall ill as those who did not. Later lessons formalize this calculation; the outbreak setting gives you an early, concrete reason to want it.
Section 5 of 9

Introduction to Sampling

⏱ Estimated reading time: 22 minutes

Lesson 2 · Part 2 · HSCI 341

The Bridge Between Question and Data

Who enters your study determines validity, precision, and every inference you later draw.

Section 1 of 4

Introduction to Sampling

Population hierarchy, sampling frames, and the probability theory that makes inference from a sample possible.

First distinction

Census vs. sample

Census

Every individual measured. Only measurement error applies.

Sample

A subset measured. Both measurement error and sampling error apply, but a well-designed sample provides nearly the same information at a fraction of the cost.

Canada runs both: the Census of Population (every 5 years) and surveys like the CCHS and CHMS that produce population estimates with sampling error attached.

Two study purposes

Descriptive vs. analytic

Descriptive

Characterises population attributes: frequency, prevalence, average. Requires representative sampling.

Analytic

Estimates the magnitude of an exposure-outcome association. Comparison is the priority.

Three levels

The population hierarchy

Study Sample Source Population Target Population Internal validity External validity

The sampling frame is the list of units in the source population. In Canadian practice it is often the Statistics Canada Address Register, an LFS area frame, or a provincial health insurance registry.

Probability foundations

Random variables, expected value, variance

Random variable

A numerical outcome of a random process: e.g., the number of new TB cases in a health region next week.

Expected value (μ)

The long-run average across many repetitions. The quantity we are usually trying to estimate.

Variance (σ²)

How spread out the values are around the mean. Spread controls the precision of sample-based estimates.

Standard error of the mean
\[ \color{#0B7B6B}{SE} = \frac{\color{#C2410C}{\sigma}}{\sqrt{\color{#6D28D9}{n}}} \]
SE standard error of the meanσ population standard deviationn sample size
Key distributions

Probability distributions in public health

Discrete

Bernoulli(p): single yes/noBinomial(n, p): count of successesPoisson(λ): rare event counts

Continuous

Normal(μ, σ²): symmetric measurementsExponential(λ): skewed waiting timesUniform(a, b): equal probability over a range
The engine of inference

The Central Limit Theorem

Sampling distribution of the mean
\[ \color{#0B7B6B}{\bar{X}} \sim N\!\left(\color{#C2410C}{\mu},\; \frac{\color{#6D28D9}{\sigma^2}}{\color{#1D4ED8}{n}}\right) \text{ for large } n \]
sample meanμ population meanσ² population variancen sample size

No matter how skewed or irregular the underlying population, the distribution of sample means approaches Normal as n grows. This is what permits confidence intervals and hypothesis tests even with non-Normal raw data.

The interactive simulator below this walkthrough lets you watch the theorem in action across six different population shapes.

Carry forward

What to take into the next section

  • Target → source → study is the hierarchy; external and internal validity describe how well each transition holds.
  • The sampling frame is where exclusions and selection bias typically enter, not the questionnaire.
  • The Central Limit Theorem and \( SE = \sigma/\sqrt{n} \) are the foundation of frequentist inference.

Introduction and Overview

An earlier lesson closed with the counterfactual framework: causal inference requires us to compare groups under different exposure conditions. This lesson asks the immediate next question: which people should be in those groups? Sampling sits at the foundation of every study you'll design in this course because the choice of who ends up in the study determines almost everything else: the validity of the results, the generalisability of conclusions, and the bias inventory you spent an earlier course learning to recognise. The four content sections move from broad principles to concrete formulas: this section sets up the population hierarchy, sampling frame, and the probability theory that makes inference from a sample possible; a later section covers types of error and non-probability sampling; a later section details the major probability sampling designs; a later section closes with how to analyse data from complex surveys and how to determine sample size in advance.

Learning Objectives

  • Distinguish between a census and a sample.
  • Contrast descriptive and analytic studies.
  • Describe the hierarchy of populations (target, source, study sample).
  • Explain the concepts of internal and external validity.
  • Define a sampling frame and explain its importance.
  • Describe the foundational ideas of probability theory (random variables, expected value, and variance) that underpin statistical inference from a sample.
  • Recognise common probability distributions (Bernoulli, Binomial, Poisson, Normal, Exponential, Uniform) and identify the public-health phenomena they typically describe.
  • Explain what a sampling distribution is and state the Central Limit Theorem in plain language.

Census vs. Sample

When we conduct research, we need data from either all individuals in a population or a subset of them. The process of obtaining this data is called measurement.

In a census, every individual in the population is evaluated. In a sample, data are collected from only a subset. Sampling is generally more convenient and less costly than conducting a full census. Interestingly, even a census can be viewed as a kind of sample: it captures the population at one point in time, making it a "sample" of the population over time.

Key Distinction

In a census, the only source of error is the measurement itself. With a sample, you contend with both measurement error and sampling error. However, a well-planned sample can provide virtually the same information as a census at a fraction of the cost.

Canadian Examples: Census vs. National Health Surveys

Canada runs both kinds of data collection at the population scale, and you will encounter all of them in public health practice:

  • Census of Population (Statistics Canada, every 5 years; most recent 2021). A near-complete enumeration of every household in Canada. The short-form census goes to all households; the long-form census goes to a 25% mandatory sample. Provides the denominators behind almost every population health rate you will calculate.
  • Canadian Community Health Survey (CCHS): Statistics Canada / Health Canada / PHAC. A continuous cross-sectional sample survey (~65,000 respondents per cycle) covering self-reported health, behaviours, and health-care use. The flagship descriptive survey for population health surveillance.
  • Canadian Health Measures Survey (CHMS): Statistics Canada / Health Canada / PHAC. A multi-stage sample survey that adds direct physical measurements (blood pressure, biomarkers, fitness) and a biobank to self-report data. Smaller (~5,700 respondents per cycle) but anchors objective measurement of population health.
  • National Population Health Survey (NPHS): the longitudinal predecessor (1994–2011) to the CCHS, still used for life-course research.

The Census gives you a denominator and demographic context; CCHS and CHMS give you population estimates of health states with sampling error attached. Choosing among them is the first applied sampling decision a public-health analyst makes.

Descriptive vs. Analytic Studies

Samples support two fundamental types of studies:

Descriptive Studies (Surveys)

A descriptive study aims to describe population attributes such as the frequency of disease or the prevalence of an exposure. Surveys answer questions like: "What proportion of people had diarrhea over a 1-month period?" or "What is the average BMI of students in Grade 12?"

The focus is on characterizing the current state of a population rather than establishing cause-and-effect relationships.

Analytic Studies

An analytic study is designed to estimate the magnitude of an association between exposures and outcomes. These studies contrast groups and seek explanations for differences between them.

Examples: "Is water source associated with the incidence of diarrhea?" or "How does time spent playing video games affect the BMI of Grade 12 students?"

Establishing an association is the first step to inferring causation, as discussed in an earlier lesson.

Choosing between a census and a sample is the first decision; choosing between a descriptive and an analytic purpose is the second. Both decisions sit on top of an even more fundamental concept: the relationship among the different populations a single study touches.

Hierarchy of Populations

Understanding the different populations involved in a study is essential for evaluating validity. There are three key populations to consider, each nested inside the next. The diagram and accordion below define them in turn; the labels that appear next to the arrows (external validity, sampling frame, internal validity) are the technical vocabulary you will use to talk about how a sample inherits or loses information from the broader population it was meant to represent.

Target Population The broadest group to which you want to generalize results External validity Source Population The population from which study subjects are drawn Sampling frame Study Sample The individuals who actually participate and provide data Internal validity

Figure 3.1. Hierarchy of populations in epidemiologic research. The target population is the broadest; the source population is the accessible subset; the study sample consists of those who actually participate.

Target Population

The target population is the population to which you want to extrapolate your results. It is often not clearly defined and may vary depending on the perspective of the person interpreting the study. For example, researchers studying rainwater cisterns in Pernambuco State, Brazil might define the target as that state, while someone else may want to generalize the findings to all semi-arid regions of Brazil.

Source Population

The source population is the population from which study subjects are actually drawn. All units in the source population should be "listable" and have a non-zero probability of being included in the study. For example, in a diarrhea study in Brazil, the source population included families from households participating in the One Million Cisterns Project (OMCP).

Study Sample

The study sample (or study group) consists of the individuals who actually end up in the study. It is typically a subset drawn from the source population. Researchers determine the necessary sample size, draw their sample, collect data from eligible subjects, and the final study sample consists of those who agreed to participate and whose data met quality requirements.

Validity: Internal and External

Internal validity refers to whether the study results are valid for members of the source population. It indicates whether the study obtained the "correct" answer for that population. Much of epidemiology is dedicated to methods that ensure internal validity.

External validity involves a subjective assessment of whether results can be generalized to the broader target population. It is generally easier to generalize results from analytic studies (which evaluate associations) than from descriptive studies (which estimate prevalence).

The Sampling Frame

The sampling frame is the list of all sampling units in the source population. Sampling units are the basic elements that will be sampled (e.g., households, individuals). A complete list of all sampling units is required for drawing a simple random sample, though some other methods do not require such a complete listing.

Example: Brazil Diarrhea Study

In a study of water cisterns and diarrhea in Brazil, a suitable sampling frame was the list of all households eligible for the One Million Cisterns Project. Once households were selected, a separate strategy was used for selecting individuals within each household.

Canadian Examples: Sampling Frames You Will Actually Use

National-scale Canadian surveys rarely have a single tidy list of every person. Instead, they assemble a frame from administrative listings:

  • Statistics Canada Address Register (AR): the dwelling-level frame used by the Census and many StatCan household surveys.
  • Labour Force Survey (LFS) area frame: CCHS draws part of its sample from the LFS area frame (which is itself based on Census enumeration areas).
  • Provincial health insurance registries (e.g., the BC Medical Services Plan client registry), close to a population census of residents and the backbone of administrative-data research at Population Data BC (PopData BC).
  • Disease registries such as the Canadian Cancer Registry or provincial reportable-disease lists serve as case frames for surveillance.

Notice how each frame has different coverage: a registry-based frame misses people without provincial coverage; an LFS area frame excludes residents on First Nations reserves and in institutions. The frame, not the questionnaire, is usually where exclusions and selection bias enter.

The vocabulary of populations and sampling frames is necessary but not sufficient. The deeper question is why we can learn about a whole population from a small sample at all. The answer comes from probability theory.

Probability Theory: Why Sampling Works

Every quantity we estimate from a sample, whether a prevalence, a mean, or a risk ratio, is the value of a random variable. If we drew a different sample tomorrow, the value would be slightly different. Probability theory is the formal language we use to describe how those values vary, and it is what lets us turn a single sample into a defensible statement about a population.

Three concepts do most of the work in introductory biostatistics:

Three Foundational Ideas

1. A random variable is a numerical outcome of a random process. For example, the number of new TB cases reported in a health region next week, or the systolic blood pressure of the next adult who walks into a clinic.

2. The expected value (also called the mean, written μ or E[X]) is the long-run average of a random variable across many repetitions of the random process. It is the value we are usually trying to estimate.

3. The variance (σ2) and its square root, the standard deviation (σ), measure how spread out the random variable is around its mean. Spread, not the mean alone, is what controls how precisely we can estimate things from a sample.

A second idea worth naming is independence. Two observations are independent when knowing the value of one tells you nothing about the value of the other. Independence is the assumption that lets a small random sample stand in for a much larger population, and it is the assumption most often violated in real public-health data (clustered households, repeated measures on the same person, contagion in infectious disease).

Why You Should Care

Whenever you compute a confidence interval, run a hypothesis test, or quote a margin of error, you are doing arithmetic on a probability distribution, usually a Normal distribution, that describes what would happen if you repeated the study many times. If you don’t know what distribution your statistic comes from, you can’t honestly attach uncertainty to it.

Types of Probability Distributions

A probability distribution describes the values a random variable can take and how likely each value is. Distributions split first into discrete (countable outcomes such as 0, 1, 2 cases) and continuous (any value on a range, such as height or blood pressure, time). The handful below are the ones you will encounter again and again in public health.

Discrete distributions

DistributionWhat it modelsPublic-health example
Bernoulli(p) A single yes/no trial with success probability p. Mean = p, variance = p(1−p). Whether one randomly chosen adult currently smokes.
Binomial(n, p) Number of "successes" in n independent Bernoulli trials. Mean = np, variance = np(1−p). Number of smokers in a CCHS sample of 1,000 adults.
Poisson(λ) Number of rare events in a fixed interval of time, area, or person-time. Mean = variance = λ. New cases of measles per week in a public-health unit; ER visits per hour.

Continuous distributions

DistributionWhat it modelsPublic-health example
Uniform(a, b) Every value between a and b is equally likely. The "default ignorance" distribution. Random-digit dialling within an area code; random selection from a list.
Normal(μ, σ2) The classic bell curve: symmetric, with most mass within 2σ of the mean. The default for many continuous biological measurements and, importantly, for sample means (see CLT below). Adult height; systolic blood pressure; standardized test scores.
Exponential(λ) Time between independent events occurring at a constant rate λ. Right-skewed, memoryless. Mean = 1/λ. Time between successive ED arrivals; survival time under a constant hazard.
Log-normal / Right-skewed Variables that are positive and span several orders of magnitude. The log of the variable is approximately Normal. Household income; hospital length of stay; viral loads.

How to Read a Distribution

Every distribution is summarized by two things: a shape (symmetric? right-skewed? bimodal?) and a small set of parameters that control its location and spread. When you see "BMI ~ Normal(27, 42)", read it as: BMI is approximately Normal, centred at 27 kg/m2, with a standard deviation of 4. About 95% of the population falls within 2 SDs of the mean, roughly 19 to 35.

Binomial(20, 0.3) Discrete · symmetric-ish Poisson(3) Discrete · counts of rare events Normal(μ, σ) Continuous · symmetric bell Exponential(λ) Continuous · right-skewed waits

Figure 3.2. Four common distributions you will meet in public-health data. Discrete distributions assign probability to whole-number outcomes (left two); continuous distributions describe a smooth curve over a real-valued measurement (right two).

🔥 Try it Yourself: Distribution Simulator

What you'll do: use the simulator below to play with each distribution's parameters and watch the shape change in real time, then run the Central Limit Theorem demo to see why sample means from any population eventually look Normal. What to take away: distributions describe specific public-health phenomena rather than serving as textbook abstractions, and the CLT is what lets us trust confidence intervals and power calculations even when the underlying data are skewed.

Aim for at least 10–15 minutes of play; this is the kind of intuition you will draw on for every confidence interval and power calculation in the rest of the course. Use the tabs to switch between exploring single distributions, the CLT demonstration, and a side-by-side comparison.

How to use this

Choose any of the six common distributions on the left. Adjust its parameters to see how the shape, mean, and spread respond. Then click "Draw a sample" to take a random draw and watch the empirical histogram converge on the theoretical curve as your sample grows.

Distribution

Number of successes in n independent yes/no trials. The natural model for "how many out of N have the condition?"

Parameters

Sampling

Tip: Bigger samples produce histograms that look more like the theoretical curve. The "law of large numbers" in action.
Theoretical distribution · Binomial(n = 20, p = 0.30)
The probability mass (discrete) or density (continuous) of the chosen distribution. Red dashed line marks the mean.
Theoretical mean (μ)
6.000
Theoretical SD (σ)
2.049
Sample mean (x̄)
Sample SD (s)

About Binomial: n = 20, p = 0.30

Binomial(n, p) counts the number of successes in n independent Bernoulli trials with the same success probability p.

Mean = np = 6.00, SD = √np(1−p) = 2.05.

  • For large n and moderate p, the binomial looks Normal; this is one of the oldest examples of the CLT.
  • Used for sample-size formulas around proportions and prevalence estimates.
  • Assumes independence and constant p; both can fail in clustered data (households, schools).

Central Limit Theorem demonstration

Pick a population shape (the heavily skewed ones show the effect most clearly). Set a sample size n, then click Draw 1,000 sample means. Each sample of size n is summarised by its mean, and we plot those means. Watch the histogram of means converge on a Normal curve as n grows, even when the underlying population is wildly non-Normal.

Population shape

The CLT says it does not matter which one of these you pick; sample means converge on a Normal curve.

Sample size

Each "trial" draws n values from the population and computes their mean.
Compare: Try n = 1 (the means just are the population). Then bump to n = 5, 30, 100. Notice the histogram becomes Normal-shaped and narrower.
Population distribution
This is the underlying "true" distribution we are sampling from. It can be anything: symmetric, skewed, or bimodal.
Population mean (μ)
Population SD (σ)
Mean of sample means
SD of sample means (SE)
Predicted SE = σ/√n
Trials drawn
0
Ready. Draw sample means to see the CLT in action.

Side-by-side: shapes at a glance

Six distributions, plotted on the same axes. Use this view to remember which shape goes with which name, and to compare how parameters change appearance. Hover for tooltips.

Display options

Different distributions can describe similar-looking data, so choosing the right one depends on the underlying mechanism as much as the shape.
Comparison of common distributions (rescaled to relative density)
Each curve is rescaled so peak density is 1, to make shape comparison easier. Shape, not absolute height, is what matters here.

How to choose a distribution in practice

  • Binary outcome (yes/no for one person)? Bernoulli. Count of "yes"s in a fixed sample? Binomial.
  • Counting rare events in time/space (cases per week, ER visits per hour)? Poisson.
  • Continuous, symmetric biological measurement (BP, height, lab values)? Normal, or check first whether the variable is approximately Normal.
  • Time until next event under a constant hazard? Exponential. Survival under varying hazard? Weibull or Gamma (advanced).
  • Strictly positive, right-skewed, multiplicative process (income, length of stay, viral load)? Log-normal; analyze on the log scale.
  • No prior information about likely values? Uniform on a sensible range.

Each distribution above describes the underlying behaviour of a single random variable. But when we draw a sample from a population, the quantity we typically care about is a summary of that sample: a mean, proportion, or rate. To make inferences from such summaries, we need one more layer of theory.

Sampling Distributions and the Central Limit Theorem

So far we have talked about distributions of individuals in a population. Now consider the distribution of a statistic, for example the mean of a random sample of n people. Because each sample of size n would give a slightly different mean, the mean itself has a distribution. We call it the sampling distribution of the mean.

Two facts about that sampling distribution drive almost all of frequentist inference:

The Two Pillars

1. Standard error. If individual observations have standard deviation σ, then sample means of size n have standard deviation σ/√n. This quantity, the SD of a statistic, is called the standard error (SE). Quadrupling your sample size halves the standard error. Keep the two ideas apart: the standard deviation tells you how much individual people differ from one another, while the standard error tells you how much a summary such as the mean would bounce around if you repeated the whole study.

2. The Central Limit Theorem (CLT). For sufficiently large n, the sampling distribution of the mean is approximately Normal, no matter what the underlying population looks like, even if the population is heavily skewed, bimodal, or discrete. "Sufficiently large" is often around n = 30 for moderately skewed distributions, and much smaller for symmetric ones.

The CLT is the hidden engine behind the bell curve that shows up everywhere in statistics. It is the reason a 95% confidence interval can be written as estimate ± 1.96 · SE: the 1.96 comes from the Normal distribution that the CLT promises us applies to the estimate, even when we have no idea what shape the underlying population has. The theorem traces from de Moivre (1733) through Laplace's 1812 binomial approximation to Lyapunov's 1901 general proof (Central limit theorem, Wikipedia).

Population (skewed) e.g., household income draw samples of size n Means at n=2 Still skewed Means at n=10 Approaching Normal Means at n=50 Tightly Normal

Figure 3.3. The Central Limit Theorem in action. The population can be wildly skewed, but as n grows the sampling distribution of the mean becomes increasingly Normal and increasingly narrow (its SD = σ/√n).

Caveat: The CLT Is About Means, Not Individuals

A common student misconception is that “a large enough sample makes the data Normal.” It does not. Income data with 1,000 observations is still right-skewed. What becomes Normal is the distribution of the sample mean across hypothetical repeated samples; that is what we use to construct confidence intervals around the mean. Inference for medians, proportions, ratios, or extreme values relies on the CLT or its analogues in different ways and may need larger n or different methods (bootstrapping, exact methods).

Key Takeaways

  • A census measures everyone; a sample measures a subset; both involve measurement, but samples also introduce sampling error.
  • Descriptive studies characterize populations; analytic studies evaluate associations between exposures and outcomes.
  • The three populations (target, source, study sample) form a hierarchy, each linked to different aspects of study validity.
  • The sampling frame is the list of all units from which the sample is drawn; in Canadian practice this often means a StatCan address register, an LFS area frame, or a provincial health insurance registry.
  • Random variables, expected values, and variances are the language we use to describe how sample-based estimates vary; they are the foundation of every confidence interval and hypothesis test that follows.
  • A small set of probability distributions (Bernoulli, Binomial, Poisson, Normal, Exponential, log-normal) describes the bulk of public-health phenomena you will encounter.
  • The standard error of the mean is σ/√n, and by the Central Limit Theorem the sampling distribution of the mean is approximately Normal for sufficiently large n; this is the engine behind frequentist inference.
Knowledge Check: this section

1. What is the key difference between a census and a sample?

A census collects data from every individual in the population, while a sample collects data from a subset. Samples introduce sampling error but are more practical and cost-effective.

2. An analytic study differs from a descriptive study in that it:

Analytic studies contrast groups and seek to estimate the strength of associations between exposures and outcomes, while descriptive studies aim to characterize population attributes.

3. Internal validity refers to whether:

Internal validity concerns whether the study produced the "correct" answer for the source population. External validity concerns generalizability to the broader target population.

4. Which distribution would best describe the number of new measles cases reported per week in a public-health unit, where cases are rare and arrive roughly independently?

The Poisson distribution describes the number of rare, independent events occurring in a fixed interval of time, area, or person-time: the canonical model for case counts in surveillance.

5. The Central Limit Theorem says that, for sufficiently large n:

The CLT is about the sampling distribution of the mean, not the raw data. Even if the underlying population is heavily skewed, the distribution of the sample mean across many hypothetical samples becomes approximately Normal as n grows.

6. If a measurement has population standard deviation σ = 12, and you draw a random sample of n = 144, what is the standard error of the sample mean?

The standard error of the mean is σ/√n = 12/√144 = 12/12 = 1. Note that quadrupling n halves the SE; precision grows with the square root of sample size, not linearly.

✦ Pass the knowledge check with 100% to continue

Section 6 of 9

Types of Error & Non-Probability Sampling

⏱ Estimated reading time: 12 minutes

Section 2 of 4

Types of Error & Non-Probability Sampling

Type I and Type II error, statistical power, and the designs that abandon random selection.

Two ways to be wrong

Type I and Type II error

Type I error (α)

Reject the null when it is true. A false positive. Conventional threshold: α = 0.05.

Type II error (β)

Fail to reject the null when it is false. A false negative. Usually set at β = 0.20.

Decision matrix
\[ \begin{array}{l|cc} & \text{Effect present} & \text{No effect} \\ \hline \text{Reject } H_0 & \text{Correct} & \color{#C2410C}{\alpha}\text{ error} \\ \text{Fail to reject} & \color{#6D28D9}{\beta}\text{ error} & \text{Correct} \end{array} \]
α error Type I: reject a true null (false positive)β error Type II: fail to reject a false null (false negative)
Power

Power = 1 − β

Statistical power
\[ \color{#0B7B6B}{\text{Power}} = 1 - \color{#6D28D9}{\beta} \]
Power probability of detecting a true effectβ Type II error rate

A study with 80% power (β = 0.20) will detect a true effect of the specified size eight times out of ten.

  • To increase power: increase sample size, reduce measurement error, or target a larger effect size.
  • A negative finding is only informative if power was adequate. Most underpowered studies are inconclusive, not reassuring.
No random selection

Three non-probability designs

Judgement

Investigator selects participants believed to be representative. Subject to unmeasured selection bias.

Convenience

Whoever is easiest to reach. Fast and cheap; rarely generalisable.

Purposive

Deliberate targeting of specific individuals or groups. Common in qualitative research and case series.

Hidden populations

Snowball and hybrid designs

Snowball sampling (Goodman, 1961): participants recruit peers. No sampling frame needed. Used for hidden populations.

Respondent-driven sampling

Heckathorn (1997). Uses structured referral chains to estimate population parameters. Assumes equilibrium across chains.

Time-location sampling

Systematic sampling at known venues. Used in HIV behavioural surveillance. Requires venue attendance to be unrelated to the outcome.

Carry forward

What to take into the next section

  • Type I (α) = false positive; Type II (β) = false negative; power = 1 − β.
  • Non-probability samples are generally inappropriate for descriptive studies: without known selection probabilities, prevalence estimates cannot be generalised.
  • Selection bias enters when the selection mechanism is correlated with the outcome, not merely with exposure.

Introduction and Overview

An earlier section set up the population hierarchy and the probability theory behind sampling. This section turns to two practical consequences. First: when we draw conclusions from a sample, we make systematic types of mistake (Type I and Type II errors), and these mistakes are quantifiable. Second: there are sampling strategies that abandon probability theory altogether (convenience samples, judgement samples, snowball samples) with predictable consequences for inference. Both halves of this section give you the vocabulary to evaluate when those approaches are acceptable and when they are not.

Learning Objectives

  • Explain the two types of statistical error (Type I and Type II).
  • Define the null hypothesis, P-values, and statistical power.
  • Describe three non-probability sampling methods and their limitations.

Types of Error

In any study based on a sample, the variability of the outcome, measurement error, and sample-to-sample variability all affect results. When making inferences based on sample data, they are subject to error. Within hypothesis testing in analytical studies, there are two key types of error:

Table 3.1. Types of Error

Conclusion of AnalysisEffect Truly PresentEffect Truly Absent
Effect present (reject null)CorrectType I (α) error
No effect (accept null)Type II (β) errorCorrect
Type I (α) Error

A Type I error occurs when you conclude that the outcomes in the groups are different (i.e., that an association exists), when in fact they are not. In other words, you falsely reject the null hypothesis. The probability of a Type I error is denoted α.

Statistical tests are aimed at disproving the null hypothesis (that there is no difference between groups). When P ≤ 0.05, we are "reasonably sure" that any detected effect is not due to chance, but there remains a 5% chance of making a Type I error.

Type II (β) Error

A Type II error occurs when you conclude that there is no association between the exposure and outcome, when in fact there is. You fail to reject the null hypothesis when you should have. The probability of a Type II error is denoted β.

Reasons a study might fail to find a real effect include: the exposure truly had no effect, the study design was inappropriate, the sample size was too small (low power), or simply bad luck.

Statistical Power

Power is the probability that you will find a statistically significant difference when a real difference of a defined magnitude exists. Mathematically, power = 1 − β.

For example, if a study has 80% power, it has an 80% chance of detecting a true effect of the specified size. To increase power, you need to increase the sample size. So-called negative findings (failure to find a difference) are less commonly reported in the literature, partly because many studies lack adequate power.

🎲 Interactive: Sample Size & the Law of Large Numbers

What you'll do: pick a population, set a sample size n, draw repeated samples, and watch the distribution of sample means concentrate around the true population mean as n grows. What to take away: the standard error shrinks as 1/√n; this is the formal reason why “more data” means better precision and higher statistical power. The intuition you build here drives every confidence interval and sample-size calculation in the rest of the course.

Population Distribution

The "true" distribution we're sampling from. Red line = true mean (μ).

0.506.50μ=3.50Fair die
Sampling Distribution of the Mean

Each bar is the count of sample means falling in that range. Yellow = most recent sample's mean.

Draw samples to see distribution of means0.506.50Sample mean (x̄)
True mean (μ)
3.50
Mean of sample means
Std. error (observed)
Theoretical SE = σ/√n
0.54
Samples drawn
0
Last sample mean
Try this: set n = 1, draw 100 samples, then bump n to 50 and draw 100 more (after Reset). The spread of the yellow histogram shrinks dramatically, the 1/√n precision gain in action.

The error framework above assumes a probability sample. The next subsection covers what happens when investigators forgo formal random selection altogether, and what that costs them.

Non-Probability Sampling

Samples drawn without an explicit method for determining each individual's probability of selection are known as non-probability samples. Whenever there is no formal process for random selection, the sample should be considered non-probability. Sample selection that is unrelated to the outcome of interest leaves inference intact, but selection that depends on unmeasured determinants of the outcome produces selection bias, a form of specification error formalised by Heckman (1979) and reviewed for hidden populations by Sudman & Kalton (1986). There are three main types:

Click each card to learn more:

Judgement
Sample
Click to learn more
Convenience
Sample
Click to learn more
Purposive
Sample
Click to learn more

Important Limitation

Non-probability samples are generally inappropriate for descriptive studies because you cannot generalize prevalence estimates to the source population without knowing each individual's probability of being included. However, non-probability methods are commonly used in analytical studies where comparing exposure groups is the priority.

Chain-Referral and Hybrid Designs

Snowball sampling, first formalised by Goodman (1961), recruits hidden-population members through peer referrals and is widely used when no sampling frame exists. Two newer hybrids partially recover probability-style inference: respondent-driven sampling (RDS), introduced by Heckathorn (1997) and extended with unbiased estimators by Salganik & Heckathorn (2004); and time-location (venue-based) sampling, applied at national scale for HIV behavioural surveillance by MacKellar and colleagues (2007). Magnani, Sabin, Saidel, & Heckathorn (2005) review when each design is appropriate for hard-to-reach populations.

Key Takeaways

  • Type I (α) error means falsely concluding there is an effect; Type II (β) error means missing a real effect.
  • Power (1 − β) is the probability of detecting a true effect; increasing sample size increases power.
  • Non-probability samples (judgement, convenience, purposive) lack a formal random selection process and are primarily used in analytic studies.
Knowledge Check: this section

1. A Type I (α) error occurs when you:

A Type I error is a "false positive": you reject the null hypothesis and conclude there is an association when there really is not.

2. Statistical power is defined as:

Power = 1 − β, where β is the Type II error rate. A study with 80% power will detect a true effect 80% of the time.

3. Why are non-probability samples generally inappropriate for descriptive studies?

Without knowing the probability of selection for each individual, you cannot reliably estimate population parameters like prevalence or incidence.

✦ Pass the knowledge check with 100% to continue

Section 7 of 9

Probability Sampling Methods

⏱ Estimated reading time: 15 minutes

Section 3 of 4

Probability Sampling Methods

Simple, systematic, stratified, cluster, multistage, and targeted designs, and how to choose among them.

The defining property

What makes a sample a probability sample

Every element in the sampling frame has a known, non-zero probability of selection, assigned through a formal random process.

This is what separates probability sampling from all non-probability designs. The known probability is what allows you to compute unbiased estimates and valid standard errors.

Random does not mean haphazard. A formal, reproducible process is required: computer-generated numbers or a random number table.

The building blocks

Simple random and systematic sampling

Simple random

Equal probability for all. Needs complete list. All standard analyses apply directly.

Systematic

Pick a random start, then every jth unit. Practical for sequential records. Watch for periodicity.

Sampling interval
\[ \color{#0B7B6B}{j} = \frac{\color{#C2410C}{N}}{\color{#6D28D9}{n}} \]
j sampling interval (take every j-th unit)N population sizen desired sample size
Dividing first

Stratified random sampling

Source population Stratum A Stratum B Stratum C Stratum D SRS or systematic within each stratum; Neyman allocation for optimal precision
Sampling groups

Cluster sampling

Primary sampling unit = the cluster (household, class, clinic).

Randomly select clusters; measure everyone inside each selected cluster.

Advantages: only a cluster list is needed; fewer locations to visit.

Limitation: within-cluster similarity inflates variance relative to SRS for the same n.

Selected Cluster Cluster Cluster
Two-stage selection

Multistage sampling

Select primary sampling units (PSUs), then randomly sample individuals within each PSU rather than taking everyone.

CCHS design

Stratify by health region → cluster by dwelling within each region → select one person per dwelling (Kish grid).

Key requirement

Equal selection probabilities require probability-proportional-to-size (PPS) at the PSU stage, or a constant sampling fraction within each PSU.

Carry forward

Targeted sampling and the design decision

Targeted (risk-based) sampling focuses on high-risk strata: efficient for rare diseases but requires prior knowledge of risk factors and cannot yield standard risk ratios without external data.

Decision question

Does your frame support random selection? If yes, choose a probability design. If no, document the non-probability design and its analytic limits.

Complexity cost

Stratification, clustering, and multistage designs require design-corrected analysis in Section 4. The analysis must match the design.

Introduction and Overview

An earlier section closed by walking through what non-probability sampling looks like and why it forfeits inferential validity. This section returns to probability sampling and details the major variants you'll encounter in real surveys. The six tabs below are not interchangeable; each design buys a different combination of cost, complexity, and statistical efficiency. Read the comparison table at the end of the section as a decision aid you can return to whenever you're choosing a design.

Learning Objectives

  • Define a probability sample and explain why random selection is essential.
  • Describe simple random, systematic random, stratified random, cluster, multistage, and targeted (risk-based) sampling methods.
  • Identify the advantages and disadvantages of each method.

What Is a Probability Sample?

A probability sample is one in which every element in the population has a known, non-zero probability of being included. This implies that a formal process of random selection has been applied to the sampling frame. The key advantage is that probability samples allow for valid statistical inferences about the source population.

Random ≠ Haphazard

Random selection uses a formal, reproducible process (e.g., computer-generated random numbers, random number tables); it is not the same as selecting participants haphazardly or arbitrarily.

Types of Probability Sampling

Simple Random Sample

In a simple random sample, every study subject in the source population has an equal probability of being included. A complete list of the source population is required, and a formal random process is used to select individuals.

Example: To study wait times in a hospital emergency room, you need 1,000 records from 13,000 admissions over the past year. You randomly generate 1,000 numbers between 1 and 13,000 and pull those records.

Advantage: Conceptually simple; all standard statistical analyses apply directly.

Limitation: Requires a complete list of the entire source population.

Systematic Random Sample

In a systematic random sample, a complete list is not required; you only need an estimate of the total population and sequential access to individuals. The sampling interval (j) is computed as the population size divided by the desired sample size.

How it works: Randomly pick a starting point between 1 and j, then select every jth subject after that.

Example: To sample 1,000 from 13,000 emergency patients, the sampling interval is 13. Randomly pick a number between 1 and 13 for your starting patient, then select every 13th patient thereafter. So if your random start happens to be 7, you would sample patients 7, 20, 33, 46, and so on down the list.

Caution: Bias may occur if the factor you are studying is related to the sampling interval (e.g., periodic patterns in admissions).

Stratified Random Sample

The population is divided into mutually exclusive strata based on factors likely to affect the outcome. Then, within each stratum, a simple or systematic random sample is chosen. The mathematical foundations of stratification, including the now-standard optimum (Neyman) allocation rule for assigning sample size across strata, were laid out by Neyman (1934) in his landmark Royal Statistical Society paper that effectively founded probability sampling theory.

In proportional stratified sampling, the number sampled from each stratum is proportional to that stratum's share of the total population.

Three key advantages:

  • Ensures all strata are represented in the sample.
  • Can produce more precise overall estimates than a simple random sample because between-strata variation is removed.
  • Allows estimation of stratum-specific outcomes.

Example: If hospital wait times differ between males and females, stratify records by sex and randomly sample within each group.

Cluster Sampling

A cluster is a natural grouping of study subjects with one or more common characteristics (e.g., a household is a cluster of people; a classroom is a cluster of students; a clinic is a cluster of patients).

In cluster sampling, the primary sampling unit (PSU) is the cluster itself, and it is often larger than the unit of concern. Every individual within a selected cluster is included in the sample.

Example: To estimate smoking prevalence among Grade 12 students, randomly select 10 of 47 Grade 12 classes and survey all students in those 10 classes.

Advantage: Easier when getting a list of clusters is simpler than listing all individuals. Often cheaper to visit fewer locations.

Limitation: Individuals within a cluster tend to be more alike, increasing sampling variation for a given sample size compared to SRS.

Important: A sample is only a "cluster sample" if the group is the sampling unit and the individuals within it are the unit of concern. If the group itself is the unit of concern (e.g., "does anyone in the household smoke indoors?"), it is not a cluster sample.

Multistage Sampling

Multistage sampling is similar to cluster sampling, except that after selecting primary sampling units (PSUs), a sample of secondary sampling units (individuals) is drawn within each PSU rather than surveying everyone.

Example: To study smoking among students, first randomly select 10 classes (PSUs), then randomly select 5 students from each class rather than surveying all students in every class. Within-household selection in face-to-face surveys is most often done using a Kish grid, the objective respondent-selection procedure introduced by Kish (1949).

To ensure all individuals have the same probability of being selected, either choose PSUs proportional to their size, or use a constant sampling proportion within each PSU; the latter requires PPS (probability-proportional-to-size) selection at earlier stages.

The number of individuals per cluster (ni) can be optimized by balancing within-cluster and between-cluster variance against the costs of sampling groups versus individuals.

Targeted (Risk-Based) Sampling

Targeted sampling stratifies the source population based on characteristics associated with the probability of disease occurrence, then focuses sampling on strata where disease is most likely to be found.

Individuals are assigned point values based on their probability of having the disease of interest, and sampling proceeds until a predetermined number of points have been sampled. This is an unequal probability sampling strategy; some individuals may even have a zero probability of inclusion.

Advantage: Requires a much smaller sample to detect rare diseases when key risk characteristics can be identified.

Limitation: Key epidemiological parameters (e.g., risk ratios) may not be known for the study population and must be estimated from other evidence.

Comparison of Sampling Methods

MethodRequires Complete List?Key AdvantageKey Limitation
Simple RandomYesSimple; all standard analyses applyNeeds complete population list
SystematicNo (needs estimate)Practical; easy to implementPeriodic bias if factor linked to interval
StratifiedYes (within strata)More precise; ensures representationNeeds to know stratum membership
ClusterList of clusters onlyCheaper; no need to list individualsHigher variance than SRS for same n
MultistageList of PSUs onlyFlexible; cost-effectiveComplex design; needs more subjects
TargetedNo (risk-based)Efficient for rare diseasesNeeds prior knowledge of risk factors

Worked Example: How the CCHS Combines These Methods

The Canadian Community Health Survey illustrates a real multistage probability design in action:

  1. Stratification: the population is first stratified by health region (about 110 health regions across Canada), and a target sample size is allocated to each so that every region produces stable estimates.
  2. Clustering: within each health region, dwellings are sampled from the LFS area frame (groups of dwellings that share a geographic boundary). This is the cluster stage.
  3. Selection within cluster: one person is randomly selected from each chosen household to complete the interview.
  4. Top-up samples: an RDD (random digit dialling) telephone frame fills in coverage for areas where the area frame is sparse.

The result is a probability sample where every Canadian resident has a known, non-zero chance of selection, but the selection probability differs by region, household size, and frame. That is why CCHS data must be analysed with survey weights and bootstrap replicate weights (covered on the next page).

Reflection

Think of a health research question you are interested in. Which sampling method would be most appropriate, and why? What practical constraints (cost, time, available lists) would influence your choice?

Model answerA defensible answer names the question (e.g., prevalence of food insecurity in BC post-secondary students) and matches design to it. Stratified random sampling is the right default: stratify by institution type (research-intensive vs. teaching-focused vs. community college) and within each stratum draw an SRS proportional to enrolment, oversampling smaller strata to ensure precise estimates. Practical constraints: cost favours administrative-data sampling frames over door-to-door; time favours online survey delivery; available lists (institution registrar files) determine the stratification variables that are feasible. If institutional lists are unavailable, fall back to cluster sampling on courses or classes, accept the design effect, and inflate n accordingly. Document refusal rates and run weighted analysis to address differential non-response.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • Probability samples give every element a known, non-zero chance of selection, enabling valid statistical inference.
  • Simple random sampling requires a complete list; systematic sampling needs only sequential access.
  • Stratified sampling improves precision by removing between-strata variation.
    Stratified versus cluster sampling: strata are internally varied and all are sampled; clusters are internally similar and only some are sampledStratifiedsample from every stratumeach stratum internally varied;all strata sampled → gain precisionClustersample whole clusterseach cluster internally similar;only some clusters sampled (outlined)→ lose precision (design effect)
    Stratified versus cluster sampling. Strata are heterogeneous internally and every stratum is sampled, which removes between-stratum variation and improves precision. Clusters are relatively homogeneous internally and only a subset is sampled, so cases within a cluster carry overlapping information and precision falls (the design effect).
  • Cluster and multistage sampling are practical when listing all individuals is impractical, but they require more subjects for the same precision.
  • Targeted sampling is efficient for rare outcomes but requires prior knowledge of risk characteristics.
Knowledge Check: this section

1. What defines a probability sample?

A probability sample ensures every member of the population has a known, non-zero probability of being selected through a formal random process.

2. A key advantage of stratified random sampling over simple random sampling is that it:

By dividing the population into homogeneous strata and sampling within each, the between-strata variation is explicitly removed from the overall estimate, potentially improving precision.

3. In cluster sampling, why is sampling variation typically greater than in simple random sampling for the same sample size?

Individuals within a cluster (e.g., students in the same class) are more similar to each other than to individuals in other clusters, which increases sampling variation compared to drawing individuals independently.

4. Targeted (risk-based) sampling is most useful when:

Targeted sampling focuses on high-risk strata to efficiently detect rare diseases, requiring prior knowledge of risk characteristics and a smaller sample size than other methods.

✦ Complete the reflection and pass the knowledge check with 100% to continue

Section 8 of 9

Analysing Survey Data & Sample Size

⏱ Estimated reading time: 15 minutes

Section 4 of 4

Analysing Survey Data & Sample Size

Design-correct analysis, the design effect, and the formulae for determining sample size in advance.

Three adjustments

Stratification, weights, clustering

Stratification

Combine stratum-specific estimates. Can reduce overall SE when the stratifying variable predicts the outcome.

Sampling weights

Weight = 1 / p(selection). Corrects for unequal selection probabilities. May change both point estimate and SE.

Clustering

Within-cluster similarity inflates variance. Adjust SE at the PSU level. Ignoring it produces SE values that are 30–80% too small for CCHS data.

Horvitz-Thompson weight
\[ \color{#0B7B6B}{w_i} = \frac{1}{\color{#C2410C}{\pi_i}} \]
wi survey weight for unit iπi selection probability of unit i
Quantifying the cost of complexity

The design effect (deff)

Design effect (Kish, 1965)
\[ \color{#0B7B6B}{\text{deff}} = \frac{\color{#C2410C}{\text{Var}_{\text{complex}}}}{\color{#6D28D9}{\text{Var}_{\text{SRS}}}} \]
deff design effectVarcomplex variance under the actual designVarSRS variance under simple random sampling

deff = 1

Complex design is as precise as SRS.

deff > 1

Less precise than SRS. Brazil diarrhea study: deff = 4.43. The variance was 4.43 times the SRS equivalent.

Two corrections

Finite population correction and bootstrap weights

Finite population correction (FPC)
\[ \color{#0B7B6B}{\text{FPC}} = \frac{\color{#C2410C}{N} - \color{#6D28D9}{n}}{\color{#C2410C}{N} - 1} \]
FPC finite population correctionN population sizen sample size

Apply when sampling fraction > 10%, in simple or stratified designs only. Not applicable in multistage sampling.

CCHS / CHMS practice: Statistics Canada releases 500 bootstrap replicate weights (Rao & Wu, 1988). Run your analysis 500 times; use survey or srvyr in R. Ignoring the weights typically makes standard errors 30–80% too small.

Four formulae

Key sample-size formulae

Estimate a proportion
\[ \color{#0B7B6B}{n} = \frac{\color{#C2410C}{Z_{\alpha}^{2}}\, \color{#6D28D9}{p}\, \color{#1D4ED8}{q}}{\color{#BE185D}{L^{2}}} \]
n required sample sizeZα z-value for the confidence levelp expected proportionq 1 − pL margin of error
Estimate a mean
\[ \color{#0B7B6B}{n} = \frac{\color{#C2410C}{Z_{\alpha}^{2}}\, \color{#6D28D9}{\sigma^{2}}}{\color{#BE185D}{L^{2}}} \]
n required sample sizeZα z-value for the confidence levelσ² population varianceL margin of error
Compare two proportions (per group)
\[ \color{#0B7B6B}{n} = \frac{\left[\color{#C2410C}{Z_{\alpha}}\sqrt{2\bar{p}\bar{q}} - \color{#6D28D9}{Z_{\beta}}\sqrt{p_1 q_1 + p_2 q_2}\right]^{2}}{(\color{#1D4ED8}{p_1} - \color{#BE185D}{p_2})^{2}} \]
n sample size per groupZα z-value for significanceZβ z-value for powerp1 proportion in group 1p2 proportion in group 2
Clustering inflates n

Design-effect and FPC adjustments

Clustering adjustment
\[ \color{#0B7B6B}{n'} = \color{#C2410C}{n} \times \bigl[1 + \color{#6D28D9}{\rho}\,(\color{#1D4ED8}{m} - 1)\bigr] \]
n' cluster-adjusted sample sizen unadjusted sample sizeρ intracluster correlationm average cluster size
Finite population correction adjustment
\[ \color{#0B7B6B}{n'} = \frac{1}{\dfrac{1}{\color{#C2410C}{n}} + \dfrac{1}{\color{#6D28D9}{N}}} \]
n' corrected sample sizen uncorrected sample sizeN population size

Brazil example: unadjusted n = 685 per group → after clustering (ρ = 0.45, m = 6) → n = 2,230 per group. More than three times the unadjusted estimate.

Carry forward

What to take into the final section

  • Design-correct analysis requires accounting for stratification, weights, and clustering. Ignoring them produces false precision.
  • deff > 1 means clustering has cost you precision; inflate the required sample size accordingly with \( n' = n \times [1 + \rho(m-1)] \).
  • Sample-size calculations are assumption documents: precision, variance, confidence level, power, plus adjustments for clustering, attrition, and finite populations.

Introduction and Overview

An earlier section walked through the major probability designs. Real surveys (the CCHS being the canonical Canadian example) typically combine several of these: stratification at the top level, clustering within strata, multistage selection within clusters. That combination is precisely what makes the analysis non-trivial: a complex sample design demands a complex analysis. The first half of this section covers how to do that analysis correctly. The second half closes the loop by walking through how to determine sample size before the data are collected.

Learning Objectives

  • Explain how stratification, sampling weights, and clustering affect the analysis of survey data.
  • Define the design effect and the finite population correction.
  • Describe the key factors that determine sample size.
  • Apply basic sample-size formulae for estimating proportions and means.

Analysing Complex Survey Data

When data come from a complex sampling design (involving stratification, weighting, or clustering), the analysis must account for these features. Ignoring them can lead to incorrect point estimates and underestimated standard errors.

Accounting for Stratification

If the population was divided into strata before sampling, this must be reflected in the analysis. Stratification provides stratum-specific estimates and can reduce the standard error of the overall estimate if the stratifying variable is related to the outcome.

However, stratification alone does not change the overall point estimate; it primarily affects precision. The total population size in each stratum must be known to compute appropriate sampling weights.

Sampling Weights

Not all individuals in a probability sample necessarily have the same probability of selection. The sampling weight for each individual is the inverse of their overall selection probability; this inverse-probability weighting underlies the Horvitz-Thompson estimator introduced by Horvitz & Thompson (1952), which produces unbiased totals and means from any probability sample with known inclusion probabilities.

The probability of selection depends on multiple stages. For example, in a household survey:

p(selection) = (n/N) × (m/M)

where n = households in sample, N = households in source population, m = individuals selected per household, and M = total people in that household.

Multiplying the two stages captures a simple idea: your overall chance of being sampled is the chance your household is chosen times the chance you are then picked within it.

The sampling weight = 1/p(selection). This weight reflects how many people in the source population each sampled individual "represents." Incorporating weights may change both the point estimate and the standard error.

Accounting for Clustering

In cluster and multistage sampling, individuals within groups are usually more alike than randomly chosen individuals. This means observations are not independent, and standard errors must be adjusted upward.

The most common approach is to identify the primary sampling unit (PSU) and adjust all standard error calculations for clustering at that level. The technique called variance linearisation is widely used for this purpose and requires a large number of PSUs to be reliable.

The Design Effect (deff)

The design effect (deff) summarizes the overall impact of the sampling plan on precision. It is the ratio of the variance from the complex sampling design to the variance that would have been obtained from a simple random sample of the same size. The concept and the term were coined by Leslie Kish in his classic textbook Survey Sampling (Kish, 1965) and remain the standard summary statistic for complex-design efficiency.

Interpreting the Design Effect

A deff > 1 means the complex design produces less precise (larger variance) estimates than a simple random sample would. For example, in the Brazil diarrhea study, the deff was 4.43, meaning the variance of the incidence estimate was 4.43 times larger than what a simple random sample of the same size would have produced.

Example: Impact of Survey Design on Estimates

Type of AnalysisIncidence EstimateSE
Simple random sample (assumed)0.14620.0061
+ Stratification0.14620.0059
+ Stratification + Weights0.17510.0091
+ Clustering0.14620.0088
All features combined0.17510.0128

Notice how incorporating all features of the sampling plan changes both the point estimate (from 14.62% to 17.51%) and dramatically increases the standard error (from 0.0061 to 0.0128). Ignoring the sampling design would give a misleadingly precise, and potentially incorrect, result.

Canadian Practice: Bootstrap Weights for the CCHS and CHMS

Statistics Canada distributes the CCHS and CHMS with a set of 500 bootstrap replicate weights rather than releasing the underlying cluster identifiers (which would risk re-identification). The rescaling-bootstrap method that produces these weights was developed by Rao & Wu (1988). To get correct standard errors you re-run your analysis 500 times, once with each replicate weight, and combine the results.

Most analysts use survey or srvyr in R, svy commands in Stata, or SAS PROC SURVEY* procedures. If you ignore the bootstrap weights and just analyse the CCHS as if it were a simple random sample, your standard errors will typically be 30–80% too small, and your confidence intervals and p-values become meaningless.

Finite Population Correction (FPC)

When the proportion of the population sampled is relatively large (>10%), precision improves beyond what would be expected from an "infinite" population. The finite population correction adjusts the estimated variance downward:

FPC Formula

FPC = (N − n) / (N − 1)

where N is the population size and n is the sample size. The FPC should not be applied in multistage sampling even if the number of PSUs sampled exceeds 10% of the total PSUs. It is only applicable to descriptive studies using simple or stratified random sampling.

The intuition is straightforward: once you have already measured a large share of the population, little of it is left to be uncertain about, so the estimate is more precise than the standard (infinite-population) formula assumes.

Analysing data correctly is necessary but not sufficient. Just as important is making sure you collected enough data to begin with; an underpowered study cannot be rescued by clever analysis. The remainder of this section covers sample-size calculation.

Sample-Size Determination

Choosing the right sample size involves both statistical and non-statistical considerations. Non-statistical factors include available resources (time, money, personnel) and the nature of the sampling frame. Statistical considerations include:

Precision of the Estimate

The more precise you need your estimate to be, the larger the sample you need. If you want to know diarrhea prevalence within ±5%, you need more subjects than if ±10% is acceptable. Precision is denoted L (the "allowable error" or half the desired confidence interval width).

Expected Variation in the Data

For proportions, variance = p × q (where q = 1 − p). You need a rough estimate of the proportion to calculate the required sample size. For continuous variables like BMI, you need an estimate of the population variance (σ²). One approach: estimate the range that covers 95% of values, divide by 4 to get σ, then square it for σ².

Level of Confidence

The confidence level (typically 95%) determines how sure you want to be that the confidence interval includes the true population value. This is linked to the Z-value: for 95% confidence, Zα = 1.96. Higher confidence requires a larger sample.

Power (for Analytic Studies)

In analytical studies, you also need to specify the desired power (often 80%). Power determines the sample size needed to detect a specific effect size. For 80% power, Zβ = −0.84. Greater power requires a larger sample.

Key Sample-Size Formulae

ObjectiveFormulaVariables
Estimate a proportionn = Zα² × p × q / L²p = expected proportion; L = precision
Estimate a meann = Zα² × σ² / L²σ² = population variance; L = precision
Compare 2 proportionsn = [Zα√(2pq) − Zβ√(p1q1 + p2q2)]² / (p1−p2p = (p1+p2)/2; n = per group
Compare 2 meansn = 2[(Zα−Zβ)² × σ²] / (μ1−μ2σ² = population variance; n = per group
FPC adjustmentn′ = 1 / (1/n + 1/N)n = initial estimate; N = population size
Clustering adjustmentn′ = n × [1 + ρ(m−1)]ρ = intra-class correlation; m = cluster size

Worked Example: Comparing Two Proportions

Suppose you want to determine if rainwater cisterns reduce the monthly risk of diarrhea from 15% to 10%. With 95% confidence and 80% power:

p1 = 0.15, p2 = 0.10, p = 0.125, q = 0.875

Applying the formula yields n = 685 per group, so you would need 1,370 total individuals (685 with cisterns, 685 without).

If the outcome is clustered within households (ρ = 0.45, average household size m = 6), the clustering adjustment increases the requirement to 2,230 per group, more than triple the unadjusted estimate.

That multiplier, 1 + ρ(m - 1) = 3.25, is itself the design effect for this clustered design: because responses within a household are correlated, each additional person in a cluster carries less new information than an independent draw, so a larger overall sample is needed to reach the same precision.

R Activity: Sampling designs, survey weights, and sample size

The companion R script r-activities/HSCI_341_Lesson_3_Sampling.R walks through three blocks: (A) drawing simple random, stratified, and cluster samples in base R; (B) computing weighted prevalence with the survey package; and (C) running sample-size calculations with power.prop.test and power.t.test, then adjusting for clustering via a design effect.

# PART A -- three probability sampling designs from a 1,000-row frame
set.seed(341)
N <- 1000
frame <- data.frame(id = 1:N,
                    province  = sample(c("BC", "AB", "ON", "QC"), N, replace = TRUE),
                    household = sample(1:300, N, replace = TRUE))

srs   <- frame[sample(N, 100), ]                        # simple random
strat <- do.call(rbind, by(frame, frame$province,
                          function(d) d[sample(nrow(d), 25), ])) # stratified
sel_hh <- sample(unique(frame$household), 30)
clust  <- frame[frame$household %in% sel_hh, ]               # cluster

c(SRS = nrow(srs), Stratified = nrow(strat), Cluster = nrow(clust))

# PART B -- design-corrected prevalence with the survey package
library(survey)
dat <- data.frame(province = sample(c("BC","AB","ON","QC"), 2000, replace = TRUE),
                  smoker   = rbinom(2000, 1, 0.18),
                  weight   = runif(2000, 800, 2200))
des <- svydesign(ids = ~1, strata = ~province, weights = ~weight, data = dat)

mean(dat$smoker)                                # naive (unweighted)
svymean(~smoker, design = des)                  # design-corrected
confint(svymean(~smoker, design = des))         # 95% CI

# PART C -- sample-size calculations + design-effect adjustment
power.prop.test(p1 = 0.15, p2 = 0.10,
                power = 0.80, sig.level = 0.05)        # two proportions
power.t.test(delta = 5, sd = 14,
             power = 0.80, sig.level = 0.05)           # two means (SBP)

n_srs <- 685; rho <- 0.45; m <- 6
ceiling(n_srs * (1 + rho*(m - 1)))                # cluster-adjusted n

What you should be able to do after this activity: draw each of the three probability samples, fit a survey design and report a weighted prevalence with its CI, and compute a sample size for two proportions, two means, and a cluster design.

R Reflect on what you just ran

Use the questions below to interpret the actual numbers you produced. Look at your console output before answering.

1. The line c(SRS = ..., Stratified = ..., Cluster = ...) printed three sample sizes. What were the three numbers, and which design gave you the most variable sample size on a re-run? Why is the cluster sample size NOT exactly 100 or 200 here?

Model answerThe three numbers are typically: SRS = 100 (exact), Stratified = 100 (exact, by construction of the strata sizes), Cluster = around 96–120 depending on the seed (varies). Cluster has the most variable sample size on re-run because it samples clusters, not individuals; once you pick a cluster, you take everyone in it, so the total n depends on the actual sizes of the sampled clusters. SRS and stratified explicitly draw fixed individual counts, so their n is deterministic.

2. Compare mean(dat$smoker) with svymean(~smoker, design = des). Were they nearly the same, and why does that make sense given that the weights came from runif(800, 2200) with no relationship to province or smoker?

Model answermean(dat$smoker) and svymean(~smoker, design = des) were nearly identical because the weights drawn from runif(800, 2200) are independent of both province and smoker status. Weights only matter when they are correlated with the variable being estimated (or with selection probability); under random weights with no informative structure, the weighted mean equals the unweighted mean in expectation. The simulation's point: weights fix design-induced bias only when there is design-induced bias to fix.

3. power.prop.test(p1 = 0.15, p2 = 0.10, power = 0.80, ...) returned an n per group, and the cluster adjustment (rho = 0.45, m = 6) multiplied 685 by roughly 3.25 to give ~2,227. In one sentence, what does that ratio tell you about the price of cluster sampling vs. SRS?

Model answerThe cluster sample needs 2,227 / 685 ≈ 3.25 times as many participants as an SRS to achieve the same statistical power. That ratio is the design effect = 1 + (m−1)ρ = 1 + 5(0.45) = 3.25, confirming the formula. In plain terms: when responses within clusters are correlated (ρ = 0.45 is large), each additional person in a cluster adds less new information than a fresh independent draw; you pay for the operational convenience of cluster sampling by needing a much larger overall n.
Saved.

Reflection

Why do you think it is important to account for clustering when determining sample size? What would happen to your study conclusions if you ignored the clustering effect?

Model answerIgnoring clustering treats every observation as an independent draw, but cluster-correlated data violate that assumption; nearby students share the same teacher, the same socioeconomic context, the same outbreak exposure. The result is under-stated standard errors: CIs are too narrow, p-values too small, and you reject the null when you should not. Concretely, a study analyzing 1,000 students from 50 classes as if they were 1,000 independent observations would report SE that is too tight by a factor of √DE ≈ 1.8 (for ρ = 0.45, m = 6). Conclusions would be over-confident: false-positive associations declared as real, replication fails, and policy decisions made on a sandcastle. The fix: cluster-robust standard errors, multilevel/mixed models, or generalised estimating equations, chosen by data structure and inferential target.

Minimum 20 characters required.

✓ Reflection saved

Key Takeaways

  • Complex survey analyses must account for stratification, sampling weights, and clustering to produce correct estimates and valid standard errors.
  • The design effect (deff) quantifies how much less precise a complex design is relative to a simple random sample.
  • Sample size depends on desired precision, expected variance, confidence level, and (for analytic studies) power.
  • Clustering can dramatically increase the required sample size, especially when the intra-class correlation is high.
  • The finite population correction reduces sample size requirements when sampling a large fraction (>10%) of the population.
Knowledge Check: this section

1. What does a design effect (deff) of 4.43 indicate?

The design effect is the ratio of variance from the complex design to variance from a simple random sample. A deff > 1 means less precision (more variance) than SRS.

2. Sampling weights are computed as:

Sampling weights = 1/p(selection). They represent how many individuals in the source population each sampled individual "represents."

3. Which of the following increases the required sample size?

Greater precision means a smaller L (allowable error), which increases the required sample size. Higher confidence levels, greater variance, and clustering also increase sample size requirements.

✦ Complete the reflection and pass the knowledge check with 100% to continue

Section 9 of 9

Final Assessment

⏱ Estimated time: 30 minutes

Bringing It All Together

This module joined two lessons that sit at the front of the course for the same reason: everything that follows depends on having data to work with, and on knowing what that data can and cannot say about the population it came from. The first half set up the institutional and methodological vocabulary of surveillance: the four system types, the Canadian product layer, the five dimensions of data quality, the CDC 10-step framework, FIORP, and a worked outbreak. The duty epidemiologist who answered the phone in the opening scenario was doing the same kind of reasoning the rest of this course formalizes, but under time pressure and inside a multi-jurisdictional architecture of legislation, agencies, and dashboards.

The second half turned to sampling, which is where the outbreak workflow already pointed. When the population at risk cannot be enumerated, an investigation samples; and whenever it samples, the questions of the later sections apply: which population is the target and which is the source, what frame was used, what sampling error and design effect the estimates carry, and how large a sample the question demands. The target → source → study hierarchy, the central limit theorem, Type I and Type II error and power, the probability and non-probability designs, and the sample-size formulae are the machinery that lets a sample stand in for the population it was drawn from.

The final assessment below asks you to integrate across all eight sections: recognizing surveillance system types from descriptions, choosing the right surveillance product for a question, walking an investigation through the framework, defining populations and frames, matching a sampling design to a research question, and reasoning about error, weights, and design effects. What you take away sets up the next stretch of the course: a later lesson on questionnaire design turns the hypothesis-generating interview into a measured instrument, and later lessons formalize the disease-frequency and association measures you computed by hand here.

Key Takeaways from this lesson

  • Surveillance is defined by closing the loop: ongoing data collection only counts as surveillance if analysis and dissemination produce decisions, alerts, or programs.
  • The four system types (passive, active, sentinel, syndromic) trade off coverage, cost, depth, and timeliness; almost every modern surveillance product is a hybrid.
  • Canadian surveillance is multi-layered and provincially anchored, and every source can be interrogated on five dimensions of data quality: timeliness, completeness, representativeness, sensitivity, and predictive value positive.
  • Outbreak investigation is the CDC 10-step framework applied under FIORP's federal/provincial division of labour; it previews, under time pressure, the cohort, case-control, and 2×2 tools that later lessons formalize.
  • Sampling is the bridge between a research question and feasible data collection: the target → source → study hierarchy makes internal and external validity a precise distinction rather than a slogan.
  • The central limit theorem is what makes inference from a sample to a population work; Type I (α), Type II (β), and power (1−β) are design parameters you set deliberately, not after-the-fact diagnostics.
  • Probability designs (simple, systematic, stratified, cluster, multistage) trade off precision, cost, and feasibility; non-probability designs are defensible only for specific purposes; complex designs require weighting and a design effect in analysis.
  • Sample-size calculations are explicit assumption documents: precision or effect size, variance, confidence level, and adjustments for clustering, attrition, and finite populations.

Core Concepts Reviewed

Section 1: Langmuir's working definition of surveillance and the action loop (data → information → action); the five purposes (detect, characterize, monitor, evaluate, plan); the four system types (passive, active, sentinel, syndromic) with Canadian examples; the notifiable-disease reporting flow from clinic to MOH to province to PHAC to WHO.

Section 2: The federal product layer (CNDSS, FluWatch, CCDSS, CVSD, specialty systems); the BC provincial layer (BCCDC dashboards, IRIS, Panorama); long-running data infrastructure (vital statistics, cancer registries, DAD/NACRS, CCHS, wastewater); five dimensions of surveillance data quality (timeliness, completeness, representativeness, sensitivity, predictive value positive).

Section 3: Outbreak vs cluster vs epidemic vs pandemic; the CDC 10-step investigation framework; the Canadian FIORP and its three federal partners (PHAC, CFIA, Health Canada); the standing tension between speed and accuracy; the equity dimension of surveillance.

Section 4: The 2017–18 romaine lettuce E. coli O157:H7 outbreak as a worked example; the role of whole-genome sequencing in cluster detection; the case-case design as a pragmatic alternative; a working preview of the methods the rest of this course will formalize (cohort and case-control designs, 2×2 tables, risk ratios), applied here under time pressure.

Section 5: Census vs. sample; descriptive vs. analytic studies; the target → source → study population hierarchy and its mapping onto internal and external validity; the sampling frame; random variables, probability distributions, sampling distributions, and the central limit theorem.

Section 6: Sampling error and bias; Type I (α) and Type II (β) error and statistical power (1−β); the non-probability designs (judgement, convenience, purposive), snowball and hybrid designs, and when each is defensible.

Section 7: What makes a sample a probability sample; simple random, systematic, stratified, cluster, multistage, and targeted sampling; how each design trades precision against cost and feasibility.

Section 8: Analysing complex survey data with stratification, weights, and clustering; the design effect (deff); the finite population correction and bootstrap weights; sample-size formulae for common objectives, with design-effect and finite-population adjustments.

The final reflection below asks you to step out of method-mode and carry the two halves of the module into one plan. There is no single right answer; the goal is to leave the lesson with an articulated stance on how surveillance signals and sampling design fit together, because the settings you encounter in the rest of this course (and beyond) will keep pushing on it.

Reflection

You are the duty epidemiologist from the opening scenario. The school-cafeteria outbreak is over, and the regional health authority now wants to know how common the same pathogen is across the whole region, in and out of outbreak settings, over the coming year. Describe how you would move from the surveillance and outbreak machinery of the first half of this module to a defensible sampling plan for that question: which surveillance signals you would draw on, how you would define the target, source, and study populations, which sampling design you would use and why, and what would drive your sample-size calculation (including any design effect). Name one limitation of your plan that you would acknowledge in the report.

Model answerStart from surveillance: routine passive reporting (notifiable-disease counts flowing through the province to CNDSS) gives the baseline and the case definition, PulseNet Canada subtyping shows which clusters belong to the same strain, and syndromic or wastewater signals show where under-detection is likely. Passive counts alone under-estimate the burden, so the year-ahead question needs a designed sample rather than a count. Define the target population as all residents of the region over the year; the source population as residents reachable through the chosen frame (for example, households on the provincial health-registry list, or patients attending a sentinel network of clinics); and the study population as those who complete testing and the questionnaire. Because the region has natural clusters (communities, clinics, schools) and travel costs are real, use a stratified multistage cluster design: stratify by health-service area and urban/rural setting, sample clusters with probability proportional to size, then sample households or patients within clusters. Sample size starts from the expected prevalence (a conservative value near the surveillance estimate, or 0.5 if it is unknown), the desired precision (for example a 95% confidence interval of ±3 percentage points), and the confidence level; then inflate by the design effect 1 + ρ(m − 1) for clustering and by the expected non-response, and apply the finite population correction if a small community is sampled almost completely. Analyse with survey weights, stratification, and cluster-robust variance. The main limitation to acknowledge: the frame and the response process both reach some residents more easily than others, so representativeness (one of the five data-quality dimensions from the first half of the module) has to be assessed directly, for instance by comparing respondents with the census and with the surveillance case mix; and a sentinel-clinic frame trades coverage for timeliness and depth exactly as sentinel surveillance does.

Minimum 20 characters required.

✓ Reflection saved

Final Knowledge Assessment

Complete the following 15-question assessment, which draws on all eight sections of the module. A score of 100% is required to complete the lesson. You may retake the assessment as many times as needed.

Final Assessment: Surveillance and Sampling (15 Questions)

Question 1: Which of the following is the strongest argument that surveillance does analytical work beyond “data collection”?

Langmuir's classic point is that surveillance is defined by closing the loop: collection without analysis-and-dissemination-for-action is bookkeeping, not surveillance.

Question 2: A regional health authority sets up a system in which paramedics flag the chief complaint of every 911 call into a real-time dashboard, with anomaly-detection algorithms triggering review when call volume for a syndrome exceeds the baseline. This is best described as:

Real-time, non-clinical, pre-diagnosis signals fed into anomaly detection are the defining features of syndromic surveillance. The tradeoff is high sensitivity, often low specificity.

Question 3: Which Canadian surveillance product is the primary national source of weekly influenza-like-illness rates from a designated network of family-medicine clinicians?

FluWatch combines a sentinel network of clinicians, lab partners, and provincial outbreak counts to produce the weekly national respiratory-virus picture.

Question 4: A surveillance dashboard reports a weekly count of laboratory-confirmed STEC infections in BC. The lag from symptom onset to appearance on the dashboard is roughly 14 days. Which dimension of surveillance data quality is most directly described by this lag?

Timeliness measures the lag from event to data product. The 14-day lag here describes the speed at which information becomes available for action, not who is captured (completeness/representativeness) or whether alerts are real (PVP).

Question 5: Which of the following best distinguishes an outbreak from a cluster?

A cluster is an aggregation of cases that may or may not be unusual. An outbreak is a cluster judged by the public-health authority to exceed the expected baseline; the threshold combines statistical and operational considerations.

Question 6: The shape of an epi curve where cases rise sharply, peak briefly, and fall away over a period roughly equal to one incubation period is most consistent with:

A point-source outbreak (everyone exposed in a single brief window) produces an epi curve whose shape mirrors the incubation-period distribution: sharp rise, brief peak, tail roughly one incubation period long.

Question 7: In a closed-population outbreak (e.g., a wedding), the appropriate analytic study design is typically:

When the population at risk is fully enumerable (everyone who attended the event), the natural design is a retrospective cohort. Case-control is reserved for open-population settings where enumerating the source population is not feasible.

Question 8: Reflecting on the lesson as a whole, which is the most accurate characterization of the relationship between routine surveillance and outbreak investigation?

Surveillance and outbreak investigation are complementary: routine surveillance produces the signals that warrant investigation, and outbreak investigation applies case-control, cohort, and 2×2 methods (formalized in later lessons of this course) in the operational setting where speed and uncertainty are highest.

Question 9: The source population is best described as:

The source population is the accessible population from which study subjects are drawn. All units should have a non-zero probability of being included.

Question 10: External validity refers to:

External validity is a subjective assessment of whether results from the source population can be generalized to the broader target population.

Question 11: A Type II (β) error occurs when:

A Type II error is a "false negative": you accept the null hypothesis and conclude there is no effect when one actually exists.

Question 12: A convenience sample is characterized by:

A convenience sample is chosen because it is easy to obtain, for example by selecting households close to a research centre.

Question 13: A key advantage of stratified random sampling is:

By dividing the population into homogeneous strata and sampling within each, stratified sampling explicitly removes between-strata variation from the overall estimate.

Question 14: In cluster sampling, the primary sampling unit (PSU) is:

In cluster sampling, the PSU is the cluster itself (e.g., household, classroom, clinic), which is typically larger than the unit of concern (the individuals within).

Question 15: The clustering adjustment formula n′ = n[1 + ρ(m−1)] shows that the required sample size increases when:

When ρ is high (individuals within clusters are very similar) and m is large (many individuals per cluster), the correction factor [1 + ρ(m−1)] becomes large, substantially increasing the required sample size.
✦ Complete every section knowledge check and reflection, including the final reflection above, before submitting

Lesson complete

You have successfully completed this lesson: Surveillance and Sampling. Continue to a later lesson: Questionnaire Design.

You have set up the vocabulary that the rest of this course builds on: the four surveillance system types, the major Canadian products, the CDC 10-step outbreak framework and the FIORP machinery, a hands-on line-list analysis, the target → source → study population hierarchy, the central limit theorem, Type I and Type II error and power, the probability and non-probability sampling designs, and the sample-size and design-effect calculations that make a sample defensible.

A later lesson, Questionnaire Design, picks up the next link in the design chain: once you have decided who will be in the study, you have to decide what to ask them and how. Later lessons then formalize measures of disease frequency and association, screening, validity, and confounding.