Study Designs
Fundamental Epidemiological Concepts and Approaches
Learning objectives for this lesson:
- Distinguish observational from experimental study designs and descriptive from analytic studies, and identify the direction of inquiry and the appropriate measure of association for each design
- Describe the key features, strengths, and limitations of cross-sectional, cohort, and case-control studies, and select among them for a given research question
- Explain the logic and limitations of ecological studies, define the ecologic fallacy and recognize when it may occur, and describe how systematic reviews synthesize evidence across study types
- Describe the key features of six hybrid study designs (case-crossover, self-controlled case-series, case-case, case-only, case-cohort, and case-case-control), including the logic of using cases as their own controls across time and the choice between unidirectional and bidirectional referent selection
- Describe two-stage sampling designs, explain when they enhance the efficiency of cross-sectional, cohort, and case-control studies, and design the basic sampling strategy for a two-stage case-control study
- Apply hybrid design concepts to research questions involving rare exposures, transient triggers, or expensive covariates
- Design a controlled trial that produces a valid and efficient evaluation of an intervention: state its objectives, specify the target and source populations, and describe the phases of clinical research from Phase 0 through Phase IV
- Allocate subjects using simple, stratified, cross-over, factorial, cluster, and split-plot randomisation; distinguish single, double, and triple blinding and the bias each prevents; and compute sample size requirements, including the inflation factor for cluster randomised trials
- Compare intent-to-treat and per-protocol analyses, define direct, indirect, and total vaccine efficacy, and apply the CONSORT 2010 reporting standards to plan and report a randomised trial
- Match the full portfolio of designs, from cross-sectional studies to randomised trials, to the research questions and practical constraints each is built to address
This course was developed by Dr. Kiffer G. Card, Faculty of Health Sciences, Simon Fraser University based on Dohoo, I. R., Martin, S. W., & Stryhn, H. (2012). Methods in Epidemiologic Research. VER Inc.
Glossary: Key Terms, People & Concepts
📚 Reference page, available throughout the lesson
This glossary collects the key concepts, people, and ideas you will meet in this lesson. Use it as a reference while you work through the material, or as a review before assessments. Type in the search box to filter entries.
Review of Observational Study Designs
⏱ Estimated reading time: 15 minutes
Introduction and Overview
Earlier lessons of this course built the building blocks: causal logic, surveillance, sampling, questionnaire design, frequency, screening, and association. This lesson is a checkpoint that consolidates the observational study designs you first met in an earlier course (earlier lessons) and explicitly maps each design to the measures of association from an earlier lesson. The three content sections walk through this in order: cross-sectional studies (this section), cohort and case-control designs (a later section), and ecological studies plus an introduction to evidence synthesis via systematic reviews (a later section). Once the standard observational designs are consolidated, later lessons will introduce hybrid and controlled designs that combine or extend these basics.
Learning Objectives
- Distinguish between observational and experimental studies.
- Differentiate descriptive from analytic study designs.
- Describe the features and uses of cross-sectional studies.
- Understand the role of study design in the hierarchy of evidence.
Observational vs. Experimental Studies
All epidemiologic studies can be broadly classified into two categories based on whether the investigator assigns the exposure. In experimental studies (also called intervention studies or controlled trials), the researcher deliberately allocates exposure, for example assigning participants to receive a new treatment or a placebo. In observational studies, the researcher simply observes and measures exposures and outcomes as they occur naturally, without intervening.
This distinction matters enormously for causal inference. Experimental designs, particularly randomized controlled trials (RCTs), can establish exchangeability through random assignment. Exchangeability means the exposed and unexposed groups are alike in every other factor that affects the outcome, so a difference in outcomes can be credited to the exposure itself rather than to some background difference between the groups. This strengthens our ability to attribute differences in outcomes to the exposure (Concato, Shah, & Horwitz, 2000). Observational studies, by contrast, must contend with the possibility that exposed and unexposed groups differ in ways that also affect the outcome; that is, they must account for confounding. Vandenbroucke (2008) argues that observational and experimental research serve distinct epistemic purposes, discovery and evaluation, rather than sitting on a single ladder of rigour.
Why Observational Studies Remain Essential
Despite the inferential strengths of experiments, most epidemiologic research is observational. Many important exposures, including smoking, occupational hazards, environmental pollutants, and dietary patterns, cannot ethically be assigned to participants. Others would be impractical to study experimentally because the outcomes are rare or take decades to develop. Observational designs therefore remain the backbone of epidemiologic evidence for most health questions, and Ioannidis (2005) reminds us that even at the top of the hierarchy, study findings depend on power, prior probability, and analytic flexibility, not on design alone.
Descriptive vs. Analytic Studies
Within observational epidemiology, a further distinction exists between descriptive and analytic study designs.
Descriptive Studies
Descriptive studies characterize the distribution of disease or health states in a population. They address questions of who, what, where, and when. Common descriptive designs include case reports, case series, and surveys that estimate disease prevalence or describe the demographic characteristics of affected populations.
Descriptive studies are often the first step in epidemiologic investigation. They generate hypotheses about potential risk factors but do not formally test causal associations. For example, a case series describing an unusual cluster of pneumonia cases among young men in Los Angeles in 1981 was among the first reports that led to the identification of HIV/AIDS.
Analytic Studies
Analytic studies go beyond description to evaluate associations between exposures and outcomes. They incorporate a comparison group (exposed vs. unexposed, cases vs. controls) and use statistical methods to estimate the strength and direction of associations. The major analytic observational designs are cohort studies, case-control studies, and cross-sectional analytic studies.
Analytic studies are designed to test hypotheses. They ask why and how questions: Does smoking increase the risk of lung cancer? Is a high-sodium diet associated with hypertension? By including a comparison group and measuring the magnitude of association, they move us closer to causal inference.
Key Differences at a Glance
Descriptive: No formal comparison group. Generates hypotheses. Focuses on patterns of disease distribution (person, place, time). Examples: case reports, prevalence surveys, ecological studies (when purely descriptive).
Analytic: Includes a comparison group. Tests hypotheses. Estimates measures of association (risk ratios, odds ratios, rate ratios). Examples: cohort studies, case-control studies, cross-sectional analytic studies.
The boundary between these categories is not always rigid. A cross-sectional survey that only reports prevalence is descriptive, but the same survey becomes analytic when it compares prevalence across exposure groups and estimates odds ratios.
Cross-Sectional Studies
A cross-sectional study measures exposure and outcome status simultaneously in a defined population at a single point in time (or over a short period). It provides a "snapshot" of the population, allowing researchers to estimate the prevalence of disease and the prevalence of various exposures.
Key Characteristics
- Timing: Exposure and outcome are assessed at the same time, with no follow-up period.
- Measure of disease frequency: Prevalence (not incidence), because you are measuring how many people currently have the condition.
- Measure of association: The prevalence ratio or prevalence odds ratio. Because exposure and outcome are measured simultaneously, temporal sequence is often unclear.
- Sampling: Participants are typically sampled from a defined population regardless of exposure or disease status.
Example: Cross-Sectional Study of Diabetes and Physical Activity
Researchers survey 5,000 adults in a city, measuring their current level of physical activity and whether they have been diagnosed with type 2 diabetes. They find that 12% of sedentary adults have diabetes compared to 4% of active adults. The prevalence ratio is 12/4 = 3.0, suggesting sedentary adults are three times as likely to currently have diabetes.
However, can we conclude that being sedentary caused diabetes? Not necessarily; some people may have become sedentary because of their diabetes. This is the core limitation of cross-sectional designs: the inability to establish temporal sequence.
Strengths and Limitations
- Relatively quick and inexpensive to conduct.
- Can study multiple exposures and outcomes simultaneously.
- Useful for estimating disease prevalence and planning health services.
- Good for generating hypotheses that can be tested with analytic designs.
- Can be based on existing data sources (e.g., national health surveys).
- Cannot establish temporal sequence: you do not know whether the exposure preceded the outcome.
- Measures prevalence, not incidence, so results are influenced by disease duration (conditions that last longer are over-represented).
- Subject to prevalence-incidence bias (also called Neyman bias or survival bias): people with rapidly fatal or quickly resolved conditions may not be captured.
- Susceptible to confounding, and the temporal ambiguity makes it harder to identify and control for confounders.
Prevalence vs. Incidence: Why It Matters
Cross-sectional studies capture prevalent cases, that is, people who currently have the disease. This pool of cases over-represents conditions that are long-lasting and under-represents conditions that are rapidly fatal or quickly cured. As a result, associations observed in cross-sectional data may not reflect the actual causes of disease onset. For etiologic research, designs that measure incidence (new cases over time), such as cohort studies, are generally preferred.
Key Takeaways
- Observational studies observe naturally occurring exposures; experimental studies assign them.
- Descriptive studies characterize disease distribution; analytic studies test hypotheses about associations.
- Cross-sectional studies provide a snapshot of prevalence at a single point in time.
- The inability to establish temporal sequence is the primary limitation of cross-sectional designs.
- Prevalence-based measures can be distorted by disease duration, making cross-sectional studies less suitable for etiologic inference.
1. What is the key difference between observational and experimental studies?
2. A cross-sectional study measures:
3. The primary limitation of cross-sectional studies for causal inference is:
✦ Pass the knowledge check with 100% to continue
Review of Cohort & Case-Control Studies
⏱ Estimated reading time: 18 minutes
Introduction and Overview
An earlier section reviewed cross-sectional designs, useful for prevalence estimation but limited for causal inference. This section turns to the two analytic designs that can support causal claims: cohort studies (sampling on exposure, following forward to disease) and case-control studies (sampling on disease, looking backward at exposure). Both produce specific measures of association you computed by hand in an earlier lesson.
Learning Objectives
- Describe the design, direction of inquiry, and key features of cohort studies.
- Describe the design, direction of inquiry, and key features of case-control studies.
- Identify the appropriate measures of association for each design.
- Compare the strengths and limitations of both designs.
Cohort Studies
A cohort study begins by identifying a group of individuals (a cohort) who are free of the outcome of interest, classifying them by exposure status, and following them over time to observe whether the outcome develops. The direction of inquiry moves from exposure to outcome, the same temporal direction as causation itself.
Types of Cohort Studies
Prospective (Concurrent) Cohort Study
Participants are enrolled in the present, exposure status is assessed, and they are followed forward in time to observe the development of outcomes. This is the classic cohort design. The Framingham Heart Study, which began enrolling participants in 1948 and continues to follow their descendants, is one of the most famous examples (Mahmood, Levy, Vasan, & Wang, 2014).
Advantage: Exposure is measured before the outcome occurs, minimizing recall bias and clearly establishing temporal sequence.
Disadvantage: Can be very expensive and time-consuming, especially for diseases with long latency periods or low incidence.
Retrospective (Historical) Cohort Study
The investigator uses historical records (e.g., employment records, medical charts, or registry data) to reconstruct a cohort whose exposure status was determined in the past. The outcomes may have already occurred or can be ascertained in the present.
Advantage: Faster and less expensive than a prospective study because the waiting period for outcome development has already elapsed.
Disadvantage: Relies on the quality and completeness of existing records. Important variables may not have been measured or may be recorded inconsistently.
Direction of Inquiry
Figure 8.1. Direction of inquiry in cohort vs. case-control studies. Cohort studies move from exposure to outcome; case-control studies start with outcome and look back at exposure.
Measures of Association in Cohort Studies
Before the formulas, it helps to fix the notation. Nearly every measure in this lesson is read off a 2×2 table that cross-classifies each person by exposure (the rows) and by disease (the columns). The four counts are labelled a, b, c, and d:
| Disease + | Disease − | Row total | |
|---|---|---|---|
| Exposed | a | b | a + b |
| Unexposed | c | d | c + d |
So a is the number of exposed people who develop the disease, and a + b is everyone who was exposed. That makes a/(a+b) the proportion of the exposed who get the disease and c/(c+d) the same proportion among the unexposed. With that picture in mind: because cohort studies follow participants over time and observe new cases, they can directly estimate incidence, which allows the calculation of:
- Risk Ratio (Relative Risk): The ratio of cumulative incidence in the exposed group to cumulative incidence in the unexposed group. RR = (a/(a+b)) / (c/(c+d)).
- Rate Ratio: The ratio of incidence rates (person-time denominators) in exposed vs. unexposed groups.
- Risk Difference (Attributable Risk): The absolute difference in incidence between exposed and unexposed groups.
Why Cohort Studies Are Well Suited to Causal Questions
The ability to measure incidence directly is the central strength of the cohort design. It establishes temporal sequence (exposure precedes outcome), allows calculation of multiple measures of association, and can study multiple outcomes associated with a single exposure. For rare exposures, cohort studies are particularly efficient because you can intentionally over-sample exposed individuals.
Case-Control Studies
A case-control study begins by identifying individuals who have the outcome of interest (cases) and a comparable group who do not (controls). The investigator then looks backward in time to compare the exposure histories of the two groups. The direction of inquiry moves from outcome to exposure.
Selecting Cases and Controls
The validity of a case-control study depends critically on appropriate selection of cases and controls (Wacholder, McLaughlin, Silverman, & Mandel, 1992):
Cases should be clearly defined using consistent diagnostic criteria. They may be incident cases (newly diagnosed) or prevalent cases (existing), though incident cases are preferred to reduce survival bias. Cases are typically identified from hospitals, disease registries, surveillance systems, or population-based sampling frames.
Controls should come from the same source population that gave rise to the cases. They should be people who, had they developed the disease, would have been identified as cases in the study. Common sources include hospital-based controls (patients admitted for other conditions), population-based controls (random samples from the community), and neighbourhood or friend controls. The choice of control group is often the most scrutinized aspect of a case-control study.
Matching involves selecting controls that are similar to cases on specific characteristics (e.g., age, sex, geography). Matching helps control for confounding by these variables and can improve statistical efficiency. However, matched variables can no longer be evaluated as risk factors, and matched analyses require conditional logistic regression.
Measures of Association in Case-Control Studies
Because case-control studies do not follow a defined cohort over time, they cannot directly calculate incidence. Therefore, the risk ratio cannot be computed directly. The deeper reason is that the investigator decides in advance how many cases and how many controls to enrol, so the fraction of the sample that is diseased is fixed by that design choice rather than by how common the disease actually is; a proportion like a/(a+b) therefore does not estimate a genuine risk. Instead, the primary measure of association is the odds ratio (OR):
The Odds Ratio
The odds ratio compares the odds of exposure among cases to the odds of exposure among controls: OR = (a/c) / (b/d) = ad/bc. Under certain conditions, particularly when the disease is rare in the source population (the "rare disease assumption"), the odds ratio approximates the risk ratio. This is why case-control studies are especially useful for studying rare diseases: the rare disease assumption is most likely to hold, and the design efficiently identifies a sufficient number of cases.
Here is the intuition for why rarity matters. The risk of disease in the exposed is a/(a+b), while the odds of disease is a/b. When the disease is rare, very few of the exposed are cases, so a is tiny next to b; that makes a+b almost equal to b, and so the risk a/(a+b) almost equals the odds a/b. The same holds among the unexposed, so the odds ratio and the risk ratio nearly coincide. When the disease is common, a is no longer negligible and the two measures pull apart, with the odds ratio landing farther from 1.
Comparing Cohort and Case-Control Designs
| Feature | Cohort Study | Case-Control Study |
|---|---|---|
| Direction of inquiry | Exposure → Outcome | Outcome → Exposure |
| Starting point | Defined by exposure status | Defined by disease status |
| Measures incidence directly? | Yes | No |
| Primary measure of association | Risk ratio, rate ratio | Odds ratio |
| Best suited for | Rare exposures, multiple outcomes | Rare diseases, multiple exposures |
| Temporal sequence | Clearly established | Relies on retrospective data |
| Cost and time | Often expensive and lengthy | Generally less expensive and faster |
| Key biases | Loss to follow-up, selective attrition | Recall bias, selection bias in controls |
Choosing the Right Design: A Practical Example
Suppose you want to study whether exposure to a specific industrial solvent increases the risk of a rare liver cancer. A prospective cohort study would require following thousands of exposed and unexposed workers for decades, which is extremely expensive and slow. A case-control study, by contrast, could identify 200 liver cancer cases from a cancer registry, select 400 matched controls, and assess past occupational exposure through interviews and employment records, producing results in months rather than years.
For rare diseases, case-control studies are usually the design of choice because they can efficiently identify enough cases to detect meaningful associations.
Key Takeaways
- Cohort studies follow exposed and unexposed groups forward in time to observe outcome development; they can directly calculate incidence and risk ratios.
- Case-control studies compare exposure histories of cases and controls; they estimate the odds ratio, which approximates the risk ratio when the disease is rare.
- Cohort studies are preferred for rare exposures and when multiple outcomes are of interest; case-control studies are preferred for rare diseases and when multiple exposures are of interest.
- Both designs are susceptible to specific biases: cohort studies to loss to follow-up, case-control studies to recall bias and control selection bias.
1. In a cohort study, the direction of inquiry is:
2. The odds ratio is the primary measure of association in case-control studies because:
3. Case-control studies are especially efficient for studying:
✦ Pass the knowledge check with 100% to continue
Ecological Studies & Evidence Synthesis
⏱ Estimated reading time: 15 minutes
Introduction and Overview
Earlier sections covered designs that operate at the individual level. This section makes one final move, up to the group level, with ecological studies, and then steps back to ask how evidence from many such studies can be synthesized. The ecological fallacy you met in an earlier course reappears here as a specific limitation; systematic reviews are the formal answer to “what does the evidence as a whole say?”
Learning Objectives
- Describe the design and purpose of ecological studies.
- Explain the concept of group-level variables and the ecologic fallacy.
- Discuss how systematic reviews synthesize evidence across study types.
- Understand where ecological studies and systematic reviews fit in the hierarchy of evidence.
Ecological Studies
An ecological study (also called a group-level or aggregate study) uses groups, rather than individuals, as the unit of analysis. Instead of measuring exposure and outcome in each person, the investigator compares exposure levels and disease rates across defined populations, such as countries, regions, or time periods.
Types of Ecological Studies
Click each card to learn more:
ComparisonClick to learn more
(Temporal)Click to learn more
DesignClick to learn more
Group-Level Variables
A distinctive feature of ecological studies is their reliance on group-level variables. These can be classified into three types:
Aggregate measures are summaries of individual-level data for a group, for example the mean blood pressure of a country's population, or the percentage of a county's residents who smoke. These are the most common type of group-level variable and are derived from individual measurements.
Environmental measures are physical or chemical characteristics of the environment shared by all members of a group, for example ambient air pollution levels in a city, fluoride concentration in a water supply, or average annual temperature. These are inherently group-level and cannot be measured at the individual level.
Global measures are attributes of the group that have no individual-level analogue, for example population density, the existence of a specific public health law, type of healthcare system, or the Gini coefficient (income inequality). These contextual factors can only be measured at the group level and are particularly interesting because they may represent structural or policy-level determinants of health.
The Ecologic Fallacy
The ecologic fallacy is the most important limitation of ecological studies. It occurs when an association observed at the group level is incorrectly assumed to hold at the individual level (Robinson, 1950/2009; Morgenstern, 1995). Just because countries with higher fat consumption have higher breast cancer rates does not mean that the individuals within those countries who eat more fat are the same individuals who develop breast cancer.
Classic Example: Durkheim and Suicide
Emile Durkheim observed that regions with higher proportions of Protestant residents had higher suicide rates than predominantly Catholic regions. However, this group-level association does not necessarily mean that Protestant individuals were more likely to commit suicide. It is possible that the social environment of predominantly Protestant regions, perhaps characterized by greater individualism or weaker social networks, affected everyone living there, including Catholics. Attributing the group-level finding to individual Protestants would be an ecologic fallacy.
The ecologic fallacy arises because within-group variation is hidden when data are aggregated. Individuals who are exposed and individuals who develop disease may be entirely different people within the same group.
Why Ecological Studies Are Still Valuable
Despite the ecologic fallacy, ecological studies serve several important purposes. They are often the only feasible design when individual-level data are unavailable (e.g., comparing disease rates across countries using national statistics). They are useful for studying the effects of policies, environmental exposures, and other group-level variables that cannot be measured individually. They are also inexpensive and quick, making them excellent for hypothesis generation. The key is to interpret ecological associations cautiously and seek corroboration from individual-level studies.
Canadian Data Infrastructure for Cohort, Case-Control, and Ecological Designs
Most epidemiological work in Canada does not involve enrolling a fresh cohort. Instead, researchers reuse data that have already been collected through health-system encounters, surveys, environmental monitoring, and registries. Three pieces of national infrastructure show up repeatedly:
Population Data BC (PopData BC)
A platform that links de-identified individual-level administrative data across the BC Ministry of Health, Vital Statistics, PharmaNet (every dispensed prescription), the BC Cancer Registry, MSP physician billings, hospital discharges (DAD), and education and social-services data. Researchers obtain a study-specific extract under a data access agreement.
Designs supported: retrospective cohorts, nested case-control, case-crossover, ecological/spatial analyses, intervention evaluations using natural experiments. Other provinces have analogous systems: ICES (Ontario), MCHP (Manitoba), HDNS (Nova Scotia), IRSPUM (Quebec).
Health Data Research Network Canada (HDRN Canada / SPOR DSP)
A federation of provincial data centres (PopData BC, ICES, MCHP, etc.) that supports multi-jurisdictional studies under a single application. Each centre runs the analysis behind its own firewall and only summary results are shared, so individual-level data never crosses provincial lines. Useful when you need national power or generalisability across health systems.
CANUE: Canadian Urban Environmental Health Research Consortium
A national repository of standardised, postal-code- and DA-level environmental exposures: air pollution (NO2, PM2.5, O3), greenness (NDVI), walkability, noise, climate, neighbourhood SES indices. CANUE indicators can be linked to any cohort with postal codes (including PopData BC extracts and CCHS shared files), turning subject-level health data into ecological or multilevel exposure–outcome studies.
Worked Example: A PopData BC + CANUE Cohort Study
To estimate the effect of long-term PM2.5 exposure on incident cardiovascular disease in BC adults, a researcher could:
- Define the cohort using the PopData BC Consolidation File, that is, everyone with active MSP coverage on 1 Jan 2010 (a near-complete census of BC residents).
- Pull baseline covariates from MSP and DAD (chronic conditions, comorbidity scores).
- Link each person's six-character postal code to CANUE annual PM2.5 estimates to assign exposure.
- Follow forward in DAD/Vital Stats to ascertain incident MI, stroke, and CV death (the outcomes).
- Apply Cox regression with a time-varying exposure, adjusting for area-level deprivation (also from CANUE).
This is a retrospective cohort with an environmental exposure, sitting at the boundary of cohort and ecological designs. The data come from three different stewards but no new participant was ever recruited.
Evidence Synthesis: Systematic Reviews
No single study, regardless of design, can definitively establish a causal relationship. Scientific knowledge accumulates through the synthesis of evidence across multiple studies, populations, and designs. Systematic reviews are the formal method for achieving this synthesis.
What Is a Systematic Review?
A systematic review uses a pre-specified, transparent, and reproducible protocol to identify, appraise, and synthesize all available evidence on a specific research question (Higgins et al., 2019). Unlike a narrative review (where the author selects studies subjectively), a systematic review:
- Defines a precise research question (often using the PICO framework: Population, Intervention/Exposure, Comparison, Outcome).
- Conducts comprehensive, systematic searches of multiple databases.
- Applies explicit inclusion and exclusion criteria.
- Critically appraises the quality and risk of bias of each included study.
- Synthesizes findings quantitatively (meta-analysis) or qualitatively.
Meta-Analysis
When studies are sufficiently similar, their results can be combined statistically in a meta-analysis to produce a pooled estimate of effect (DerSimonian & Laird, 1986). Meta-analysis increases statistical power, provides more precise estimates, and can explore heterogeneity across studies (e.g., do results differ by study design, population, or exposure definition?).
The Hierarchy of Evidence
Systematic reviews and meta-analyses of well-conducted studies sit at the top of the traditional evidence hierarchy (Sackett, Rosenberg, Gray, Haynes, & Richardson, 1996). Below them are randomized controlled trials, then cohort studies, then case-control studies, then cross-sectional studies, then ecological studies, and finally case reports and expert opinion. However, the hierarchy is a guideline, not an absolute rule: a well-designed cohort study may provide stronger evidence than a poorly conducted RCT, and Concato, Shah, & Horwitz (2000) showed empirically that observational results often agree closely with RCTs on the same question.
The companion R script r-activities/HSCI_341_Lesson_8_Review_of_Study_Design_Concepts.R defines a small choose_measure() helper that returns the design-appropriate measure (PR for cross-sectional, RR for cohort, OR for case-control) from a single 2x2 table with cells a = 194, b = 1588, c = 303, d = 1314. You then call it three times to see how the same cells produce different measures depending on what design you claim to have run.
# 2x2 cells
a <- 194; b <- 1588
c <- 303; d <- 1314
choose_measure <- function(a, b, c, d, design) {
switch(design,
"cross-sectional" = c(PR = (a/(a+b)) / (c/(c+d))),
"cohort" = c(RR = (a/(a+b)) / (c/(c+d))),
"case-control" = c(OR = (a*d) / (b*c)),
stop("design must be one of: cross-sectional, cohort, case-control"))
}
choose_measure(a, b, c, d, "cross-sectional")
choose_measure(a, b, c, d, "cohort")
choose_measure(a, b, c, d, "case-control")
# side-by-side
designs <- c("cross-sectional", "cohort", "case-control")
sapply(designs, function(d_) choose_measure(a, b, c, d, d_))
What you should be able to do after this activity: pick the design-appropriate measure of association from a 2x2 table, explain why PR and RR are numerically identical, and articulate when (and why) an OR diverges from an RR.
R Reflect on what you just ran
Use the questions below to interpret the three numbers choose_measure() produced. Look at your console output before answering.
1. Report the three returned measures (PR, RR, OR) from the same cells. Which two are numerically identical, and why does the formula make that the case?
2. The OR differs from the RR. By how much? Is the rare-disease assumption reasonable given a = 194 vs. b = 1588 and c = 303 vs. d = 1314, and how does the cell pattern explain the gap?
3. Why is it INCORRECT to report a prevalence ratio (or risk ratio) from a case-control study, even though the formula technically runs in R? What is being held fixed by the case-control sampling design that breaks that interpretation?
Reflection: Thinking Across Study Designs
Consider a public health question you care about (e.g., the effect of air pollution on childhood asthma, or the impact of a sugar tax on obesity rates). Which study designs would be most appropriate for different aspects of this question? Why might you need evidence from multiple designs to draw a convincing causal conclusion?
Minimum 20 characters required.
Key Takeaways
- Ecological studies use groups as the unit of analysis and can examine aggregate, environmental, or global measures.
- The ecologic fallacy occurs when group-level associations are incorrectly attributed to individuals.
- Ecological studies are valuable for hypothesis generation, policy evaluation, and studying group-level exposures, but results must be interpreted cautiously.
- Systematic reviews use transparent, reproducible methods to synthesize evidence across studies.
- Meta-analysis can pool results statistically for greater precision and power.
- No single study design is sufficient for establishing causation; converging evidence from multiple designs strengthens causal inference.
1. The ecologic fallacy occurs when:
2. Population density is an example of which type of group-level variable?
3. A systematic review differs from a narrative review primarily because it:
✦ Complete the reflection and pass the knowledge check with 100% to continue
Introduction & Time-Based Case-Only Designs
⏱ Estimated reading time: 18 minutes
Introduction and Overview
An earlier lesson consolidated the four standard observational designs. This lesson introduces the variants that combine or extend them. Hybrid designs are the answer to specific limitations of the standard four, such as rare or expensive exposures, transient triggers (Suissa, 1995), and surveillance data without obvious controls. The three content sections move from time-based case-only designs (this section: case-crossover and self-controlled case-series), through case-only comparison designs (a later section), to case-cohort and two-stage sampling designs (a later section) that subsample from larger cohorts to make biomarker-heavy studies affordable.
Learning Objectives
- Understand why hybrid study designs were developed and how they fit alongside the traditional cohort, case-control, and cross-sectional designs.
- Describe the design logic of case-crossover studies and identify when they are appropriate.
- Distinguish between unidirectional and bidirectional referent selection strategies.
- Describe the self-controlled case-series design and recognise the contexts in which it is most useful.
What Are Hybrid Study Designs?
By now you are familiar with the classic observational designs: cohort, case-control, and cross-sectional studies. Hybrid designs are variants of these classic designs that have been developed to address particular methodological challenges such as expensive covariates, rare exposures, and transient triggers, or surveillance data where traditional control selection is problematic.
This lesson covers six hybrid designs plus one important sampling strategy. Four of the hybrid designs use only cases (no separate control group), while two use a control series. The two-stage sampling design, by contrast, is a strategy that can be layered onto any of the traditional designs to enhance efficiency.
Why a Family of Hybrid Designs?
Each hybrid design solves a specific problem. Case-crossover studies eliminate the difficulty of choosing controls for transient exposures. Case-cohort studies allow one comparison group to support the study of multiple outcomes. Case-only studies allow inferences about gene-environment interactions when a control group is impractical. Two-stage designs let researchers spend money on detailed measurement only where it matters most. Knowing the “problem” each design was created to solve makes it much easier to remember when to use it.
The Six Hybrid Designs at a Glance
Click any card to see a brief description of the design and its key feature.
Case-Crossover Studies
The case-crossover study is the observational analogue of the experimental crossover design. Each case serves as its own control by contrasting exposure during a defined time window before the event with exposure during one or more comparison time windows.
Maclure (1991) introduced the design to answer the “why now” question, in contrast to the “why me” question answered by traditional case-control studies. By using the same person as both case and control, the design automatically controls for all time-invariant confounders, including ones the investigator never measured or even thought of. Maclure and Mittleman (2000) review a decade of applications.
When Is a Case-Crossover Design Appropriate?
Three Conditions Must Hold
1. The exposure must be transient. Stable exposures (such as smoking status or chronic medication use) cannot be evaluated because they would be present in all time windows.
2. The outcome must be acute. The event must happen close in time to the exposure if a causal relationship exists. Diseases with long induction periods are unsuitable.
3. The exposure must not be affected by the outcome. If experiencing the event changes future exposure (e.g., a heart attack alters subsequent activity), bidirectional control selection is problematic.
Defining the Risk Period and Control Period
Two design choices drive the validity of a case-crossover study: the length of the risk period (sometimes called the case-risk window) and the strategy for selecting control periods (sometimes called referent periods).
The risk period is the time during which the exposure, if causal, would have produced the event. Choosing a risk period that is too long increases the chance of detecting spurious associations; too short, and real associations may be missed. For physical exertion and myocardial infarction the risk window might be a few hours; for mobile phone use and motor vehicle crashes it might be five minutes; for air pollution effects on respiratory hospitalisations it is typically one day.
Figure 11.1. A symmetric bidirectional case-crossover design. One control window is selected before the event and one after, balancing potential time trends in exposure.
Strategies for Selecting Control Periods
Unidirectional (Backward) Referent Selection
Control periods are chosen only from time before the event. This was the original case-crossover approach. It is the appropriate choice when the event itself alters future exposure, for example when a leg injury changes subsequent training distance, or food poisoning alters what someone eats afterward.
Limitation: If exposure prevalence changes over time (a long-term trend), comparing only earlier control periods with the case-risk period can produce biased estimates.
Symmetric Bidirectional Referent Selection
Control periods are selected both before and after the case event, often equally spaced. The intent is that, if exposure is trending, the higher and lower exposure values from the two flanking control periods will roughly cancel out. This is now the most widely used approach.
Limitation: Cases that occur very early or very late in the study period may have only one control period feasible. Bidirectional selection is only valid if the event itself does not affect future exposure.
Time-Stratified Referent Selection
Janes, Sheppard, & Lumley (2005) proposed this method when shared-exposure data (such as daily air pollution measurements) are available across the entire observation period. The study period is stratified a priori (e.g., by month). When a case occurs, say on a Wednesday in July, all the other Wednesdays in July serve as control periods. This effectively matches on day-of-week and month and avoids the need to specify a single lag time.
Advantage: Eliminates the controversy over how to choose the spacing between case and control periods. Naturally accommodates shared exposure data.
Example 11.1: Weather Events and Waterborne Disease Outbreaks
Thomas and colleagues (2006) studied 92 waterborne disease outbreaks in Canada between 1975 and 2001. They hypothesised that extreme rainfall and warm spring conditions might trigger outbreaks. For each outbreak, the six weeks immediately before onset served as the case-risk period. The 27-year period was stratified into six time windows, and within each non-case window a six-week control period was selected, matched to the case on month, day, and ecozone. Conditional logistic regression identified warmer temperatures and extreme rainfall as plausible contributors.
Notice how the design eliminates the need to find “control communities” that did not have an outbreak, a perennial difficulty in waterborne disease epidemiology. Each outbreak community is its own control.
Example 11.2: Salmonella Outbreak in Long-Term Care
Haegebaert and colleagues (2003) used a case-crossover design within a foodborne Salmonella outbreak that affected mostly residents of chronic-care institutions. Food exposures during the three days before illness onset were compared with food exposures during a control period three days long, ending two days before the case-risk period. Because the illness itself would change subsequent food intake, only earlier (unidirectional) control periods were used. Mantel-Haenszel matched-pair odds ratios were calculated for each meat product. Notice that the design avoided the difficult problem of selecting institutionalised “controls” whose food intake would otherwise have to be matched.
Analysis of Case-Crossover Data
Because each case is matched to one or more control periods within the same individual, the data are analysed as if from a matched case-control study. With one control period per case, the data fit a 2×2 table and McNemar's test applies. Intuitively, this comparison ignores occasions when exposure was the same in both windows and asks, among the occasions where it differed, whether exposure fell in the risk window more often than in the control window. With multiple control periods, conditional logistic regression is the standard approach, and the exponentiated coefficient represents the change in odds of the event associated with a one-unit short-term increase in exposure.
When daily exposure data are available for the entire observation period (the “shared exposure” setting common in air pollution studies), the data can equivalently be analysed as a Poisson time series. The two analytical frameworks are mathematically linked when time-stratified referents are used.
Self-Controlled Case-Series Studies
The self-controlled case-series design (often shortened to “case-series” in this literature, but distinct from the descriptive case-series of clinical reports) was developed by Farrington (1995) largely for vaccine safety research. It is a close cousin of the case-crossover design but generalises the comparison from discrete control periods to all of an individual's observation time outside the risk window. Whitaker, Farrington, Spiessens, & Musonda (2006) provide an accessible tutorial.
The Logic of the Design
For each individual who has experienced the outcome of interest, an observation period is defined, namely a calendar window during which exposure history and event occurrence are tracked. Within that observation period, one or more risk periods are designated based on the biology of the exposure (e.g., 6–35 days after vaccination for febrile conditions). All remaining time within the observation period constitutes the control period.
The analysis compares the rate of events during risk time with the rate during control time, after adjusting for the duration of each. As with the case-crossover design, this is a within-person comparison: every time-invariant characteristic of the case (genetics, sex, baseline health) is automatically controlled by design. Age and season can be adjusted for analytically because they vary across the observation period.
Figure 11.2. The observation period for a single case is partitioned into risk periods (after each exposure) and control periods (everything else). The number of events and the duration of each period type drive the relative incidence estimate.
The SCCS is just a conditional Poisson model where each case acts as its own stratum. Below: build a long-format dataset where each case contributes one row of risk-period person-time and one row of control person-time, then fit the model with gnm.
# install.packages(c("gnm", "SCCS"))
library(gnm)
# Long-format SCCS data: 4 cases, each with risk and control periods
sccs <- data.frame(
case = rep(1:4, each = 2),
period = rep(c("risk", "control"), 4),
events = c(1, 0, 1, 0, 2, 1, 0, 1),
pt = c(30, 335, 28, 337, 30, 335, 29, 336)
)
# Conditional Poisson on case (each case is its own intercept)
fit <- gnm(events ~ period,
offset = log(pt),
family = poisson,
eliminate = factor(case),
data = sccs)
exp(coef(fit)) # incidence rate ratio (risk vs. control)
exp(confint.default(fit))
Why this works. All time-invariant features of each case (sex, genes, baseline frailty) drop out because we condition on the case ID. What remains is the risk-vs-control rate ratio, which is exactly the design's quantity of interest. The dedicated SCCS package wraps this with helpful diagnostics.
R Reflect on what you just ran
Use the questions below to interpret the actual numbers you produced. Look at your console output before answering.
1. exp(coef(fit)) returned the incidence rate ratio (IRR) for risk vs. control periods. What value did you get, and in one sentence what does it mean for a case during their risk window?
exp(coef(fit)) returns an incidence rate ratio of about 23 for the risk vs. control period. Because the dataset is fixed (there is no random simulation here), everyone gets the same value. Interpretation: during a case's risk window, the interval right after the trigger, their event rate is roughly 23 times higher than during their control time. The self-controlled case-series lets you read this off directly, because each case serves as their own control and the ratio is a pure within-person contrast.2. From exp(confint.default(fit)), report the 95% CI. Does it cross 1, and what does that tell you about statistical significance?
3. The model used eliminate = factor(case). Explain in your own words why this design feature means you do NOT need to adjust for sex, genetics, or any other time-invariant confounder.
eliminate = factor(case) stratifies the conditional likelihood by individual: each case acts as their own control, so any characteristic that does not change within the person (sex, genetics, baseline lifestyle, chronic comorbidities) is held constant within the comparison. Mathematically, time-invariant person-level effects drop out of the conditional likelihood because both the risk and control windows belong to the same person. That is the whole point of the self-controlled case-series (and of its close relative the case-crossover): it controls perfectly for unmeasured time-fixed confounders without ever measuring them.Key Assumptions
If a febrile reaction after a first vaccine dose causes parents to skip the booster, the exposure pattern is no longer independent of outcome. One way to deal with this is to ignore post-event exposures (i.e., consider only the first vaccination). Whitaker and colleagues note that the bias from violating this assumption is often small in practice, but it should be considered explicitly.
If the outcome is death (which clearly ends observation) or a serious illness that prompts withdrawal from the study, the design's assumptions are violated. The standard self-controlled case-series is designed for outcomes that occur and resolve, allowing observation to continue.
Multiple recurrences of the outcome can be included as long as they are conditionally independent given exposure. If they are not (e.g., one event makes another more likely), only first events should be analysed.
If the observation period or risk window does not cover the full duration over which exposure can affect the outcome, any resulting estimate of relative incidence is biased toward the null. Sample size formulae are given in Whitaker, Hocine, & Farrington (2009).
Analysis
The standard analytic tool is a conditional Poisson regression model, where the outcome is the count of events in each risk and control time interval and the logarithm of the duration of each interval is the offset. The parameter of interest is the relative incidence, that is, the rate during the risk period relative to the rate during the control period.
Example 11.3: Falls and Antihypertensive Medication
Gribbin and colleagues (2011) used UK primary-care databases to study whether starting an antihypertensive medication transiently increased the risk of falls in adults aged 60 and older. They identified 9,862 falls between 2003 and 2006. For each patient, episodes of continuous medication exposure of up to 60 days were defined. After each prescription, the exposure period was further subdivided into day 0, days 1–21, and days 22–60. All remaining person-time was the unexposed baseline. Poisson regression yielded incidence rate ratios for each post-exposure period, allowing the temporal pattern of risk after initiation to be characterised.
Why is this question well-suited to a self-controlled case-series? Because the comparison is within-person, all the patient-level confounders that complicate fall risk (frailty, polypharmacy, comorbidity, age) are automatically controlled.
Key Takeaways
- Hybrid designs are variants of the classic observational designs developed to address specific methodological challenges.
- Case-crossover studies use each case as its own control by comparing exposure during a risk period with exposure during one or more control periods at other times. They are best for transient exposures and acute outcomes.
- Three referent-selection strategies for case-crossover studies are unidirectional, symmetric bidirectional, and time-stratified. The choice depends on whether the event itself alters subsequent exposure and whether time trends in exposure are likely.
- Case-crossover data are typically analysed by conditional logistic regression; with shared daily exposure data, an equivalent Poisson time-series approach is available.
- The self-controlled case-series partitions each case's observation period into risk windows (defined by exposure timing) and control time (everything else). Conditional Poisson regression yields the relative incidence.
- Both designs automatically control all time-invariant confounders, including those the investigator has not measured.
1. The case-crossover design is most appropriate for studying:
2. A unidirectional (backward-only) referent selection strategy is preferred when:
3. The chief advantage shared by both case-crossover and self-controlled case-series designs is that they:
✦ Pass the knowledge check with 100% to continue
Case-Only Comparison Designs
⏱ Estimated reading time: 17 minutes
Introduction and Overview
An earlier section covered designs that use each case as their own control across time. This section turns to designs that compare different kinds of cases to one another, useful when traditional controls are unavailable, when subtypes of disease have different aetiologies, or when surveillance data only contain ill people. Each design in this section sacrifices something in exchange for not needing a healthy control group.
Learning Objectives
- Describe the design and applications of case-case studies.
- Distinguish case-case studies from case-case-control studies.
- Explain the logic of case-only studies for evaluating gene-environment interactions.
- Recognise the assumptions and limitations of each design.
Case-Case Studies
The case-case design is a variant of the case-control design in which the comparison group consists of cases of a different disease subtype drawn from the same surveillance system. McCarthy and Giesecke (1999) proposed it as an efficient way to identify risk factors that distinguish closely related etiological subgroups using routine surveillance data.
For example, the cases might be people infected with Salmonella Typhimurium, while the “controls” might be people infected with Salmonella Heidelberg. Both groups have salmonellosis (both are cases), but the design seeks to identify exposures that distinguish one serotype from the other.
When Is the Case-Case Design Useful?
Two Common Settings
1. Identifying differential risk factors for related endemic diseases. When all subjects who appear in a surveillance system have undergone similar selection (e.g., they all sought medical care, all had stool cultured), comparing them with one another minimises selection bias and recall bias. Comparing them to community controls who never had any salmonellosis would be far more vulnerable to these biases.
2. Distinguishing outbreak cases from sporadic cases of the same organism. In an outbreak investigation, the cases are people whose isolates match the outbreak strain. The “controls” are sporadic cases of the same serotype during the same time window. The exposures that differentiate them point to the outbreak vehicle.
Strengths and Limitations
Comparable selection experience. Because both groups appear in the same surveillance system, both have passed through similar diagnostic and reporting filters. Selection bias is minimised.
Comparable recall experience. Both groups have had a similar clinical experience (an episode of gastrointestinal illness). Their motivation to recall recent food exposures is similar, reducing differential recall bias.
Efficient use of surveillance data. No new control recruitment is required; the comparison group is already in the database.
Cannot identify shared risk factors. Exposures that cause both serotypes equally (such as eating any contaminated food) will not be detected because they are present in both groups.
Surveillance limitations. Wilson and colleagues (2008) note tendencies for selection bias (only severe cases reported), information bias (data collected by people who know the diagnosis), confounding (limited covariate information), and lack of detail on exposure.
The OR is not a true risk measure. Because the “controls” are not drawn from the underlying source population, the odds ratio reflects the relative difference in exposure between two case subtypes, not the absolute risk of either disease.
If the analysis identifies poultry consumption as a stronger risk factor for S. Typhimurium than for S. Heidelberg, this does not tell us that eating poultry causes Typhimurium in absolute terms. Rather, it tells us that poultry consumption is more strongly associated with the Typhimurium subgroup than with the Heidelberg subgroup, which is useful for tracing distinct food sources or transmission routes.
To get an absolute risk estimate, a traditional case-control study with population-based controls would still be required. Case-case findings are best treated as hypothesis-generating about subtype-specific exposures.
Example 11.4: Two Campylobacter Species
Gillespie and colleagues (2002) used population-based surveillance data from England and Wales to compare the exposure histories of people with Campylobacter coli infection (the much rarer species) with those of people with Campylobacter jejuni infection. Standard structured questionnaires from the surveillance system provided the exposure data. Backward stepwise logistic regression identified differential risk factors and tested for interaction. The authors emphasised that exposures common to both species would not be detected by this design; only those that distinguish the two species could emerge.
If you wanted to know which exposures were common to both species, what design would you use instead?
Example 11.5: A Salmonella Outbreak in Germany
Krumkamp and colleagues (2008) investigated a 2003 outbreak of Salmonella 1,4,[5],12:i:- in a German district. Ten outbreak cases were compared with 97 sporadic cases of other Salmonella serotypes that occurred in the same area during the same year. Telephone interviews collected exposure histories. Fisher's exact tests and odds ratios identified meat sold from a single butcher shop as the only significant risk factor, a finding that would have been very difficult to achieve with a traditional community-based control group.
Analysis
Case-case data are analysed by the same techniques as risk-based case-control studies, typically logistic regression. The exponentiated coefficient is interpreted as the relative odds of exposure between the two case subtypes, not as a risk ratio.
Case-Case-Control Studies
The case-case-control design (Kaye, Harris, Samore, & Carmeli, 2005) was developed to overcome a specific limitation of traditional case-control studies in the context of antimicrobial resistance. The original example was vancomycin-resistant Enterococcus (VRE) versus vancomycin-susceptible Enterococcus (VSE).
The Problem the Design Solves
Suppose you want to identify risk factors for VRE infection. A traditional case-control approach would compare VRE cases with non-infected controls. But many of the exposures associated with VRE (prior antibiotic use, prolonged hospitalisation, ICU stay) are also strong risk factors for VSE, and indeed for any hospital-acquired infection. So a traditional design tells you what causes hospital infection in general, not what specifically drives the resistant phenotype.
An alternative is a case-case design comparing VRE with VSE. But Kaye and colleagues argue against this: VRE often emerges from external sources (transmission of an already-resistant strain) rather than from within-patient evolution of a susceptible strain. So contrasting VRE directly with VSE conflates “risk of acquiring a resistant strain” with “risk of selection pressure on a susceptible one.”
The Case-Case-Control Solution
The design uses two case series (resistant and susceptible) and one control series (people without infection from the same source population). Two separate logistic regression models are fitted, one comparing each case series with the controls. The risk factors are then sorted into three categories.
Figure 11.3. In the case-case-control design, two case series are each compared separately with the same control series. Comparison of the two resulting models identifies which risk factors are unique to the resistant phenotype.
Interpreting the Three Variable Categories
These are risk factors unique to the resistant phenotype. They are the variables most useful for understanding what drives resistance specifically. In the original VRE example, exposure to vancomycin itself or to a roommate carrying VRE might fall in this category.
These are risk factors unique to the susceptible phenotype. They tell us what predisposes to acquiring the susceptible strain in particular (perhaps community sources for the susceptible organism but not for the resistant one).
These are risk factors for the target organism in general, regardless of resistance status. Hospitalisation, prior antibiotic exposure, indwelling catheters, and severity of illness typically appear here. They are real risk factors but they do not distinguish resistance from susceptibility.
Design Considerations
- First positive culture per patient. Only the first positive culture should be included to avoid double-counting; for nosocomial infection studies, restrict to cultures taken >48 hours after admission.
- Source population matters. Controls should come from the same source population as the cases. For nosocomial infections, controls should be other patients hospitalised >48 hours, ideally with documented negative cultures for both phenotypes.
- Confounding control. Because there is only a single control series, restricted sampling and matching are difficult. Confounding is typically handled through multivariable unconditional logistic regression.
Example 11.6: MRSA Colonisation in an ICU
Melo and Fortaleza (2009) investigated risk factors for nasopharyngeal colonisation with methicillin-resistant Staphylococcus aureus (MRSA) in an ICU. They enrolled 122 patients who had been screened weekly for S. aureus colonisation. The two case series were patients colonised with MRSA and patients colonised with methicillin-susceptible S. aureus (MSSA). Controls were patients in whom no colonisation was detected during their ICU stay. Comparing the resulting two models revealed which exposures were specifically associated with the resistant phenotype rather than with general susceptibility to S. aureus colonisation.
Case-Only Studies
The case-only design uses only cases, with no observed control group recruited. The expected exposure distribution in the hypothetical “control population” is derived from theoretical or external sources. The design originated in genetic epidemiology, where the population frequency of common alleles can often be specified from external reference data (Khoury & Flanders, 1996).
What the Design Can and Cannot Estimate
Important Restriction
The case-only design cannot estimate main effects; it cannot tell you whether a gene or an environmental exposure independently raises the risk of disease. What it can estimate is interaction between two factors among cases, provided the two factors are independent of each other in the source population.
The intuition is this: among cases, if a genetic risk factor and an environmental exposure are independent in the source population but appear together more often than expected by chance, that excess co-occurrence is evidence of statistical interaction on the multiplicative scale. If the gene and exposure were also causally associated in the source population (not independent), this signal would be confounded.
Required Assumptions
- Independence in the source population. The exposure and the proposed effect modifier (often a gene, but can be sex, age, or another stable trait) must be independent in the population from which cases arose. For a heritable polymorphism not influenced by the environmental exposure, this is biologically plausible.
- Stable, well-defined effect modifier. Genetic variants, sex, race, and age are common choices because they don't change over time and can be measured reliably.
- The disease must be rare. Like many odds-ratio-based estimators, the case-only interaction estimate approximates the true interaction parameter most closely when the outcome is rare in the source population.
Recent Extensions Beyond Genetics
The design has been extended to study how non-genetic stable characteristics modify the effects of time-varying exposures. Armstrong (2003) and Schwartz (2005) used case-only designs to ask whether sex, race, age, or socioeconomic class modify the effect of extreme weather on mortality. Because age, sex, and socioeconomic class can reasonably be considered independent of daily weather exposures, the case-only approach yields a valid interaction estimate.
Analytic Logic
Suppose we are interested in whether sex modifies the effect of an extreme heat day on mortality. A case-only logistic regression takes the form:
The logic feels strange at first because we appear to be modelling a covariate (sex) as a function of an exposure. But this is mathematically equivalent to a Poisson model of mortality count as a function of heat, sex, and a heat×sex interaction term, where the case-only regression coefficient is the interaction term. The trick is that we never need a control group at all; the “control” expectation is built into the assumed independence of sex and heat in the source population.
Example 11.7: Effect Modifiers of Mortality from Temperature Extremes
Schwartz (2005) investigated whether sex, non-white race, or age over 85 modified the effect of extreme temperatures on mortality in Wayne County, Michigan. Weather data identified excessively hot and cold days. Demographic data on people who died came from medical records. Separate models were fitted for heat and for cold, and one-day and three-day average temperature exposures were both examined. All three covariates emerged as effect modifiers. Notice that no control group of survivors was needed; the inference depended on whether the demographic profile of cases differed between extreme-weather days and other days.
Why is this design especially appealing for studying mortality? Because building a comparable control group of “people who did not die” is conceptually awkward when daily death registry data are already complete.
Example 11.8: Heat Waves and Hospital Admissions in New South Wales
Khalaj and colleagues (2010) used a case-only design to identify which underlying medical conditions raised the risk of hospital admission during heat waves across five regions of New South Wales, Australia. Daily admission records and weather data covered the warm months of 1998–2006. The analysis fitted logistic regression models with each primary diagnosis as the “outcome” and an extreme-heat indicator as the predictor. Sine and cosine terms were included to control for season, since otherwise season could confound the interaction (some chronic conditions have stronger seasonal patterns than others).
Key Takeaways
- Case-case studies compare two related disease subtypes drawn from the same surveillance system, identifying differential risk factors while minimising selection and recall bias. The OR reflects relative differences in exposure between subtypes, not a true risk measure.
- Case-case-control studies use two case series (e.g., resistant and susceptible) compared separately with one control series, sorting risk factors into category A (unique to resistance), B (unique to susceptibility), and C (shared by the organism in general).
- Case-only studies use only cases and rely on external knowledge of the exposure distribution in “controls.” They estimate interaction between an exposure and an effect modifier, but not main effects, and require independence between the two factors in the source population.
- All three designs use cases-only or two-case data structures because constructing a satisfactory traditional control group would be impractical, biased, or uninformative for the specific research question.
1. A case-case study comparing Salmonella Typhimurium with Salmonella Heidelberg cases:
2. In a case-case-control study of vancomycin-resistant Enterococcus, a Category C variable is one that:
3. The case-only design's most important limitation is that it:
✦ Pass the knowledge check with 100% to continue
Case-Cohort & Two-Stage Designs
⏱ Estimated reading time: 18 minutes
Introduction and Overview
Earlier sections were about case-only designs. This section returns to designs that use cohorts as their backbone but subsample from them to make expensive measurements feasible. Both case-cohort and two-stage designs let you mount essentially a cohort study while only paying for biomarker measurements on a fraction of the participants, the kind of design that makes large biobanks practical.
Learning Objectives
- Describe the structure and rationale of the case-cohort design.
- Distinguish risk-based from rate-based case-cohort analyses.
- Explain why a single subcohort can support investigation of multiple outcomes.
- Describe the logic of two-stage sampling and identify when it is most efficient.
- Design a basic two-stage sampling strategy for a case-control study.
Case-Cohort Studies
The case-cohort design, introduced by Prentice (1986), combines features of cohort and case-control studies. From a defined source cohort, the investigator draws a random sample called the subcohort at the start of follow-up. Detailed exposure and covariate data are obtained on the subcohort. As follow-up proceeds, all incident cases that arise from the full source cohort, whether or not they happen to fall within the subcohort, are also studied.
The design has the same advantages as a full cohort study (clear temporal ordering, multiple outcomes, direct disease frequency estimates) but achieves them with much smaller measurement costs because expensive covariate or biomarker assays are performed only on the subcohort plus the cases, not on the entire source cohort.
Figure 11.4. The case-cohort layout. Detailed exposure and covariate data are needed only on the random subcohort plus the cases. Most of the full source cohort never requires expensive measurement.
The Big Win: Multiple Outcomes from a Single Subcohort
Why Researchers Love Case-Cohort Designs
One subcohort can serve as the comparison group for multiple disease outcomes. If researchers are interested in cardiovascular disease, several cancers, and diabetes within the same large cohort, they need only one set of expensive biomarker measurements on the subcohort. Each outcome study then adds detailed measurements on its own cases. By contrast, a nested case-control study requires fresh control selection for each outcome, and a full cohort analysis would require measuring everyone for everything.
Risk-Based vs. Rate-Based Designs
Risk-Based (Closed-Cohort) Case-Cohort
Suitable when the source cohort is closed (a fixed group followed for a defined period) and exposures are stable over follow-up. The subcohort is sampled by simple or stratified random sampling at the start of follow-up. Cases arising outside the subcohort during follow-up are added.
Analysis: Combine the two case groups (those in and those outside the subcohort) and analyse the data in the familiar 2×2 case-control format using logistic regression. The odds ratio approximates the relative risk when the disease is rare.
Example: Matsuda and colleagues (2011) studied placental abruption and placenta previa among 5,036 of 242,715 births in Japan, using multivariable logistic regression with the subcohort plus all cases.
Rate-Based (Open-Cohort) Case-Cohort
Suitable when the source cohort is open (entries and exits possible during follow-up) or when exposures change over time. At the moment a case occurs, eligible members of the subcohort are those who have not yet experienced the outcome. Their current exposure status (which may have been updated through repeated surveys or stored serial samples) is recorded.
Analysis: A weighted Cox proportional-hazards model is the standard approach. Weights account for the sampling fraction; for example, if the subcohort represents 20% of the source cohort, controls are typically up-weighted by 5. Three Cox-weighting schemes have been proposed historically; Prentice's method most closely reproduces the estimates from a full-cohort analysis.
Example: Agalliu and colleagues (2011) followed a subcohort of 1,979 men for prostate cancer risk, with exposure and supplement use updated through repeated surveys.
Practical Considerations
- Eligibility for the subcohort. Members must be willing to provide health history, lifestyle data, and (often) biological samples. Stratified sampling can ensure that the subcohort's covariate profile matches anticipated cases (e.g., over-sampling young adults if young-adult disease is the focus).
- Stored specimens. Serially stored tissue or blood samples allow detection of exposure changes over time and support post-hoc biomarker assays as new hypotheses emerge.
- Sampling adjustments for non-response. If 20% are sampled but only 80% of those agree to participate, the weighting should reflect the actual participation, not the original sampling probability.
- Robust standard errors are recommended for case-cohort analyses to account for the sampling variability.
- Clustering of cases. If cases tend to be diagnosed at the same clinic, marginal models with adjusted variances or frailty models should be used to account for within-cluster correlation.
Example 11.9: Drinking Water Quality and Stomach Cancer
Auvinen and colleagues (2005) studied radon and other radionuclides in drinking water and the risk of stomach cancer in a Finnish population of over 144,000 people who drew their water from drilled wells between 1967 and 1980. An initial subcohort of 4,590 was sampled with stratification by age and sex. Many of these did not actually meet the long-term-exposure criterion, leaving an effective subcohort of 371 long-term users. Stomach cancer cases (n=107) were identified through the cancer registry. Water samples were collected blindly with respect to case status and analysed for radionuclides. A proportional-hazards model accounted for how long each subject had been exposed to each level of radon. All hazard ratios were below 1, suggesting a protective association, a surprising finding that illustrates how case-cohort designs can efficiently support investigation of unusual exposures using stored samples and registry data.
Two-Stage Sampling Designs
A two-stage (or two-phase) sampling design is a strategy that can be layered on top of any traditional design, whether cohort, case-control, or cross-sectional. The first stage collects readily available, inexpensive data on a large group. The second stage collects more detailed (and usually more expensive) data on a strategically selected subsample.
Why Two-Stage Designs Make Sense
The Core Problem
Imagine you want to study whether occupational solvent exposure increases birth defect risk. Hospital records can give you basic information on hundreds of thousands of pregnancies cheaply, but a detailed occupational exposure assessment requires a one-hour interview at $200 per participant. Spending $200 on every pregnancy is unaffordable. A two-stage design lets you do the cheap step on everyone and the expensive step only where the information is most valuable.
Three Common Use Cases
The first stage uses an inexpensive surrogate exposure measure (e.g., job title from a registry). The second stage performs a detailed work-up (e.g., personal interview about specific solvent contacts, dose, duration) on a subsample. This is the most common application.
If the inexpensive first-stage measure has known measurement error, the second stage applies a near-gold-standard measurement to a subsample. The relationship between the two measures (the “measurement model”) is estimated, and inferences from the full first-stage data set can be corrected for the measurement error. McNamee (2002, 2005) describes optimal designs for this purpose.
When key covariate data are missing for many subjects, instead of assuming missingness at random, the missing-data subjects can be the explicit target of the second-stage data collection. This concentrates resources on filling specific gaps rather than dropping incomplete records.
How to Sample at Stage 2
The key design question in any two-stage study is: how should we choose whom to include in the expensive second stage? The optimal answer depends on the design.
| Stage 1 Design | Recommended Stage 2 Sampling | Rationale |
|---|---|---|
| Cohort | Fixed numbers of exposed and unexposed | Balanced sampling on exposure ensures precision in the exposure–outcome estimate. |
| Case-control | Fixed numbers of cases and controls | Balanced sampling on disease ensures precision; oversampling cases is efficient when the disease is rare. |
| Either, when surrogate exposure is available | Approximately equal numbers from each of the four exposure×disease cells | Optimal efficiency: extracts the most information from a fixed second-stage budget by ensuring that small cells (rare combinations) are not under-represented. |
A Worked Two-Stage Case-Control Sampling Strategy
Suppose your stage 1 data come from a hospital registry of 50,000 pregnancies. From the registry you can identify 2,000 birth-defect cases and 48,000 non-cases. A crude (and possibly mismeasured) exposure indicator, namely whether the mother held a job classified as “industrial”, is available for everyone.
Stage 1 cross-classification might look like this:
| Industrial job (surrogate exposed) | Other job (surrogate unexposed) | Total | |
|---|---|---|---|
| Cases | 120 | 1,880 | 2,000 |
| Non-cases | 1,500 | 46,500 | 48,000 |
For stage 2, balanced sampling across the four cells is the most efficient strategy. Suppose your budget allows 400 detailed interviews:
- Industrial-exposed cases (120 available): Take all 120.
- Other-exposure cases (1,880 available): Sample 100.
- Industrial-exposed non-cases (1,500 available): Sample 100.
- Other-exposure non-cases (46,500 available): Sample 80.
Because we have oversampled the small cells, weighting must be applied in analysis to recover correct association estimates; this is what the Cain & Breslow (1988) and later Flanders & Greenland (1991) methodologies handle. Hanley and colleagues (2005) provide worked examples of the adjusted odds ratio and its variance.
Example 11.10: A Two-Stage Case-Control Study of Childhood Asthma
Martel and colleagues (2009) used a two-stage design with three linked Quebec administrative health databases. Stage 1 was a nested case-control study within a cohort of pregnant women and their children: 5,226 asthmatic children (cases) and 20 non-asthmatic children per case were selected using density sampling matched to time of case occurrence. Covariate data from the administrative databases were used at this stage. Stage 2 was a mailed questionnaire to a subsample of mothers, balanced across the cells of the first-stage exposure–outcome cross-table to overrepresent small cells. Conditional logistic regression was used at stage 1; unconditional logistic regression with sample-fraction weighting at stage 2. Final corrected estimates were obtained by combining the stages.
Notice how the design exploits the cheap administrative data to identify cases and screen on rough covariates, while reserving the expensive questionnaire for the subset where new information will most improve the estimate.
Practical Pitfalls
- Budget allocation between stages. Hanley and colleagues (2005) note that tools for optimal allocation have not advanced much in the past two decades; in practice, simulation can help determine the relative number of stage 1 and stage 2 subjects to maximise precision under a fixed budget.
- Variance estimation. The variance of the final estimate depends on both stages. Naive variances that ignore stage 2 sampling will be too small. Hanley and colleagues provide details for dichotomous covariates; software exists for more complex situations.
- Sampling fractions must be known. Weighting requires knowing exactly what proportion of each stage 1 cell was sampled at stage 2. If non-response or other losses change those fractions, the realised (not planned) fractions should be used.
Reflection: Choosing the Right Hybrid Design
Consider a research question of interest to you. It might involve a transient environmental trigger of an acute event, an outbreak you would like to characterise, a gene-environment interaction, or a long cohort follow-up where biomarker measurement is expensive. Which hybrid design would you choose, and why? What assumptions would you need to defend, and what limitations would you have to acknowledge in your discussion?
Minimum 20 characters required.
Key Takeaways
- Case-cohort studies sample a random subcohort at the start of follow-up and add all incident cases. Detailed measurement is needed only on the subcohort plus the cases, not the full source cohort.
- A single subcohort can serve as the comparison group for many outcomes, making the design particularly attractive for large prospective studies with stored biological specimens.
- Risk-based case-cohort analyses combine subcohort and outside cases in a logistic regression. Rate-based analyses use weighted Cox models, with weights reflecting the inverse sampling probability.
- Two-stage sampling lets investigators pay for cheap, low-quality data on everyone and high-quality data only on a strategically chosen subsample.
- For two-stage case-control studies, the most efficient stage 2 sampling allocates approximately equal numbers across the four cells of the stage 1 exposure×disease table.
- Two-stage analyses must use weighting to recover unbiased estimates and must use variance formulae that account for both stages of sampling.
1. The key efficiency advantage of a case-cohort design over a full cohort study is that:
2. In a rate-based case-cohort analysis of an open cohort, the standard analytic approach is:
3. For a two-stage case-control study with a binary surrogate exposure measure available at stage 1, the most statistically efficient stage 2 sampling strategy is to:
✦ Complete the reflection and pass the knowledge check with 100% to continue
Foundations of RCTs & Trial Setup
⏱ Estimated reading time: 18 minutes
Introduction and Overview
Earlier lessons worked through the observational designs, including cross-sectional, cohort, case-control, ecological, and the hybrid variants. Every one of them measures exposure as it occurs naturally, and every one of them must spend effort defending against confounding. This lesson turns to the experimental alternative: when the investigator controls who is exposed via random assignment, confounding is broken by the design itself rather than by statistical adjustment. The three content sections move through an RCT in the order an investigator would build one: foundations and setup (this section), the central design choices around allocation, outcome measurement, sample size, and blinding (a later section), and finally the conduct, analysis, and reporting of the trial (a later section).
Learning Objectives
- Define a randomised controlled trial and explain why RCTs are considered the gold standard for evaluating interventions.
- Describe the five phases of clinical research from pre-clinical work through Phase IV post-marketing surveillance.
- Distinguish between target population, source population, and study group.
- Develop appropriate eligibility criteria that balance internal validity with generalisability.
- Specify an intervention with the precision needed for replication.
What Is a Randomised Controlled Trial?
A randomised controlled trial (RCT) is a planned experiment in which the investigator deliberately allocates participants to one or more interventions, then follows them to observe outcomes. Because the allocation is determined by the investigator (rather than by self-selection or by clinical circumstance), randomisation produces groups that are comparable on both measured and unmeasured factors at baseline. This is the property that gives the RCT its inferential power.
Walk through the first modern RCT, the trial that turned scarcity into science. Next ▶ advances scenes.
A 7-scene retelling of the 1948 MRC streptomycin trial: postwar TB epidemic, the scarce U.S. drug shipment, Bradford Hill's elegant solution (random allocation), the slot-machine randomization, dramatic 6-month outcomes, BMJ publication, and the rise of the RCT as gold standard.
Throughout this lesson we follow the chapter convention of using RCT to describe any planned experiment evaluating products or procedures outside the laboratory. The terms clinical trial (often restricted to therapeutic products in clinical settings) and field trial (carried out in general population settings) are used interchangeably with RCT. The factor under investigation is called the intervention; the effect of interest is called the outcome; people or groups participating are called subjects or participants. The lineage of the controlled comparison reaches back to James Lind's 1753 shipboard trial of citrus for scurvy, but it was the 1948 MRC streptomycin trial that introduced formal random allocation as the design's defining feature.
Why RCTs Are the Gold Standard
RCTs allow much better control of potential confounders than observational studies and reduce bias from selection and misinformation. By randomly assigning the intervention, the investigator breaks any link between exposure and unmeasured confounders, a feat that no observational design can match. As Lavori and Kelsey put it, the RCT is at present the unchallenged source of the highest standard of evidence used to guide clinical decision-making. That said, a single RCT is rarely sufficient to answer questions about complex interventions, and concerns persist that some trials lack relevance to real-world practice.
CONSORT and Trial Registration
The Consolidated Standards of Reporting Trials (CONSORT) statement was developed to improve the quality of trial reporting. It was first published in 1996, updated in 2001, and again in 2010 (Schulz, Altman, & Moher, 2010). We will use its main headings as the structural template throughout this lesson; the full 25-item checklist appears in a later section. As of 2005, the International Committee of Medical Journal Editors required investigators to register their trials prior to participant enrolment as a precondition for publishing in member journals. The WHO operates an International Clinical Trials Registry Platform (ICTRP), although registration remains effectively voluntary in many jurisdictions.
The Phases of Clinical Research
While controlled trials are valuable for assessing a wide range of factors affecting health, one of their most common uses is to evaluate pharmacological products. Before any trial in humans, extensive pre-clinical studies are conducted in vitro (test tube or cell culture) and in vivo (animal) using wide-ranging doses to obtain preliminary efficacy, toxicity, and pharmacokinetic information.
Click any phase to see its purpose, typical sample size, and key design features.
First-in-Human
Safety
Efficacy Signal
Pivotal Efficacy
Post-Marketing
Background, Objectives, and Trial Design
The objectives of a trial must be stated clearly and succinctly. A good objective describes the intervention, the allocation design (parallel, factorial, cross-over, etc.), and the primary outcome(s). Each trial should have a limited number of objectives plus, if needed, a small number of secondary outcomes. Increasing the number of objectives complicates the protocol, jeopardises compliance, and sacrifices statistical power.
Most trials contrast two groups, intervention and comparison, and are sometimes referred to as two-arm studies. Trials with more arms can be efficient when factorial designs are used (a later section). The comparison group might receive a placebo, no treatment, the usual treatment, or a different dose of the same product.
Choosing the Comparator: A Consequential Decision
Placebos are ideal when there is no established alternative intervention; where possible, a placebo is preferred to “no treatment.” However, when an effective standard treatment exists, withholding it from the comparison arm may be unethical; randomisation is ethically defensible only when there is genuine uncertainty in the expert community about which arm is superior, a state Freedman (1987) called clinical equipoise. In these settings, the standard of care serves as the comparator, and the trial may take the form of a non-inferiority trial, which aims to show the new intervention is no worse than the existing standard by more than a clinically unimportant margin (delta). Determining the appropriate value of delta is one of the most consequential design decisions in a non-inferiority trial.
Example 12.1: A Trial of Prostate Cancer Screening
The Andriole and colleagues (2012) trial randomised 38,340 men aged 55–74 to a screening intervention and 38,345 to usual care across 10 USA screening centres between 1993 and 2001. Men in the intervention arm were offered annual PSA tests for six years and digital rectal examination for four years. Follow-up extended through 2009 or 13 years from trial entry. The primary analysis was an intention-to-screen comparison of prostate cancer-specific mortality. The trial illustrates how a clearly stated objective (mortality reduction from screening) can be embedded in a simple two-arm parallel design that can nonetheless run for two decades.
Notice that the comparator was “usual care,” which sometimes included opportunistic screening. This pragmatic choice (Tunis, Stryer, & Clancy, 2003; Loudon et al., 2015) makes the results applicable to the real US health system but complicates interpretation of the “true” effect of organised screening.
Participants: Defining the Study Group
Three nested populations must be distinguished in any trial:
Figure 12.1. The target population is the group to which results should generalise. The source population is the subset that is eligible and reachable. The study group is the smaller subset that meets eligibility criteria and consents to participate.
Defining the Three Populations
- Target population: the population to which you want results to apply. Stating this explicitly helps ensure that conclusions remain practical and relevant.
- Source population: representative of the target population and consisting of eligible subjects from whom the study group is drawn. The setting and location of the source population should be described.
- Study group: the actual subjects who fit inclusion/exclusion criteria and agree to participate. Volunteers are unavoidable; how well they represent the source and target populations must be considered when extrapolating results.
Unit of Concern: Individuals or Clusters?
An early decision is the level at which the intervention will be applied. Some interventions can only be delivered to groups (e.g., medication added to drinking water; a school-wide curriculum). When the intervention is applied at the group level and the outcome is measured at the group level, this is a group-level study. When the outcome is measured on individuals within those groups, this is a cluster randomised study, which is discussed in a later section.
Eligibility Criteria
Adequate records should be available to verify previous health history, prior treatments, and any condition relevant to the trial. This is especially important when the intervention may interact with prior therapies.
For trials of therapeutic agents, a precise definition of the disease being treated is essential. Subjects who do not actually have the condition will dilute any treatment effect and reduce statistical power.
For trials of preventive products, healthy subjects are required, and procedures must be in place to confirm and document health status at the start of the trial.
Restricting a trial to subjects most likely to benefit increases statistical power but may limit generalisability. Subjects at high risk for adverse effects should generally be excluded both for ethical reasons and to protect the validity of the safety assessment.
The Width of Eligibility Criteria: A Trade-off
A narrow set of eligibility criteria yields a more homogeneous response and increases statistical power, but reduces generalisability. A broad set increases the pool of potential participants and can reveal subgroup variation, but introduces greater background variability that may hurt overall power. The recommended balance: use criteria reflecting the breadth of subjects who would receive the intervention in real-world practice if it proves effective.
Specifying the Intervention
The nature of the intervention and how it is administered must be clearly defined, with enough detail that another investigator could replicate it. Interventions vary widely:
- Medical interventions (e.g., creatine supplementation, antibiotic regimens)
- Surgical techniques (e.g., single-layer vs. double-layer uterine closure)
- Devices or instruments (e.g., progressive lenses for presbyopia)
- Screening programmes (e.g., PSA testing for prostate cancer)
- Behavioural programmes (e.g., adolescent smoking cessation)
A fixed intervention (one with no flexibility) is appropriate for assessing new products in Phase III trials. A more flexible protocol is appropriate for products that have been in use long enough that some clinical judgment has accumulated. Whenever possible, the initial treatment assignment should remain masked so clinical decisions are not influenced by knowledge of group allocation. Clear instructions are critical when participants administer some or all of the intervention themselves (such as instructions for taking medication). A monitoring system should be in place to verify that the intervention is delivered as planned.
Example 12.2: A Sequential Trial of Creatine in ALS
Groeneveld and colleagues (2003) recruited ALS patients from neuromuscular outpatient clinics in Utrecht and Amsterdam. The two interventions were creatine monohydrate and a matching placebo, designated A and B. An independent physician, masked to assignment, instructed the research pharmacist which medication to dispense. Patients were seen at 1 month, 2 months, and every 4 months thereafter. Reasons for withdrawal (serious adverse events, withdrawal of consent) were documented; patients who stopped trial medication remained in the intent-to-treat analysis. The careful specification of this masking-and-allocation chain, from the masked physician through the pharmacist to the patient, illustrates how an unambiguous intervention specification protects the integrity of the trial.
Key Takeaways
- RCTs are the gold standard for evaluating interventions because random allocation breaks the link between exposure and unmeasured confounders.
- Clinical research progresses through pre-clinical studies, then Phase 0 (first-in-human), Phase I (safety), Phase II (efficacy signal), Phase III (pivotal efficacy), and Phase IV (post-marketing surveillance).
- The CONSORT 2010 statement structures both trial design and reporting; trial registration is required by major journals.
- Three nested populations matter: the target population (who results should apply to), the source population (eligible and reachable), and the study group (eligible + consenting).
- Eligibility criteria balance internal validity against generalisability. Recommend using criteria that mirror who would receive the intervention in real practice.
- Interventions must be specified with the precision needed for replication, and a monitoring system should verify delivery.
1. Phase II clinical trials are primarily designed to:
2. The target population in an RCT refers to:
3. A non-inferiority trial differs from a standard superiority RCT in that the null hypothesis is that:
✦ Pass the knowledge check with 100% to continue
Allocation, Outcomes & Sample Size
⏱ Estimated reading time: 22 minutes
Introduction and Overview
An earlier section covered the up-front decisions: what an RCT is, what phase of clinical research it falls under, what the trial design is meant to test, and who gets included. This section turns to the four central design decisions that follow: how the outcome is measured, how large the sample needs to be to detect the effect, how participants are allocated to groups, and who is masked from the allocation. Each of these is a place where a poorly run RCT can lose its inferential advantage.
Learning Objectives
- Identify primary and secondary outcomes that are clinically relevant and measurable.
- Understand the inputs to sample size calculations and how cluster randomisation, sequential designs, and adaptive designs change them.
- Distinguish simple, stratified, cross-over, factorial, cluster, split-plot, and multicentre allocation strategies.
- Compare single, double, and triple blinding and identify the bias each is designed to prevent.
Measuring the Outcome
A controlled trial should be limited to one or two primary outcomes and a small number (one to three) of secondary outcomes. Having too many outcomes leads to multiple-comparisons problems (covered in a later section) and risks inflating the false-positive rate. Composite outcomes, which combine several events into a single measure, are sometimes used but remain controversial; for our purposes a limited number of primary and secondary hypotheses is preferred.
Outcome Scales
Dichotomous Outcomes
The outcome is yes/no: occurrence of disease, death, recovery. This is the most common type in medical trials. Dichotomous outcomes generally require larger sample sizes than continuous outcomes for the same effect size, because they convey less information per subject.
Results should be reported in both absolute terms (risk difference, number needed to treat) and relative terms (risk ratio).
Continuous Outcomes
The outcome is measured on a numeric scale: blood pressure, FEV1, quality-of-life score. Continuous outcomes are often more statistically efficient than dichotomous ones because each subject contributes more information.
Adjustment for baseline (pre-intervention) values can substantially improve precision when the correlation between baseline and follow-up exceeds 0.5.
Time-to-Event Outcomes
Survival analysis, time to disease occurrence, time to relapse. These designs can be more powerful than simple occurrence-or-not in a defined follow-up window because they use all of the timing information. The accuracy of the actual time of event matters, although Korn and colleagues note that non-differential errors in event timing rarely have a major impact on treatment-effect estimation.
Choosing Clinically Relevant Outcomes
Outcomes that can be assessed objectively are preferred, but sometimes subjective outcomes are unavoidable (e.g., self-reported symptoms). When the outcome is not assessed by a near-gold-standard procedure, the impact of the intervention on the true outcome may differ from the impact on the surrogate. Intermediate outcomes (e.g., antibody titres in a vaccine trial) can illuminate mechanism but should not replace clinically relevant primary endpoints. Clinically relevant outcomes typically include:
- Diagnosis of a particular disease: requires a clear case definition
- Mortality: objective but still requires criteria for cause and time of death
- Severity scores: difficult to develop reliably
- Objective clinical measures (rectal temperature, blood biomarkers)
- Quality-of-life and other patient-reported outcome measures
Sample Size
The size of a trial is determined through formal sample-size calculations that incorporate the estimated intervention effect, Type I error rate, and Type II error rate. Power is conventionally set to 90 percent. The sample sizes do not need to be equal in both arms.
Sample Size for Cluster Randomised Trials
When subjects are randomised in clusters (families, schools, clinics), the analysis must account for within-cluster similarity, summarised by the intra-cluster correlation coefficient (rho or ICC) and the cluster size (m). The required sample size is inflated by the design effect:
The design effect tells you how many times larger a clustered sample must be to carry the same information as a simple random sample of unrelated individuals. The reason is that people in the same cluster tend to resemble one another, so a second person from a cluster you have already sampled adds less that is new than a wholly independent person would.
Even when rho is small, large clusters cause substantial inflation. Notably, the power of a cluster trial does not increase appreciably once the number of subjects per cluster exceeds 1/rho, so adding more individuals to existing clusters yields diminishing returns. Adding more clusters is often more efficient than adding more individuals per cluster.
Sequential and Adaptive Designs
A sequential design (sometimes called a monitored study) allows hypothesis tests to be conducted on a number of occasions as data accumulate. Sample size is not fixed; instead, prespecified stopping rules halt the trial when efficacy, harm, or futility becomes clear. Sequential designs can be efficient but tend to lack power on a per-subject basis. Stopping early for benefit can produce overestimates of treatment effect, although the bias is often modest. Interim analyses should not be conducted unless the trial design accommodates them.
Adaptive designs allow the trial design to change as the study progresses. The most common adaptation is modifying the second-stage sample size based on first-stage power. Other adaptations include dropping or adding treatment arms, changing the primary endpoint, or even switching from non-inferiority to superiority. Outcome-adaptive designs use accumulating evidence to assign more subjects to the better-performing intervention (e.g., “play-the-winner”); these are appropriate only when the result of the intervention is identifiable shortly after treatment. The platform trial, an extension that evaluates multiple treatments simultaneously against a shared control under a master protocol, adding and dropping arms over time, has become an important variant in oncology and infectious-disease research (Berry, Connor, & Lewis, 2015).
Recruitment time matters. If season influences treatment response, the recruitment window should span a full calendar year. Loss to follow-up, non-compliance, and competing risks should be anticipated and the sample size adjusted upward to preserve power.
The three classic outcome scales (continuous / dichotomous / time-to-event) each have a one-line analysis in R. Below: simulate an RCT with all three outcome types and run the standard test for each.
set.seed(341)
n <- 200
arm <- factor(rep(c("control", "treatment"), each = n/2))
## (1) Continuous outcome: blood pressure
sbp <- rnorm(n, mean = ifelse(arm == "treatment", 128, 135), sd = 10)
t.test(sbp ~ arm)
## (2) Dichotomous outcome: cure (yes/no)
cure <- rbinom(n, 1, prob = ifelse(arm == "treatment", 0.45, 0.30))
tab <- table(arm, cure)
chisq.test(tab)
prop.test(table(arm, cure)) # with 95% CI for the difference
## (3) Time-to-event outcome: relapse (with right censoring)
library(survival)
time <- rexp(n, rate = ifelse(arm == "treatment", 1/24, 1/14))
event <- rbinom(n, 1, 0.7)
survdiff(Surv(time, event) ~ arm) # log-rank test
One file, three deliverables. Continuous → t.test; dichotomous → chisq.test/prop.test; time-to-event → survdiff. Pre-specifying the test statistic in your protocol, before unblinding, is a key defence against p-hacking.
R Reflect on what you just ran
Use the questions below to interpret the actual numbers from your three RCT analyses. Look at your console output before answering.
1. From t.test(sbp ~ arm), report the difference in mean SBP between treatment and control and the p-value. Did the treatment lower SBP by roughly the 7 mmHg effect that was simulated?
2. From chisq.test(tab) (cure outcome), what p-value did you get? The simulated cure probabilities were 0.45 vs. 0.30 - did your observed proportions land near those values?
chisq.test(tab) on the cure outcome gives p typically around 0.01–0.05 with observed cure proportions close to the simulated 0.45 (treatment) vs. 0.30 (control), usually 0.42–0.48 and 0.28–0.32 depending on seed. The 15 percentage-point gap is statistically detectable at n = 100 per arm, in line with the simulated truth.3. From survdiff(Surv(time, event) ~ arm), report the chi-square statistic and p-value. Given that the simulated mean relapse times were 14 days (control) vs. 24 days (treatment), does the log-rank result agree with what you would expect?
Allocation of Study Subjects
Once enrolled, subjects must be allocated to interventions. Formal randomisation is the strongest method; without it, bias is very likely to distort findings (Schulz & Grimes, 2002a). Random allocation should occur as close to the start of the intervention as possible to minimise withdrawals between assignment and treatment, and concealment of the upcoming allocation from those enrolling participants is essential to prevent selection bias (Schulz & Grimes, 2002b).
It helps to keep two ideas apart. Allocation concealment protects the moment of assignment: it hides the upcoming allocation so that whoever enrols a participant cannot steer sicker or healthier people toward one arm. Blinding, discussed later, protects what happens after assignment, so that knowing the allocation does not colour treatment, behaviour, or outcome assessment. A trial can conceal allocation even when blinding is impossible, as in a trial of surgery versus physiotherapy.
Alternatives to Randomisation
Historical Controls and Systematic Assignment
Historical control trials compare outcomes after an intervention with outcomes from a pre-intervention period. For validity, four conditions must hold: predictable outcome, complete and accurate databases, constant diagnostic criteria, and no environmental changes for subjects. Rarely are all four met. Blinding is impossible in this design.
Systematic assignment (every other subject) can be reasonable in field settings (e.g., a vaccine clinic) and is often as effective as randomisation when outcome assessment is blinded. Randomise the very first subject's assignment to avoid predictability. Never give the intervention to the first half of subjects and the comparison to the second half, which introduces secular confounding.
Pre-generating and version-controlling the allocation sequence is one of the cleanest defences against allocation-concealment failures. Below: simple, stratified, and permuted-block randomisations.
set.seed(202509) # commit this seed before enrolment opens
N <- 120; arms <- c("Drug", "Placebo")
## (1) Simple randomisation
simple <- sample(arms, N, replace = TRUE)
table(simple)
## (2) Stratified by site (3 sites, 40 each)
strat <- unlist(lapply(c("Site_A", "Site_B", "Site_C"), function(s)
paste(s, sample(rep(arms, 20)), sep = ":")))
## (3) Permuted-block randomisation (block size 4)
block <- unlist(replicate(N/4, sample(rep(arms, 2))))
head(block, 12) # every block of 4 has 2 of each arm
## Save the list - this is your audit trail
# write.csv(data.frame(id = 1:N, arm = block), "allocation_list_2025-09-01.csv",
# row.names = FALSE)
Why pre-generate? If the allocation is generated only when each patient enrols, an unblinded coordinator could (consciously or not) game who enrols when. A locked, pre-generated list removes the temptation entirely. The block design also prevents prolonged imbalance during early enrolment.
R Reflect on what you just ran
Use the questions below to interpret the allocation lists you just generated. Compare the output of table(simple) and head(block, 12) before answering.
1. Look at table(simple). With simple randomisation, did you get exactly 60 Drug and 60 Placebo? Why is some imbalance expected under simple randomisation but not under permuted-block randomisation?
2. From head(block, 12) (the first three blocks of size 4), confirm by eye that every block of 4 contains exactly 2 Drug and 2 Placebo. Why is this an attractive feature if a safety committee plans an interim analysis at n = 60?
head(block, 12) confirms two D's and two P's in each block (e.g., DPDP, PDDP, DPPD). At an n=60 interim analysis the data-safety committee will see exactly 30 Drug and 30 Placebo participants, and balanced exposure makes the interim treatment-effect estimate efficient and unbiased by allocation imbalance. Without blocked randomisation, the interim might see 35 vs. 25 and produce a noisier effect estimate, complicating stopping decisions.3. Why MUST set.seed(202509) be committed alongside the analysis script before enrolment opens? What happens to reproducibility (and to your audit trail) if it is forgotten?
set.seed(202509) before enrolment opens guarantees that anyone with the script and the data can re-run the analysis and get the same allocation, statistical tests, and CIs. Without the seed pre-committed, the allocation could be reproduced by trial-and-error to favour any particular outcome, destroying the audit trail and inviting accusations of post hoc data manipulation. Regulators and journals increasingly require pre-registration and seed commitment because it is the cheapest, strongest defence against accidental or motivated post hoc shifting of results.The Family of Random Allocation Designs
Random allocation does not mean haphazard allocation. A formal process (a computer-based random number generator, sealed-envelope randomisation, or even a coin toss) must be used (Schulz & Grimes, 2002a). Click any design to see its features and an example.
Randomisation
Randomisation
Design
Design
Randomisation
Design
Cluster Randomisation: Why It Is Less Efficient
Cluster randomised trials are statistically less efficient than individually randomised trials of the same total sample size (Campbell, Elbourne, & Altman, 2004). The clustering of subjects within groups must be accounted for in the analysis. The best follow-up scenario is to monitor all individuals for the duration of the study; if not, following a randomly selected cohort is the next most powerful approach. In some settings, repeated cross-sectional samples within each cluster must be used. Matched-cluster designs may be appropriate when the number of clusters is small, although “breaking the matches” can sometimes improve statistical efficiency.
Example 12.3: A Cluster RCT of Adolescent Smoking Cessation
Dalum and colleagues (2012) randomised 22 continuation schools in Denmark by coin toss to deliver a smoking cessation intervention or to act as control. The randomisation was blocked so that each county contained both intervention and control schools, balanced across school types (commercial vs. social-and-health). Smoking status was self-reported in surveys at baseline, week 11/2005 (short-term), and week 11/2006 (long-term). Analyses were intent-to-treat, with school as a random factor in logistic regression.
Why was cluster randomisation appropriate here? Because the intervention was delivered at the school level, individual randomisation would have introduced contamination, since intervention and control students in the same hallway would influence each other.
Example 12.4: A 2×2×2 Factorial Caesarean-Section Trial
The CAESAR trial (CAESAR study collaborative group, 2010) randomised women aged >15 undergoing their first Caesarean section to three independent factors: single- vs. double-layer uterine closure, closure vs. non-closure of the peritoneum, and liberal vs. restricted use of a subsheath drain. Telephone randomisation with a minimisation algorithm balanced participating centre, labour status, and pregnancy multiplicity. The primary outcome was maternal infectious morbidity (any of: antibiotic use for febrile morbidity, endometritis, or treated wound infection). The 3,500 women required to detect a 12% to 9% reduction with 80% power illustrate how factorial designs efficiently address multiple research questions in one trial.
Example 12.5: A Split-Plot Shoulder-Pain Trial
Watson and colleagues (2008) conducted a pragmatic split-plot trial across UK general practices. Physicians in 91 practices (the whole plot) were randomised to additional training in shoulder injection or to no additional training. Within the practices, 215 patients with acute shoulder pain were then randomised to receive either a corticosteroid or a lignocaine injection (the split plot). The main outcome was the British Shoulder Disability Questionnaire score. Notice how this design lets the trial answer two questions simultaneously: does training help, and does corticosteroid outperform lignocaine?
Example 12.6: A Multicentre Cross-Over Trial of Progressive Lenses
Boutron and colleagues (2008) compared two generations of progressive lenses for presbyopia at five primary-care optical dispensaries. 127 patients aged 43–60 were randomised to wear one lens for four weeks, then cross over to the other for four weeks, blinded to the lens sequence. Patients and the statistical analyst were both blinded; all equipment was assembled in one laboratory to ensure consistency. The primary outcome was patient preference at week 8. The cross-over design is appropriate here because lens preference is reversible and short-acting; conditions are stable, and switching has no carry-over.
Multicentre Trials
If an adequate sample is not available at one site, a multicentre trial is required. Within-centre and between-centre variances must be accounted for in design and analysis. Multicentre trials enhance generalisability (because of the broader geographic and clinical reach) and create opportunities to detect interaction effects across sites. For statistical efficiency, the number of subjects per centre should be approximately equal. Example 12.6 was conducted across five centres.
Masking (Blinding)
Blinding (or masking) refers both to the methodological principle of withholding information from individuals to prevent bias and to the specific procedures used to do so (Schulz & Grimes, 2002c). Terms can be used inconsistently in the literature, so the specific masking mechanisms always need to be described, and ideally pilot-tested in larger trials.
Figure 12.2. The three nested layers of blinding. Each additional layer is added on top of the prior layers, prevents an additional source of bias, and is harder to achieve in practice.
What Each Level of Blinding Prevents
Single-Blind: Participant Unaware
In a single-blind study the participant does not know which intervention they are receiving. This helps:
- Reduce response bias (subjects reporting symptoms differently based on what they think they are getting)
- Prevent the placebo effect
- Equalise differential attrition and non-compliance
- Reduce co-intervention bias and follow-up bias
Double-Blind: Participant + Treatment/Outcome Personnel Unaware
In a double-blind study, both the participants and the people administering the intervention or assessing the outcome are unaware of allocation. This adds protection against:
- Patient–provider interaction effects on the placebo response
- Differential attrition, non-compliance, or co-intervention driven by clinician knowledge
- Selective decisions and referrals based on treatment knowledge
- Observer bias and diagnostic bias when assessing outcomes
Triple-Blind: + Data Analysts Unaware
In a triple-blind study, the people analysing the data are also unaware of group identity (often coded as A vs. B). This is designed to ensure unbiased analytic decisions (choice of subgroups, handling of outliers, modelling decisions) that might otherwise be subtly influenced by knowledge of which group is the “new” treatment.
The success of blinding should be evaluated rather than assumed. Methods for assessing blinding success have been published.
The Role of Placebos
A placebo is a product indistinguishable from the active intervention, administered to the comparison group. In drug trials, the placebo is often the vehicle without the active ingredient. Even apparently inert placebos can have positive or negative effects (e.g., a placebo vaccine without antigen can still induce some immunity through adjuvant). These issues should be addressed before the trial begins. In some situations, blinding cannot be achieved with placebos alone, but masking should be implemented wherever feasible.
The Take-Away on Blinding
Each additional layer of blinding prevents a different bias, but each is harder to achieve in practice. Whenever feasible, design the trial so that as many sources of error as possible are prevented from the start; this also reduces the impact of any differential errors that do arise. The goal is not maximum blinding for its own sake, but rigorous masking matched to the biases that most threaten the specific trial.
Key Takeaways
- Limit a trial to 1–2 primary outcomes and 1–3 secondary outcomes; report effects in both absolute and relative terms when outcomes are dichotomous.
- Sample size depends on the expected effect, Type I and Type II error rates, and the outcome scale. Cluster randomisation requires inflation by the design effect 1 + ρ(m − 1).
- Sequential designs allow stopping for efficacy, harm, or futility but tend to lack power per subject; adaptive designs offer flexibility but require careful pre-specification.
- Random allocation is the strongest assignment method. The major design types are simple, stratified, cross-over, factorial, cluster, and split-plot, plus multicentre extensions.
- Single-blind protects against participant response bias; double-blind adds clinician and outcome-assessor blinding; triple-blind adds analyst blinding. Each layer prevents a different category of bias.
1. In a cluster randomised trial with intra-cluster correlation ρ = 0.05 and average cluster size m = 41, the design effect (sample-size inflation factor) is:
2. A cross-over design is appropriate when:
3. The chief additional bias prevented by triple-blinding (over and above double-blinding) is:
✦ Pass the knowledge check with 100% to continue
Conduct, Analysis & Special Topics
⏱ Estimated reading time: 22 minutes
Introduction and Overview
Earlier sections settled the design. This section walks through what happens once the trial is running: how to track participants through follow-up, how to analyse the data (including the intention-to-treat principle and the special case of vaccine efficacy), and how to report the trial honestly via the CONSORT checklist. The reporting framework here connects directly back to the integrity material from an earlier lesson.
Learning Objectives
- Implement effective follow-up and compliance monitoring during a trial.
- Distinguish intent-to-treat from per-protocol analysis and choose the appropriate approach.
- Identify the sources of multiple comparisons in RCTs and apply the Bonferroni adjustment.
- Compute and interpret direct, indirect, and total vaccine efficacy.
- Use the CONSORT 2010 checklist to plan and report a randomised trial.
Follow-Up and Compliance
One of the most important practical issues is ensuring that all groups are followed rigorously and equally. The follow-up period must be long enough to capture all outcomes of interest. Some loss is inevitable through drop-out or non-compliance; for trials with long follow-up, the status of all subjects should be ascertained at regular intervals. The CONSORT statement strongly recommends a flow diagram showing participant numbers at allocation, intended intervention, protocol completion, and outcome assessment.
Strategies to Minimise Loss and Maximise Compliance
Frequent contact, such as reminder messages, newsletters, and study updates, reduces attrition. Incentives may be provided, including study-related information that participants would not otherwise have, or public recognition of their contribution (subject to confidentiality).
For participants who drop out, information may still be available through routine databases if the participant consents. Documenting reasons for withdrawal allows comparison of withdrawn and remaining subjects, helping characterise potential bias.
Compliance can be assessed through interviews, biological samples (drug or metabolite levels), or indirect indicators such as collecting empty pill containers, vials, and packaging. Compliance data are essential for interpreting the difference between intent-to-treat and per-protocol results.
Statistical Methods and Analysis
Outcomes might be analysed on a continuous scale, as categorical (often dichotomous) data, or as time-to-event measurements. Time-to-event analyses can have greater power than simple occurrence-or-not in a defined window. Whatever the analysis, results should report both the effect size and its precision (typically a 95 percent confidence interval), and dichotomous outcomes should appear in both absolute (risk difference) and relative (risk ratio) terms.
Intent-to-Treat vs. Per-Protocol Analysis
This distinction is one of the most important in RCT analysis. Click the tabs to compare the two approaches.
Intent-to-Treat (ITT) Analysis
All subjects assigned to a specific intervention are analysed in that group, regardless of whether they completed the study or complied with the protocol. ITT yields a conservative estimate of the intervention effect; if anything, it will under-estimate the maximum potential benefit by including non-compliers and dropouts in the “treated” arm.
Why use it: ITT estimates the expected response when the intervention is rolled out in similar populations, because in real-world use some non-compliance and loss to follow-up are inevitable. ITT preserves the benefits of randomisation and is the recommended primary analysis.
Per-Protocol (PP) Analysis
Only subjects who complied with and completed the study as specified in the protocol are analysed. PP yields an estimate of effect under ideal compliance.
Why caution is needed: First, non-compliance is rarely a random event; non-compliers are often systematically different from compliers, so the PP estimate is likely biased. Second, future use of the intervention will involve some non-compliance, so an “assuming 100% compliance” estimate is unrealistic. PP should be reported alongside ITT but should not replace it.
Stating Numbers and Compliance Is Essential
Whichever analysis is primary, the number of subjects in each group, and whether or not they complied, must be reported. If there are considerable losses or major adherence problems, Hernán and Hernández-Díaz (2012) suggest using inverse probability weighting to reduce potential bias.
Both ITT and PP analyses are typically reported. ITT is the primary analysis for efficacy claims; PP is a secondary or sensitivity analysis to show the effect among those who actually adhered.
Baseline Comparison and Covariate Adjustment
Analysis usually starts with a baseline comparison of group characteristics as a check on randomisation. This is an assessment of comparability rather than a statistical significance test. Differences between groups, even if not statistically significant, should be noted and may justify covariate adjustment.
For dichotomous outcomes, adjustment for covariates is recommended, ideally for strong predictors identified a priori; failing that, for variables predictive of the outcome in the trial data. Adjustment can substantially increase power or reduce required sample size. Adjustment for non-confounders does little harm provided they are not intervening variables (mediators).
For continuous outcomes, controlling for baseline (pre-intervention) values can substantially improve precision. This can be done either by analysing the change score (post minus pre) or by including baseline as a covariate. Either approach gains power, particularly when the baseline-to-follow-up correlation exceeds 0.5.
The Multiple Comparisons Problem
Multiple comparisons in RCTs arise from three sources: examining multiple outcomes, examining multiple subgroups, and performing periodic interim analyses during the trial. The problem is that the experiment-wise (family-wise) error rate is much larger than the per-test rate, making spurious “significant” findings increasingly likely as the number of tests grows.
The Bonferroni adjustment is the simplest fix: divide the desired experiment-wise alpha by the number of comparisons. It is conservative; less conservative procedures (Holm, Hochberg, Benjamini-Hochberg) are available in standard texts.
The Subgroup Analysis Trap
It is tempting to evaluate many subgroups to see whether the intervention works in any of them. This should be avoided: only subgroup analyses planned a priori should be carried out; data-driven subgroup analyses generate spurious associations at alarming rates. The recommended way to test whether an intervention's effect varies by subgroup is a single overall interaction test. Note that detecting interactions reliably typically requires a sample size at least four times larger than detecting the overall main effect; effect sizes for interactions need to be roughly twice the magnitude of the main effect to be detected with similar power.
Sequential Analyses: When to Stop
Sequential design studies plan periodic analyses throughout the trial to allow early stopping for one of three reasons:
- Clear (and statistically significant) evidence of the superiority of one intervention over the other
- Convincing evidence of harm from the intervention (regardless of statistical significance)
- Little likelihood that the trial will produce evidence of an effect even if completed (futility)
Interim analyses must be pre-specified in the trial design; ad-hoc “peeking” inflates the false-positive rate and is regarded as a serious methodological breach.
Vaccine Trials: Direct, Indirect, and Total Efficacy
Standard RCT designs need modification when the intervention is a prophylactic against a communicable organism. The reason: an effective vaccine has effects on the vaccinated and on the unvaccinated, because vaccination reduces transmission. This means study subjects are not independent, an effect Hudgens and Halloran (2008) call interference.
Why Standard Direct Efficacy Estimates Are Misleading
The standard individually randomised, placebo-controlled trial of a vaccine yields an estimate of vaccine efficacy that is confounded by the proportion of the study population vaccinated. Two identical trials in populations with different transmission levels, or different vaccination coverage levels, will report different vaccine efficacy estimates even if the underlying biological protection is the same.
Three Measures of Vaccine Efficacy
To get a fuller picture, epidemiologists distinguish three measures, each requiring information from at least two populations with different vaccination coverage. Use Iv for the incidence in the vaccinated and Inv for the incidence in the unvaccinated; subscript A denotes the higher-coverage population, B the lower-coverage population.
Example 12.7: Cholera Vaccine Effects in Bangladesh
Ali and colleagues (2005) and Hudgens and Halloran (2008) analysed an individually randomised, placebo-controlled trial of killed oral cholera vaccines in residential areas (baris) in Bangladesh. Two groups were compared: Group A (more than 50% coverage) and Group B (less than 28% coverage). The first-year risks per 1,000 were: RnvB = 7.01, RvB = 2.66, RnvA = 1.47, RvA = 1.27, RB = 4.13, RA = 1.34. Here each first-year risk R plays the role of the incidence I in the equations above, so the two symbols refer to the same quantity.
Direct effect in the high-coverage group A: VEd = (1.47 − 1.27) / 1.47 = 0.14. Looking at A alone you might conclude the vaccine has little effect.
Direct effect in the low-coverage group B: VEd = (7.01 − 2.66) / 7.01 = 0.62, so the vaccine reduces risk by 62% in the unprotected population.
Indirect effect (in the unvaccinated): (7.01 − 1.47) / 7.01 = 0.79, so herd immunity reduced the unvaccinated risk by 79%.
Total relative effect: (7.01 − 1.27) / 7.01 = 0.82. Overall effect: (4.13 − 1.34) / 4.13 = 0.68.
The lesson is striking: limiting analysis to the high-coverage population alone (Group A) would have suggested the vaccine barely worked. Looking across populations with different coverage reveals the dominant role of indirect effects.
Designing Vaccine Trials to Estimate All Three Measures
To estimate VEd, VEind, and VEtot, the design must include at least two comparable populations with different vaccination coverage. The recommended approach: in population A, randomise individuals to vaccine or placebo. In population B, leave everyone unvaccinated (or assign a lower coverage). The two populations must be (a) comparable in characteristics that affect the outcome, especially baseline transmission level, and (b) physically separated so subjects do not intermix.
An alternative when finding two similar populations is impractical is to exploit natural clustering, for example randomly assigning vaccination to half the children in a geographic area, then study spread within schools, recording the proportion of children at each school who were vaccinated. Riggs and Koopman (2005) and Longini and colleagues (1998, 2002) describe how to design and analyse such trials.
Reporting: The CONSORT 2010 Checklist
The CONSORT 2010 statement is the dominant reporting guideline for parallel-group randomised trials. Its 25 items align with the design and conduct topics covered throughout this lesson. The table below highlights items most relevant to the topics we have studied.
| Section / Item | What to Report |
|---|---|
| 1a Title | Identify the study as a randomised trial in the title (improves “searchability”). |
| 2a–b Background & objectives | Scientific rationale and explicit hypotheses or objectives. |
| 3a Trial design | Type of design (parallel, factorial, cross-over, cluster, split-plot) and allocation ratio. |
| 4a–b Participants | Eligibility criteria; settings and locations of data collection. |
| 5 Interventions | Sufficient detail for replication, including how and when administered. |
| 6a–b Outcomes | Pre-specified primary and secondary outcomes and how they were assessed. |
| 7a–b Sample size | How sample size was determined; explanation of any interim analyses or stopping guidelines. |
| 8–10 Randomisation | How the allocation sequence was generated, what type of restriction (blocking) was used, the mechanism for implementing allocation, and who did what at enrolment. |
| 11a–b Blinding | Who was blinded after assignment (participants, providers, outcome assessors); description of intervention similarity if relevant. |
| 12a–b Statistical methods | Methods for primary and secondary outcomes; methods for subgroup and adjusted analyses. |
| 13a–b Participant flow | For each group, numbers randomly assigned, receiving intended treatment, and analysed; losses and exclusions with reasons. A flow diagram is strongly recommended. |
| 14a–b Recruitment | Dates of recruitment and follow-up; reasons the trial ended or was stopped. |
| 15 Baseline data | Table of baseline demographic and clinical characteristics for each group. |
| 16 Numbers analysed | For each group, denominators in each analysis and whether by originally assigned groups. |
| 17a–b Outcomes & estimation | Effect size and precision (e.g., 95% CI); for binary outcomes, both absolute and relative effect sizes. |
| 18 Ancillary analyses | Other analyses (subgroup, adjusted), distinguishing pre-specified from exploratory. |
| 19 Harms | All important harms or unintended effects in each group. |
| 20–22 Discussion | Limitations (bias, imprecision, multiplicity); generalisability; interpretation balanced against benefits and harms. |
| 23–25 Other | Trial registration number; where the full protocol can be accessed; sources of funding and role of funders. |
CONSORT extensions exist for cluster randomised trials, non-inferiority and equivalence trials, non-pharmacological treatments, herbal interventions, and pragmatic trials. Up-to-date references are available at www.consort-statement.org.
Why Reporting Quality Matters
Poor reporting prevents readers from judging the reliability and validity of trial findings, prevents extraction for systematic reviews, and is associated with biased estimates of treatment effects. Inadequate reporting is therefore a methodological problem, with consequences beyond style. Following CONSORT during planning, and through to write-up, helps ensure that the methodological choices required for transparent reporting are actually made.
Reflection: Designing Your Own RCT
Imagine you have been funded to design a randomised controlled trial of a new mobile-app-based behavioural intervention to reduce daily smoking among young adults aged 18–25 living in rural communities. Briefly outline how you would address: (1) target population vs. source population vs. study group; (2) choice of allocation strategy (individual? cluster? factorial?); (3) choice of comparator (no app? sham app? existing standard care?); (4) how you would handle the analysis if compliance with the app turned out to be poor. Identify at least one trade-off you would face and explain how you would resolve it.
Minimum 20 characters required.
Key Takeaways
- Rigorous and equal follow-up of all groups is essential. Document compliance and reasons for withdrawal; a CONSORT flow diagram is strongly recommended.
- Intent-to-treat analysis preserves randomisation and gives a conservative, real-world estimate; per-protocol analysis estimates the effect under ideal compliance and should be reported alongside ITT, not in place of it.
- Multiple comparisons inflate the experiment-wise error rate. The Bonferroni adjustment (α / k) is the simplest correction. Subgroup analyses should be pre-specified and tested via interaction terms; data-driven subgroup analyses generate spurious findings.
- For prophylactic vaccines, three measures matter: direct (VEd), indirect or herd-immunity (VEind), and total (VEtot) efficacy. Estimating all three requires data from at least two populations with different coverage.
- The CONSORT 2010 checklist of 25 items structures both the planning and the reporting of randomised trials. Following CONSORT from the design stage through to write-up is what makes transparent reporting achievable.
1. The major reason that intent-to-treat analysis is preferred as the primary analysis is that:
2. In a vaccine trial, the indirect vaccine efficacy (VEind) is estimated by:
3. With four pre-specified primary outcome comparisons and a desired family-wise error rate of 0.05, the Bonferroni-adjusted significance threshold for each comparison is:
✦ Complete the reflection and pass the knowledge check with 100% to continue
Knowledge Check & Final Assessment
⏱ Estimated time: 25 minutes
Bringing It All Together
This module walked through every major analytic study design in epidemiology, from cross-sectional surveys to the randomised controlled trial. The first part revisited the three core observational designs (cross-sectional, cohort, and case-control) alongside ecological studies and the place of systematic reviews at the top of the evidence hierarchy. The point of returning to these designs after the analytic lessons on frequency and association is that the choice of design is rarely a separate decision from the measure of association you want to estimate or the practical constraints you face. Reporting guidance for each design family is now codified: STROBE for observational studies and CONSORT for randomised trials.
The second part moved into the family of hybrid designs that mix and match the strengths of the classics: case-crossover and self-controlled case-series (which use the case as their own control across time), case-case and case-case-control designs (which substitute one case series for the traditional control group), case-cohort designs (which let one subcohort serve multiple outcomes), case-only designs (which estimate interaction without a separate control series), and two-stage sampling (which layers cheap and expensive measurement strategically). The unifying logic is efficiency: each hybrid responds to a specific limitation of cohort or case-control designs, such as recall bias, between-person confounding, costly biomarker measurement, or rare outcomes, by changing the comparison group or the sampling rule.
The third part crossed from observational into experimental epidemiology: the rationale for randomised controlled trials, the phases of clinical research, the mechanics of allocation and masking, sample size with the inflation factor for cluster trials, and the analytic and reporting standards (intent-to-treat, CONSORT 2010) that govern modern trial conduct. The shift is conceptually small, since the investigator now assigns the exposure, but methodologically transformative, because randomisation is what gives trials their causal authority. The final assessment below asks you to treat the whole portfolio as one toolkit: each design buys you something and costs you something else, and the skill this module builds is matching the design to the question. A later lesson turns from designing studies to appraising them, examining the systematic threats to validity that determine whether any design delivers on its promise.
Key Takeaways from this lesson
- The observational vs. experimental split turns on a single question: who assigns the exposure? Descriptive vs. analytic is a different axis, namely whether comparison groups are used.
- Cohort studies follow exposed and unexposed groups forward and yield incidence and risk/rate ratios; case-control studies sample on outcome, are efficient for rare diseases, and are limited to odds ratios and vulnerable to recall and selection bias.
- Ecological studies analyse group-level data and are useful for policy evaluation, but inferences to individuals risk the ecologic fallacy; systematic reviews and meta-analyses synthesise across studies under transparent, reproducible methods.
- Case-crossover studies and self-controlled case-series use each case as their own control across time, which is ideal for transient exposures and acute outcomes; control-period selection depends on whether the event alters future exposure.
- Case-case, case-case-control, case-cohort, and case-only designs each replace or reuse a comparison group to gain efficiency, and each carries an assumption (shared surveillance source, outcome-independent subcohort, gene–environment independence) that must be defended.
- Two-stage sampling layers cheap stage-1 data on everyone with detailed stage-2 measurement on a strategically chosen subsample, an efficient way to control measurement cost in any base design.
- Randomisation is what makes the RCT the gold standard for causal inference: it balances measured and unmeasured confounders in expectation, and trial design begins with explicit target, source, and study populations and a precisely specified intervention.
- Allocation strategies (simple, stratified, cross-over, factorial, cluster, split-plot, multicentre) and masking (single, double, triple) each prevent specific biases; cluster trials must inflate sample size by 1 + ρ(m − 1).
- Intent-to-treat is the primary analysis because it preserves randomisation; vaccine trials distinguish direct, indirect, and total efficacy; and the CONSORT 2010 checklist structures both planning and reporting.
Reflection
A newly marketed oral medication for a common chronic condition is suspected of transiently raising the risk of a serious acute cardiac event in the days after each dose, and a health authority is also considering a pharmacist-led counselling programme intended to reduce that risk. Drawing on all three parts of this module, lay out a sequence of study designs that would move this question from first signal to actionable evidence. For each design you propose, state the question it answers, why it beats the alternatives for that question, the key assumption or limitation you would have to defend, and the measure of association it yields. Your sequence should include at least one classical observational design, one hybrid design, and one controlled design.
Minimum 20 characters required.
Final Knowledge Assessment
Complete the following 15-question assessment, which draws on all nine sections of the module. A score of 100% is required to complete the lesson. You may retake the assessment as many times as needed.
1. In an observational study, the investigator:
2. Which measure of association can be directly calculated from a cohort study but NOT from a case-control study?
3. Recall bias is a particular concern in:
4. A country with high average alcohol consumption also has high rates of liver cirrhosis. Concluding that the individuals who drink the most in that country are the ones with cirrhosis would be an example of:
5. Systematic reviews sit at the top of the evidence hierarchy because they:
6. The case-crossover design controls for time-invariant confounders by:
7. The self-controlled case-series design is most commonly used to study:
8. A case-case study comparing Campylobacter coli with Campylobacter jejuni would NOT be useful for identifying:
9. A major attractive feature of the case-cohort design is that:
10. A two-stage sampling design is most useful when:
11. The defining feature that distinguishes a randomised controlled trial from an observational analytic study is that:
12. In a cluster randomised trial of a school-based health programme with intra-cluster correlation ρ = 0.02 and an average cluster size of 51 students per school, the sample-size inflation factor is:
13. Triple-blinding adds which additional layer of masking compared with double-blinding?
14. The recommended primary analysis for a randomised controlled trial is:
15. The CONSORT 2010 statement is:
✦ Complete every section knowledge check and reflection, including the final reflection above, before submitting