Cohort and Case-Control Study Design: Zero to Hero Tutorial
This tutorial explains how to design cohort and case-control studies for clinical, public health, behavioural, and applied research. You will learn when each design is appropriate, how to define exposure and outcome windows, how to select comparison groups, and how to control bias and confounding.
Table of Contents
- Prerequisites and Background Concepts
- Cohort vs. Case-Control Designs
- Cohort Study Design
- Case-Control Study Design
- Measure Selection
- Bias and Confounding
- Sample Size Planning
- Using DataStatPro
- Worked Examples
- Common Mistakes and How to Avoid Them
- Quick Reference Cheat Sheet
1. Prerequisites and Background Concepts
You should understand:
- Exposure: A risk factor, treatment, behaviour, or condition.
- Outcome: The event or disease being studied.
- Incidence: New cases occurring over time.
- Risk: Probability of an event over a period.
- Odds: Event probability divided by non-event probability.
- Confounding: A third variable distorting the exposure-outcome relationship.
2. Cohort vs. Case-Control Designs
| Feature | Cohort Study | Case-Control Study |
|---|---|---|
| Starts with | Exposure status | Outcome status |
| Direction | Exposure to outcome | Outcome back to exposure |
| Best for | Rare exposures, multiple outcomes | Rare outcomes, long latency |
| Main measure | Risk ratio, rate ratio, hazard ratio | Odds ratio |
| Main challenge | Follow-up and cost | Control selection and recall bias |
3. Cohort Study Design
A cohort study follows exposed and unexposed groups over time to compare outcome occurrence.
3.1 Prospective Cohort
Exposure is measured before outcomes occur.
Strengths:
- Clear temporality.
- Direct incidence estimation.
- Multiple outcomes possible.
Limitations:
- Can be expensive.
- Loss to follow-up can bias results.
- Inefficient for rare outcomes.
3.2 Retrospective Cohort
Existing records define exposure and follow-up.
Use when reliable historical exposure and outcome data are available.
3.3 Basic Cohort Measures
Risk in exposed:
Risk in unexposed:
Risk ratio:
Risk difference:
4. Case-Control Study Design
A case-control study selects participants by outcome status, then compares prior exposure.
4.1 Case Definition
A good case definition specifies:
- Diagnostic criteria.
- Time period.
- Geographic or institutional source.
- Incident or prevalent cases.
Incident cases are usually preferable because they reduce survival bias.
4.2 Control Selection
Controls should represent the exposure distribution in the population that produced the cases.
Control sources:
- Population controls.
- Hospital or clinic controls.
- Neighbourhood controls.
- Registry controls.
4.3 Basic Case-Control Measure
Odds ratio:
In rare diseases, the odds ratio approximates the risk ratio.
5. Measure Selection
| Design | Preferred Measures |
|---|---|
| Prospective cohort | Risk ratio, risk difference, rate ratio, hazard ratio |
| Retrospective cohort | Risk ratio, rate ratio, hazard ratio |
| Case-control | Odds ratio |
| Nested case-control | Odds ratio or incidence density ratio |
Choose measures before analysis and report confidence intervals.
6. Bias and Confounding
Common threats:
- Selection bias.
- Information bias.
- Recall bias.
- Loss to follow-up.
- Confounding by indication.
- Immortal time bias.
Control strategies:
- Restriction.
- Matching.
- Stratification.
- Regression adjustment.
- Sensitivity analysis.
7. Sample Size Planning
Sample size depends on:
- Expected exposure prevalence.
- Expected outcome risk or odds.
- Detectable effect size.
- Ratio of exposed to unexposed or controls to cases.
- Desired power and significance level.
For case-control studies, increasing controls per case can improve power, but gains become small beyond about four controls per case.
8. Using DataStatPro
Use DataStatPro to:
- Build 2 x 2 tables.
- Calculate odds ratios, risk ratios, and confidence intervals.
- Run chi-square or Fisher's exact tests.
- Fit logistic regression models.
- Summarize cohort follow-up and event rates.
- Export publication-ready tables.
9. Worked Examples
Example 1: Prospective Cohort
Follow smokers and non-smokers for 10 years and compare incidence of chronic lung disease. Report risk ratio, risk difference, and confidence intervals.
Example 2: Case-Control Study
Select patients with a rare cancer and matched controls from the same region. Compare prior occupational exposure and report odds ratio.
Example 3: Retrospective Cohort
Use hospital records to compare infection risk among patients who received two different catheter types.
10. Common Mistakes and How to Avoid Them
| Mistake | Why It Matters | Better Practice |
|---|---|---|
| Using prevalent cases for etiologic claims | Survival bias | Prefer incident cases |
| Controls from wrong source population | Selection bias | Select controls from the population that produced cases |
| Measuring exposure after outcome | Reverse causality | Define exposure window before outcome |
| Ignoring loss to follow-up | Biased cohort estimates | Track and compare losses |
| Reporting odds ratio as risk ratio when outcome is common | Exaggerates effect | Use correct measure |
11. Quick Reference Cheat Sheet
| Goal | Best Design |
|---|---|
| Study rare exposure | Cohort |
| Study rare disease | Case-control |
| Estimate incidence | Cohort |
| Study multiple outcomes | Cohort |
| Study long-latency disease efficiently | Case-control |
Key formulas:
Report source population, eligibility criteria, exposure window, outcome definition, confounding strategy, and effect estimate with confidence interval.