Cohort and Case-Control Study Design

Design exposure-outcome studies with appropriate comparison groups.

Quick answer

Cohort and Case-Control Study Design in DataStatPro helps researchers understand the method, choose appropriate assumptions and outputs, and connect the analysis to publication-ready reporting. Design exposure-outcome studies with appropriate comparison groups.

Cohort and Case-Control Study Design: Zero to Hero Tutorial

This tutorial explains how to design cohort and case-control studies for clinical, public health, behavioural, and applied research. You will learn when each design is appropriate, how to define exposure and outcome windows, how to select comparison groups, and how to control bias and confounding.


Table of Contents

  1. Prerequisites and Background Concepts
  2. Cohort vs. Case-Control Designs
  3. Cohort Study Design
  4. Case-Control Study Design
  5. Measure Selection
  6. Bias and Confounding
  7. Sample Size Planning
  8. Using DataStatPro
  9. Worked Examples
  10. Common Mistakes and How to Avoid Them
  11. Quick Reference Cheat Sheet

1. Prerequisites and Background Concepts

You should understand:

  • Exposure: A risk factor, treatment, behaviour, or condition.
  • Outcome: The event or disease being studied.
  • Incidence: New cases occurring over time.
  • Risk: Probability of an event over a period.
  • Odds: Event probability divided by non-event probability.
  • Confounding: A third variable distorting the exposure-outcome relationship.

2. Cohort vs. Case-Control Designs

FeatureCohort StudyCase-Control Study
Starts withExposure statusOutcome status
DirectionExposure to outcomeOutcome back to exposure
Best forRare exposures, multiple outcomesRare outcomes, long latency
Main measureRisk ratio, rate ratio, hazard ratioOdds ratio
Main challengeFollow-up and costControl selection and recall bias

3. Cohort Study Design

A cohort study follows exposed and unexposed groups over time to compare outcome occurrence.

3.1 Prospective Cohort

Exposure is measured before outcomes occur.

Strengths:

  • Clear temporality.
  • Direct incidence estimation.
  • Multiple outcomes possible.

Limitations:

  • Can be expensive.
  • Loss to follow-up can bias results.
  • Inefficient for rare outcomes.

3.2 Retrospective Cohort

Existing records define exposure and follow-up.

Use when reliable historical exposure and outcome data are available.

3.3 Basic Cohort Measures

Risk in exposed:

RE=a/(a+b)R_E = a/(a+b)

Risk in unexposed:

RU=c/(c+d)R_U = c/(c+d)

Risk ratio:

RR=RE/RURR = R_E/R_U

Risk difference:

RD=RERURD = R_E - R_U


4. Case-Control Study Design

A case-control study selects participants by outcome status, then compares prior exposure.

4.1 Case Definition

A good case definition specifies:

  • Diagnostic criteria.
  • Time period.
  • Geographic or institutional source.
  • Incident or prevalent cases.

Incident cases are usually preferable because they reduce survival bias.

4.2 Control Selection

Controls should represent the exposure distribution in the population that produced the cases.

Control sources:

  • Population controls.
  • Hospital or clinic controls.
  • Neighbourhood controls.
  • Registry controls.

4.3 Basic Case-Control Measure

Odds ratio:

OR=a/bc/d=adbcOR = \frac{a/b}{c/d} = \frac{ad}{bc}

In rare diseases, the odds ratio approximates the risk ratio.


5. Measure Selection

DesignPreferred Measures
Prospective cohortRisk ratio, risk difference, rate ratio, hazard ratio
Retrospective cohortRisk ratio, rate ratio, hazard ratio
Case-controlOdds ratio
Nested case-controlOdds ratio or incidence density ratio

Choose measures before analysis and report confidence intervals.


6. Bias and Confounding

Common threats:

  • Selection bias.
  • Information bias.
  • Recall bias.
  • Loss to follow-up.
  • Confounding by indication.
  • Immortal time bias.

Control strategies:

  • Restriction.
  • Matching.
  • Stratification.
  • Regression adjustment.
  • Sensitivity analysis.

7. Sample Size Planning

Sample size depends on:

  • Expected exposure prevalence.
  • Expected outcome risk or odds.
  • Detectable effect size.
  • Ratio of exposed to unexposed or controls to cases.
  • Desired power and significance level.

For case-control studies, increasing controls per case can improve power, but gains become small beyond about four controls per case.


8. Using DataStatPro

Use DataStatPro to:

  • Build 2 x 2 tables.
  • Calculate odds ratios, risk ratios, and confidence intervals.
  • Run chi-square or Fisher's exact tests.
  • Fit logistic regression models.
  • Summarize cohort follow-up and event rates.
  • Export publication-ready tables.

9. Worked Examples

Example 1: Prospective Cohort

Follow smokers and non-smokers for 10 years and compare incidence of chronic lung disease. Report risk ratio, risk difference, and confidence intervals.

Example 2: Case-Control Study

Select patients with a rare cancer and matched controls from the same region. Compare prior occupational exposure and report odds ratio.

Example 3: Retrospective Cohort

Use hospital records to compare infection risk among patients who received two different catheter types.


10. Common Mistakes and How to Avoid Them

MistakeWhy It MattersBetter Practice
Using prevalent cases for etiologic claimsSurvival biasPrefer incident cases
Controls from wrong source populationSelection biasSelect controls from the population that produced cases
Measuring exposure after outcomeReverse causalityDefine exposure window before outcome
Ignoring loss to follow-upBiased cohort estimatesTrack and compare losses
Reporting odds ratio as risk ratio when outcome is commonExaggerates effectUse correct measure

11. Quick Reference Cheat Sheet

GoalBest Design
Study rare exposureCohort
Study rare diseaseCase-control
Estimate incidenceCohort
Study multiple outcomesCohort
Study long-latency disease efficientlyCase-control

Key formulas:

RR=[a/(a+b)]/[c/(c+d)]RR = [a/(a+b)]/[c/(c+d)]

RD=[a/(a+b)][c/(c+d)]RD = [a/(a+b)] - [c/(c+d)]

OR=ad/bcOR = ad/bc

Report source population, eligibility criteria, exposure window, outcome definition, confounding strategy, and effect estimate with confidence interval.