Protocol Writing and Statistical Analysis Plans: Zero to Hero Tutorial
This tutorial explains how to write research protocols and statistical analysis plans that make studies transparent, reproducible, and analysis-ready. It is useful for experiments, clinical trials, observational studies, surveys, and program evaluations.
Table of Contents
- Prerequisites and Background Concepts
- Why Protocols and SAPs Matter
- Protocol Structure
- Statistical Analysis Plan Structure
- Endpoints and Estimands
- Sample Size and Power
- Missing Data and Protocol Deviations
- Reporting and Reproducibility
- Using DataStatPro
- Worked Examples
- Common Mistakes and How to Avoid Them
- Quick Reference Cheat Sheet
1. Prerequisites and Background Concepts
You should understand:
- Protocol: The complete plan for conducting the study.
- Statistical analysis plan: The pre-specified plan for analyzing data.
- Endpoint: Outcome used to evaluate the research question.
- Analysis population: Set of participants included in a specific analysis.
- Estimand: The treatment or exposure effect being targeted.
- Protocol deviation: Departure from the approved study plan.
2. Why Protocols and SAPs Matter
Protocols and SAPs reduce:
- Selective reporting.
- Outcome switching.
- Analytical flexibility.
- Ambiguous decision rules.
- Reproducibility problems.
They also help teams align before data collection begins.
3. Protocol Structure
A strong protocol includes:
- Title and version history.
- Background and rationale.
- Objectives and hypotheses.
- Study design.
- Setting and population.
- Eligibility criteria.
- Intervention or exposure definitions.
- Outcomes and measurement schedule.
- Sample size justification.
- Data collection and management.
- Ethics and safety.
- Dissemination plan.
The protocol should explain what will happen. The SAP should explain exactly how results will be analyzed.
4. Statistical Analysis Plan Structure
A SAP should include:
- Analysis objectives.
- Primary and secondary endpoints.
- Analysis populations.
- Derived variables.
- Descriptive summaries.
- Primary model.
- Covariates.
- Subgroup analyses.
- Multiplicity strategy.
- Missing-data methods.
- Sensitivity analyses.
- Table and figure shells.
Pre-specification does not prevent judgment. It prevents hidden judgment.
5. Endpoints and Estimands
An endpoint defines what is measured. An estimand defines the effect being estimated.
Estimand elements:
- Population.
- Treatment or exposure conditions.
- Endpoint.
- Intercurrent events.
- Summary measure.
Example:
Difference in mean 12-week systolic blood pressure between assigned treatment arms, regardless of adherence, among randomized participants.
6. Sample Size and Power
Document:
- Primary endpoint.
- Effect size or precision target.
- Significance level.
- Power.
- Variance or event-rate assumptions.
- Allocation ratio.
- Attrition adjustment.
For attrition:
where is expected dropout proportion.
7. Missing Data and Protocol Deviations
Plan how to handle:
- Missing baseline variables.
- Missing outcomes.
- Dropouts.
- Nonadherence.
- Ineligible participants.
- Duplicate records.
- Outliers.
Common analysis populations:
| Population | Definition |
|---|---|
| Intention-to-treat | Analyzed according to assignment |
| Per-protocol | Followed protocol sufficiently |
| Safety | Received at least one exposure or intervention |
| Complete case | Has required variables for analysis |
8. Reporting and Reproducibility
Prepare:
- Table shells.
- Figure shells.
- Variable dictionary.
- Analysis dataset specifications.
- Version-controlled analysis scripts.
- Audit trail for protocol amendments.
Report deviations from the SAP transparently.
9. Using DataStatPro
Use DataStatPro to:
- Estimate sample size and power.
- Create planned descriptive tables.
- Run pre-specified tests and models.
- Export effect estimates and confidence intervals.
- Generate publication-ready figures.
The SAP should name the planned DataStatPro analysis module when appropriate.
10. Worked Examples
Example 1: Two-Arm Trial SAP
Primary endpoint: 12-week blood pressure. Primary model: ANCOVA adjusted for baseline blood pressure and site. Analysis population: intention-to-treat.
Example 2: Survey Protocol
Primary estimate: student satisfaction proportion. Sampling: stratified by year. Analysis: weighted proportion with 95% confidence interval.
Example 3: Observational Cohort SAP
Primary exposure: treatment received within 48 hours. Primary outcome: 30-day readmission. Analysis: adjusted logistic regression with pre-specified confounders.
11. Common Mistakes and How to Avoid Them
| Mistake | Why It Matters | Better Practice |
|---|---|---|
| SAP written after seeing outcomes | Bias risk | Finalize before analysis |
| Vague endpoints | Outcome switching | Define variable, time point, and scoring |
| No missing-data plan | Flexible conclusions | Pre-specify primary and sensitivity methods |
| Too many unplanned subgroups | False positives | Limit and justify subgroup analyses |
| No table shells | Reporting drift | Draft shells before analysis |
12. Quick Reference Cheat Sheet
| Document | Purpose |
|---|---|
| Protocol | Conduct the study |
| SAP | Analyze the study |
| Data dictionary | Define variables |
| Table shells | Predefine reporting |
| Amendment log | Track changes |
SAP essentials:
- Primary endpoint.
- Primary model.
- Analysis population.
- Covariates.
- Missing-data strategy.
- Multiplicity strategy.
- Sensitivity analyses.
Key formula:
Report protocol version, SAP version, amendment dates, and deviations.