A meta-analysis statistically combines comparable results from two or more studies. Done well, it can improve precision, show how effects vary across settings, and reveal where evidence remains uncertain. It cannot repair biased studies, incompatible outcomes, or a poorly designed review.

Meta-analysis at a glance

Step Main decision Output
1. Define the question Population, exposure or intervention, comparator, outcome, design Review question and eligibility criteria
2. Write the protocol Methods decided before results are known Registered or time-stamped protocol
3. Search and select Databases, search terms, screening rules Reproducible search and study flow
4. Extract and appraise Effect data, study features, risk of bias Verified analytical dataset
5. Harmonize effects Common effect measure and direction Effect estimate plus standard error for each study
6. Choose the model Fixed-effect or random-effects assumptions Prespecified synthesis model
7. Pool the estimates Weighting and variance estimation Summary effect and confidence interval
8. Assess variation Clinical, methodological, and statistical heterogeneity Q, I2I^2, τ2\tau^2, and often a prediction interval
9. Test robustness Alternative assumptions and influential studies Sensitivity and subgroup results
10. Report transparently Methods, results, limitations, certainty PRISMA-aligned report

1. Define a focused review question

Start with a question that determines which studies can be combined. For intervention reviews, a PICO structure is useful:

  • Population: Who is represented?
  • Intervention or exposure: What condition is being evaluated?
  • Comparator: What is the reference condition?
  • Outcome: Which definition, instrument, and follow-up time matter?
  • Study design: Which designs are eligible?

Translate the question into explicit inclusion and exclusion criteria before searching. Specify acceptable outcome definitions, time points, study designs, languages, publication types, and minimum data requirements.

A broad topic is not yet a meta-analysis question. “Does exercise help?” is too vague. “What is the effect of supervised aerobic exercise, compared with usual care, on systolic blood pressure at 12 to 24 weeks in adults with hypertension?” is much closer to an analyzable question.

2. Write and register a protocol

A protocol separates planned decisions from choices made after seeing results. It should define:

  • Eligibility criteria and information sources.
  • Primary and secondary outcomes.
  • Search and screening procedures.
  • Data items and risk-of-bias tools.
  • Effect measures and rules for choosing among multiple estimates.
  • The meta-analysis model and heterogeneity estimator.
  • Planned subgroup, meta-regression, and sensitivity analyses.
  • Methods for reporting-bias assessment and certainty of evidence.

Register the protocol in an appropriate registry when possible, or preserve a dated public protocol. Document later deviations and explain why they occurred.

3. Search systematically and select studies

Search multiple relevant bibliographic sources and consider trial registries, reference lists, grey literature, and contact with authors when appropriate. Save the exact query for every database, the date searched, and the number of records retrieved.

Use at least two independent reviewers for screening when resources permit. Resolve disagreements with a documented process. Keep exclusion reasons at the full-text stage so the study-flow diagram can be reconstructed.

The PRISMA 2020 statement provides the principal checklist and flow diagrams for transparent reporting. PRISMA is a reporting guideline; following it does not by itself guarantee that the review methods are valid.

4. Extract data and assess risk of bias

Create a prespecified extraction form. At minimum, collect:

  • Study identifier, setting, design, and recruitment dates.
  • Participant and intervention characteristics.
  • Outcome definition and measurement time.
  • Group sample sizes and events, means and standard deviations, or reported estimates and uncertainty.
  • Adjusted and unadjusted estimates, with the variables used for adjustment.
  • Funding, conflicts of interest, and information needed for risk-of-bias judgments.

Two people should independently extract critical outcome data or one should extract and another verify. Check that sample sizes, standard deviations, standard errors, confidence intervals, event coding, and treatment directions have not been confused.

Assess risk of bias by outcome, not only by study. Choose a design-appropriate tool, such as RoB 2 for randomized trials or ROBINS-I for non-randomized intervention studies. Do not convert detailed judgments into an arbitrary numerical “quality score” unless a validated method specifically requires it.

5. Choose a common effect size

Studies must address sufficiently similar questions and express results on a compatible scale. Common effect measures include:

Outcome type Common effect measures Interpretation of the null
Binary Risk ratio, odds ratio, risk difference 1 for ratios; 0 for differences
Continuous, same scale Mean difference 0
Continuous, different scales Standardized mean difference 0
Time to event Hazard ratio 1
Correlation Fisher's zz for pooling, often back-transformed to rr 0
Single proportion or prevalence Proportion on a justified transformed or generalized-model scale Depends on scale

Choose one effect measure because it answers the review question, not because it gives the smallest p value. Preserve a consistent direction: if higher scores are beneficial in one study and harmful in another, recode before pooling.

Ratio measures are usually analyzed on the logarithmic scale. For example, an odds ratio is pooled as log(OR)\log(OR) with its standard error, then exponentiated for presentation.

Avoid treating multiple effects from the same study as independent. Select one estimate using a prespecified rule or use a method that models dependency, such as multilevel meta-analysis, multivariate meta-analysis, or robust variance estimation.

6. Choose fixed-effect or random-effects meta-analysis

The model choice reflects the target of inference and assumptions about the true effects.

Fixed-effect model

A fixed-effect model assumes that all included studies estimate one common true effect and that observed differences arise from sampling error. It may be defensible when the studies are close replications and inference is restricted to that common effect.

Random-effects model

A random-effects model assumes that true effects vary across studies and estimates the mean of a distribution of effects. This is often more plausible when populations, implementation, settings, or follow-up periods differ.

Random effects is not an automatic solution to heterogeneity. It changes the estimand and weighting, and its mean may be uninformative when effects differ greatly or point in opposing directions. Do not select the model only because a heterogeneity test is or is not statistically significant. The current Cochrane Handbook chapter on meta-analysis likewise emphasizes clinical judgment, careful heterogeneity assessment, and sensitivity analysis.

For random-effects analysis, report the estimator for the between-study variance τ2\tau^2, such as restricted maximum likelihood, and the method used for the summary-effect confidence interval. With few studies, conventional Wald intervals can be too optimistic; a method such as Hartung-Knapp may be considered when appropriate.

7. Calculate the pooled effect

Most standard meta-analyses use a weighted average:

θ^=i=1kwiθ^ii=1kwi\hat\theta=\frac{\sum_{i=1}^{k} w_i\hat\theta_i}{\sum_{i=1}^{k} w_i}

where θ^i\hat\theta_i is the effect estimate from study ii and wiw_i is its weight.

Under a common inverse-variance fixed-effect model:

wi=1SEi2w_i=\frac{1}{SE_i^2}

Under a conventional random-effects model:

wi=1SEi2+τ2w_i=\frac{1}{SE_i^2+\tau^2}

The random-effects weight includes estimated between-study variance, so smaller studies often receive relatively more weight than under a fixed-effect model. The precise calculation depends on the outcome, estimator, small-sample method, and handling of sparse data.

Use the DataStatPro meta-analysis software to enter study estimates, calculate fixed-effect and random-effects models, inspect forest plots, and evaluate heterogeneity. Preserve the dataset and all model settings so the result can be reproduced.

8. Assess heterogeneity

Heterogeneity means that study results vary beyond what sampling error alone would be expected to produce. Examine three forms:

  • Clinical heterogeneity: differences in participants, interventions, comparators, outcomes, or follow-up.
  • Methodological heterogeneity: differences in design, measurement, analysis, or risk of bias.
  • Statistical heterogeneity: observed effect estimates differ more than expected from within-study sampling variation.

Common statistics answer different questions:

  • Cochran's Q: tests compatibility with a common-effect model but has low power with few studies and high power with many.
  • I2I^2: estimates the proportion of observed variation attributable to heterogeneity rather than sampling error. It is not the percentage of studies that disagree.
  • τ2\tau^2: estimates the between-study variance on the analysis scale.
  • Prediction interval: estimates a range for the true effect in a comparable future setting under the random-effects model.

Do not interpret I2I^2 from rigid cutoffs alone. Consider its uncertainty, the magnitude and direction of effects, τ2\tau^2, clinical differences, and the number of studies. A prediction interval can be especially informative when the summary effect is favorable but plausible effects in new settings include no benefit or harm.

9. Read the forest plot correctly

A forest plot usually shows each study's effect estimate as a marker, its confidence interval as a horizontal line, its weight through marker size, and the pooled estimate as a diamond.

Interpret it in this order:

  1. Confirm the effect measure, scale, and direction of benefit.
  2. Locate the no-effect value: 1 for ratio measures and 0 for difference measures.
  3. Compare study estimates and confidence intervals, not just p values.
  4. Look for important differences in direction or magnitude.
  5. Read the pooled effect and its confidence interval.
  6. Interpret heterogeneity and, when available, the prediction interval.
  7. Check subgroup labels, model type, and study weights.

A pooled diamond that excludes the null does not prove that every study, population, or future setting has a non-null effect.

10. Assess reporting bias and small-study effects

A funnel plot displays effect size against a measure of study precision. Asymmetry can arise from publication bias, selective reporting, heterogeneity, data errors, or chance; it is not a direct test of publication bias.

Statistical tests for funnel-plot asymmetry are generally unreliable with few studies. When enough comparable studies are available, select a test appropriate to the effect measure and data structure. Compare published reports with registries or protocols, search for unavailable results, and assess selective reporting alongside any funnel plot.

Methods such as trim-and-fill rely on strong assumptions and should be presented as sensitivity analyses, not as corrections that reveal the “true” effect.

11. Explore heterogeneity without overfitting

Use subgroup analysis or meta-regression only when a scientific reason suggests that an effect modifier may matter. Specify the expected direction before looking at the results when possible.

Compare subgroups with a formal interaction test. Evidence that one subgroup is statistically significant and another is not does not establish a difference between them. Meta-regression is an observational, study-level analysis and is vulnerable to confounding and ecological bias. With few studies or many candidate moderators, apparent explanations are often unstable.

Clearly distinguish prespecified analyses from exploratory analyses generated after seeing the data.

12. Run sensitivity and influence analyses

Sensitivity analysis asks whether reasonable analytical choices change the conclusion. Useful checks may include:

  • Fixed-effect versus random-effects results when both target relevant questions.
  • Alternative τ2\tau^2 estimators or confidence-interval methods.
  • Exclusion of studies at high risk of bias.
  • Alternative outcome definitions, follow-up times, or correlation assumptions.
  • Methods designed for zero events or sparse data.
  • Leave-one-out analysis and influence diagnostics.
  • Models that account for dependent effect sizes.

Do not delete an influential study solely because it weakens the pooled result. Investigate data accuracy and clinical or methodological reasons, then report analyses with and without the study when exclusion is scientifically defensible.

Worked interpretation example

Suppose 12 randomized trials compare an intervention with usual care. A random-effects model gives a risk ratio of 0.78, with a 95% confidence interval from 0.66 to 0.93, I2=48%I^2=48\%, and a 95% prediction interval from 0.54 to 1.14.

A defensible interpretation is:

Across the included studies, the intervention was associated with a lower average risk than usual care (RR = 0.78, 95% CI [0.66, 0.93]). Results varied moderately across studies (I2=48%I^2=48\%). The prediction interval (0.54 to 1.14) indicates that the true effect in a comparable new setting could range from substantial benefit to little or no benefit.

This wording distinguishes the estimated mean effect from the expected variation across settings. The review should still address risk of bias, certainty of evidence, outcome importance, absolute risks, and applicability.

When drafting the results, use the companion guide to report p values, confidence intervals, and effect sizes without reducing the synthesis to a significance label.

What should a meta-analysis report?

Use the appropriate reporting guidance and include enough detail for another analyst to reproduce the synthesis:

  1. Protocol and registration details.
  2. Complete eligibility criteria and search strategies.
  3. Screening process and PRISMA flow diagram.
  4. Data-extraction and risk-of-bias procedures.
  5. Effect measure, model, weighting method, and software version.
  6. τ2\tau^2 estimator and confidence-interval method for random effects.
  7. Rules for multiple outcomes, time points, and dependent estimates.
  8. Forest plot, study estimates, summary effect, and uncertainty.
  9. Heterogeneity statistics and prediction interval when appropriate.
  10. Prespecified subgroup, meta-regression, and sensitivity analyses.
  11. Reporting-bias assessment and certainty of evidence.
  12. Limitations, protocol deviations, funding, and conflicts of interest.

The PRISMA 2020 expanded checklist specifically asks authors to identify the model, synthesis method, heterogeneity methods, software, and relevant random-effects details.

If the included studies use several outcome types, the statistical test decision guide can help clarify the underlying effect estimates before they are converted to a common meta-analysis scale.

Common meta-analysis mistakes

  • Pooling studies because they report the same outcome name even though the populations, interventions, or time points are incompatible.
  • Choosing an effect measure or model after inspecting which result is favorable.
  • Mixing adjusted and unadjusted estimates without a protocol-based rule.
  • Counting multiple correlated outcomes from one study as independent studies.
  • Entering a standard deviation as a standard error, or reversing event coding.
  • Treating a non-significant Q test as proof of homogeneity.
  • Using I2I^2 alone to decide between fixed effect and random effects.
  • Calling funnel-plot asymmetry proof of publication bias.
  • Running many unplanned subgroup analyses and reporting only favorable findings.
  • Interpreting a precise pooled estimate without considering risk of bias or evidence certainty.

Frequently asked questions

What is a meta-analysis in simple terms?

A meta-analysis is a statistical method that combines comparable effect estimates from two or more independent studies. Larger or more precise studies usually receive more weight, but the exact weighting depends on the model.

What is the difference between a systematic review and a meta-analysis?

A systematic review uses explicit methods to find, select, appraise, and synthesize evidence. A meta-analysis is the statistical combination step that may be included in a systematic review. A valid review may decide not to pool studies when they are too different.

How many studies are needed for a meta-analysis?

Mathematically, two studies can be pooled, but that does not make the result reliable. With very few studies, heterogeneity, publication-bias tests, prediction intervals, and random-effects uncertainty are difficult to estimate. The studies must also be sufficiently comparable.

Should I use fixed effect or random effects?

Use the model that matches the intended inference and assumptions. Fixed effect targets one common effect; random effects targets a mean across varying true effects. Do not choose only from the p value of a heterogeneity test.

What does a high I-squared mean?

A high I-squared suggests that much of the observed variation is beyond within-study sampling error. Its importance depends on the size and direction of the effects, uncertainty in I-squared, the between-study variance, study characteristics, and the number of studies.

Can I do a meta-analysis without individual participant data?

Yes. Most conventional meta-analyses use aggregate effect estimates and standard errors extracted from publications or supplied by investigators. Individual participant data can support more consistent analyses but require access, harmonization, and methods that preserve clustering by study.

Is a forest plot the same as a meta-analysis?

No. A forest plot is a visualization of study estimates and, often, a pooled result. A forest plot can be shown without pooling, and a valid meta-analysis requires design, extraction, appraisal, modeling, and interpretation decisions beyond drawing the plot.

Does meta-analysis prove causation?

No. Causal interpretation depends on the designs and biases of the included studies, the review methods, consistency, directness, precision, and other evidence. Pooling biased observational estimates does not turn them into randomized evidence.

Apply this guide

Run and document a meta-analysis in DataStatPro

The DataStatPro meta-analysis workflow combines study-level estimates with transparent model settings, forest plots, heterogeneity statistics, prediction intervals, publication-bias diagnostics, and exportable results. The protocol still determines which studies and estimates belong in the synthesis.

  • Import verified effect estimates and their standard errors.
  • Choose and document the effect measure, model, and variance method.
  • Interpret the pooled effect with heterogeneity, robustness, and risk of bias.
Run a meta-analysis
Editorial policy: Read our authorship, review, and corrections policy. Examples are educational and should be adapted to the study design and destination requirements.