Why Sample Size Calculators Give Different Answers and Which Result to Trust

Interactive example

Adjust a sample size for attrition

Enter the analyzable sample and expected loss. The result is rounded up to the next whole participant.

Continue in DataStatPro

Two reputable sample-size calculators can return different numbers for what appears to be the same study. Usually, neither result is random. The calculators are solving different versions of the problem because one or more inputs, formulas, defaults, or rounding rules differ.

The right response is not to choose the larger or smaller answer automatically. It is to identify which calculation matches the planned design and estimand.

The assumptions that change sample size most

Effect size

Smaller effects require larger samples. A standardized mean difference of 0.30 and 0.50 describe meaningfully different planning targets. Your effect-size assumption should come from clinically or scientifically meaningful differences, prior evidence, pilot uncertainty, or a justified range, not merely a conventional label such as “medium.”

Power

Increasing power from .80 to .90 reduces the chance of missing the planned effect, but it increases the required sample. Verify whether the displayed number is total sample size or sample size per group.

One-sided versus two-sided testing

A one-sided test can require fewer observations, but it is defensible only when effects in the opposite direction would not lead to the same claim and the direction was specified in advance.

Allocation ratio

Equal group sizes are often most efficient. If recruitment produces a 2:1 or 3:1 allocation, the total sample normally increases for the same power.

Study design

Independent groups, paired measurements, repeated measures, clusters, survival outcomes, proportions, and noninferiority designs use different information structures. Selecting “two means” is not enough if the observations are paired or clustered.

Defaults that quietly create disagreement

Calculators may use different values for:

  • Significance level and multiplicity adjustment.
  • Expected standard deviation or baseline event rate.
  • Correlation between repeated measurements.
  • Nonsphericity correction in repeated-measures ANOVA.
  • Intracluster correlation and average cluster size.
  • Continuity correction for proportions.
  • Exact, approximate, or simulation-based formulas.
  • Numerator degrees of freedom in omnibus tests.
  • Rounding before or after unequal allocation.

Record every input shown on the results page. A sample-size statement that reports only power and alpha is often not reproducible.

Worked comparison

Imagine a two-group study planned with a two-sided alpha of .05 and power of .80. Calculator A assumes a standardized effect of 0.50 and equal allocation. Calculator B uses 0.45 after incorporating a more conservative standard deviation. Calculator B will require more participants even though both are correctly configured.

If a third calculator reports sample size per group while the others report total sample size, its number may appear to be half as large. The label, not the mathematics, explains the apparent disagreement.

Which answer should you trust?

Trust the result whose statistical model and inputs correspond to your written study plan. Use this audit sequence:

  1. Match the design and primary outcome.
  2. Confirm whether the hypothesis is superiority, equivalence, or noninferiority.
  3. Match one-sided or two-sided alpha.
  4. Verify the effect-size definition and its source.
  5. Confirm power, allocation, correlation, clustering, and number of groups.
  6. Separate analyzable sample size from recruitment target.
  7. Reproduce the calculation with an independent tool or formula.
  8. Preserve a dated export of the inputs and result.

Attrition is applied after the analyzable sample

If the analysis requires 160 complete participants and anticipated attrition is 20%, do not simply add 20% of 160. Divide by the expected retention:

Nrecruit=Nanalysis1attrition=1600.80=200.N_{recruit} = \frac{N_{analysis}}{1 - attrition} = \frac{160}{0.80} = 200.

For clustered or longitudinal studies, loss may operate at more than one level. Participant dropout, whole-cluster loss, and unusable measurements should be considered separately where relevant.

Plan with a range, not false precision

When the effect size or event rate is uncertain, calculate several plausible scenarios. A sensitivity table showing optimistic, central, and conservative assumptions is more informative than a single exact-looking number.

Use the DataStatPro sample size calculator to document the selected design and inputs. For deeper guidance, read the sample size and power analysis tutorial. The defensible sample size is the one connected to a transparent design decision, not merely the first number a calculator displays.

Frequently asked questions

Why do two sample-size calculators give different results?

They may use different tests, effect-size definitions, approximations, tails, allocation ratios, correlation assumptions, continuity corrections, rounding rules, or output conventions.

Which sample-size calculator should I trust?

Trust the calculation that matches the written study design and whose inputs and method can be documented and independently reproduced.

Is sample size reported per group or in total?

It depends on the calculator. Check the result label and allocation settings before comparing numbers across tools.

How should I adjust sample size for attrition?

Divide the required analyzable sample by the expected retention proportion. For example, with 20% attrition, divide by 0.80 and round up.

Editorial review: DataStatPro Statistical Review. Examples are educational and should be adapted to the study design and destination requirements.