Two reputable sample-size calculators can give different answers to what looks like the same question. The discrepancy is usually traceable: the tools are using a different statistical test, effect-size definition, tail, power target, approximation, allocation ratio, design correction, or output convention.
The result to trust is not automatically the largest, smallest, or most precise-looking number. Trust the result whose method and inputs match the prespecified primary analysis in your study protocol. If two tools truly use the same model and unrounded inputs, their answers should usually be identical or differ only slightly because of numerical method or rounding.
A quick example: 64, 86, or 128 participants?
Suppose a researcher plans to compare two independent group means with equal allocation, a standardized mean difference of 0.50, alpha of .05, and 80% power.
Using a basic large-sample normal approximation gives about 63 participants per group, or 126 in total. A calculator using the more exact noncentral t distribution may return about 64 per group, or 128 in total. That small difference is expected and comes from the method.
Now change just one input at a time:
| Calculator setup | Approximate result | Why it differs |
|---|---|---|
| Two-sided, 80% power, effect size 0.50 | 63 per group | Baseline normal approximation |
| One-sided, 80% power, effect size 0.50 | 50 per group | Less stringent critical value |
| Two-sided, 90% power, effect size 0.50 | 85 per group | Higher probability of detecting the target effect |
| Two-sided, 80% power, effect size 0.40 | 99 per group | Smaller target effect |
These are not four competing answers to one fully specified problem. They are answers to four different problems. The interactive calculator above lets you reproduce these contrasts using the same basic approximation.
The seven settings to compare before judging the answers
1. Statistical test and study design
A two-sample comparison of independent means is not interchangeable with a paired t test, repeated-measures model, cluster-randomized design, comparison of proportions, survival analysis, noninferiority test, or regression coefficient test. Each design uses information differently.
Check the calculator's exact module or test name first. If the planned primary analysis is a mixed model but the calculator assumes independent observations, agreement on alpha and power does not make the result valid.
2. Effect-size definition and value
An effect size can be entered as a raw mean difference, standardized mean difference, event-rate difference, odds ratio, hazard ratio, correlation, or another design-specific quantity. Even two standardized effects may use different denominators.
For two independent means, Cohen's d is commonly written as:
If one calculator uses while another derives from a larger standard deviation, the second calculator must request more participants. The effect should reflect a scientifically meaningful target and plausible variability, supported by prior evidence or sensitivity analysis rather than an arbitrary label such as “medium.”
3. Power
Power is the probability of rejecting the null hypothesis when the specified target effect is true. Increasing power from .80 to .90 increases the required sample because the study is being designed to miss that effect less often.
4. One-sided or two-sided alpha
A one-sided test places the rejection region in one direction and normally requires fewer observations. It is appropriate only when the directional hypothesis is justified in advance and an effect in the opposite direction would not support the same claim. Choosing one-sided testing merely to reduce the sample is not defensible.
Also check whether alpha is .05 for one primary test or has been adjusted for multiple primary comparisons.
5. Allocation ratio
Equal group sizes are usually most efficient for a simple two-group comparison. If recruitment will be 2:1 or 3:1, the total sample normally increases for the same power. Confirm whether a calculator asks for the ratio as treatment:control or control:treatment.
6. Design corrections and nuisance assumptions
Depending on the design, calculators may require or silently assume:
- Correlation between paired or repeated measurements.
- Intracluster correlation and average cluster size.
- Nonsphericity correction in repeated-measures ANOVA.
- Baseline event rate for a comparison of proportions.
- Accrual time, follow-up time, and censoring for survival analysis.
- Covariate adjustment or expected model fit.
- Continuity correction, finite-population correction, or multiplicity adjustment.
A small change in one of these assumptions can produce a large sample-size change.
7. What the displayed number means
One tool may report sample size per group, another the total, and another the number of complete analyzable participants before attrition. For cluster trials, the output may be people, clusters, or both. Read the result label before comparing the numbers.
When a small discrepancy is harmless
If the design and every input match, a difference of one or two participants per group can result from:
- A normal approximation versus the noncentral t distribution.
- Exact versus approximate methods for proportions.
- Rounding intermediate quantities versus rounding only the final result.
- Searching only integer group sizes rather than reporting a continuous solution.
In this situation, use the result from the method appropriate to the planned analysis and round upward. Document the software, version, statistical procedure, and all inputs so another researcher can reproduce it.
When a large discrepancy is a warning
A difference such as 64 versus 120 participants per group is unlikely to be explained by final rounding. Recheck the test, effect-size scale, power, tail, alpha, allocation, and design corrections. Also verify that one result is not total while the other is per group.
Do not average two incompatible answers. Averaging hides the disagreement without resolving which model represents the study.
Which result should you trust?
Use this decision sequence:
- Write the primary outcome, comparison, estimand, and planned statistical model before opening a calculator.
- Select a calculator module that represents that model and the dependence structure in the data.
- Enter an effect size on the correct scale and document its scientific and empirical justification.
- Match alpha, test direction, power, allocation, and all design-specific assumptions.
- Confirm whether the output is per group or total and whether it represents analyzable or recruited participants.
- Run sensitivity scenarios for uncertain inputs rather than relying on one optimistic value.
- Reproduce the preferred scenario with an independent calculator, statistical package, or published formula.
- If results still differ materially, consult the method documentation and resolve the formula difference before fixing the recruitment target.
The most trustworthy answer is therefore the most reproducible answer tied to the correct design, not simply the most conservative number. A larger sample does not repair a calculation based on the wrong outcome, test, or effect-size definition.
Adjust the analyzable sample for attrition
Apply attrition after estimating the analyzable sample. If the analysis requires 160 complete participants and anticipated attrition is 20%, divide by expected retention:
Simply adding 20% of 160 would give 192, which is expected to leave only about 154 complete participants. For clustered or longitudinal studies, participant dropout, whole-cluster loss, and unusable measurements may need separate treatment.
Report the calculation so it can be checked
A reproducible methods statement should identify the primary test or model, effect-size definition and source, alpha and tail, power, allocation ratio, design corrections, required analyzable sample, attrition allowance, final recruitment target, and calculation software or formula.
Use the DataStatPro sample size calculator to compare documented scenarios. For more background, read the sample size and power analysis tutorial. A sensitivity table with optimistic, central, and conservative assumptions is usually more informative than one exact-looking result.
Frequently asked questions
Why do two sample-size calculators give different results?
They may use different tests, effect-size definitions, approximations, tails, allocation ratios, design assumptions, rounding rules, or output conventions. Compare every setting before comparing the displayed numbers.
Which sample-size calculator should I trust?
Trust the calculation whose statistical method and assumptions match the prespecified primary analysis and can be independently reproduced. Do not choose solely by brand, result size, or number of available options.
Should I use the larger sample-size result to be safe?
Not automatically. A larger result based on the wrong model is not safer. First resolve why the results differ, then use the appropriate method and run sensitivity scenarios for uncertain assumptions.
Is sample size reported per group or in total?
It depends on the calculator. Check the output label, allocation settings, and whether the result refers to analyzable participants or the recruitment target.
How should I adjust sample size for attrition?
Divide the required analyzable sample by the expected retention proportion. For example, with 20% attrition, divide by 0.80 and round upward.
