“Do we need three samples, or thirty?” There is no magic answer determined by experiment type alone. You must know which effect you cannot afford to miss, how variable values are, and which errors you will accept.
This unit asks one question.
How do we plan the amount of information needed for a conclusion before an experiment begins?
Four quantities form one relationship
Simple planning for a mean comparison needs:
- effect
Δ: the smallest practically meaningful difference to detect - variability
σ: expected SD among independent experimental units - significance level
α: Type I error accepted when H0 is true - power
1−β: probability of detecting a specified true Δ
Choose three of these and the analysis structure, and you can calculate n. A smaller effect or larger variability lowers Δ/σ, the signal-to-noise ratio, and requires more n. A stricter α or a higher power target also raises n.
For equal allocation to two independent groups with known σ, a normal approximation makes per-group n roughly proportional to (σ/Δ)². Halving Δ needs about four times as many observations. This is why asking to find even slightly smaller differences can be expensive.
The minimum detectable difference is not a difference seen in a past sample
Planning Δ is the smallest effect the study cannot reasonably miss. Putting a chance-large pilot difference directly into the calculation can under-plan n. Justify Δ with prior evidence, process tolerances, clinical importance, measurement units, and decision costs.
σ is uncertain as well. A small-pilot SD varies greatly and can be reduced by selected conditions. Use a conservative upper value, external data, a blinded internal pilot, or a sensitivity table, and present n at several plausible σ values.
Post hoc power calculated from the observed effect in the same data largely restates the p-value. To explain a non-significant result, report the pre-specified plan, CI width, detectable effects, and the information actually collected.
Power is a curve for each effect
A sample size does not have just one power. When the effect is zero, a properly calibrated α-level test rejects near α; power rises as the effect grows. A power curve shows which effects this n can find, and how well.
A one-sided test may have more power in its pre-specified direction, but it does not test the opposite direction. Switching to one-sided after seeing data to reduce n or lower a p-value is not acceptable.
In-Silico Lab: inspect sensitivity of planned n
The Lab uses a normal approximation for two independent groups, two-sided α=.05, equal n, and a known common SD.
- Reduce Δ from 2 to .5 and inspect the increase in per-group n.
- Raise SD from 2 to 4 and confirm the squared relationship.
- Compare the cost of 80%, 90%, and 95% power.
- Before fixing the calculation as the study n, list the missing design elements.
Plan independent n with effect, spread, and power
Two independent groups, two-tailed α=.05, transparently computes the normal approximation of equal distribution. This is a starting point and not a replacement for all designs.
If this is your first time: What should I press?
- 1. Read the question firstIn the Lab title, check the one thing you will compare this time.
- 2. Change just one conditionInitially, change only one of the inputs: n, effect, or spread.
- 3. new composite specimen pressureNew synthetic data is created. The same conditions may vary depending on the sample.
- 4. Pictures and calculation results CompareWrite in one sentence what moves and what stays the same before and after the change.
If it gets stuckresetGo back to see the default results and change just one condition. This Lab is not a correct answer tester but a pattern observation tool.
Synthetic observations of the same settings
calculation result
Adding the dropout rate and reflecting the design effects of clustering and repeated measurements are separate steps. Do not repeat p-values by recalculating power with post-observation effects.
educational synthetic modelbjs-comparison-sequence-v1. Actual research judgments must separately reflect experimental units, missingness, distribution, multiplicity, pre-planning, and domain criteria.
Real sample sizes are larger than a simple formula
These may need separate adjustment, simulation, or a specialist model:
- expected attrition, analysis exclusions, and measurement failure
- unequal allocation and cost differences
- cluster design effects, such as wells within donors or patients within sites
- repeated-measure correlation and paired SD
- multiple endpoints, interim analyses, and multiple comparisons
- different estimands: non-inferiority, equivalence, survival, or proportions
- integer block sizes and a minimum number of batches
For 10% attrition, increase from the n that must remain analysable—such as n/(1−0.10)—rather than simply n×1.10. In clustered data, both the number of individuals and clusters matter; more wells cannot replace donor-level information.
In JMP output, input assumptions are part of the result
Calculate the remainder based on known values among effect, SD, α, and power.
We see that the detection probability changes continuously when the effect changes.
Independent n, test direction, distribution ratio and analysis method must be consistent with the plan.
Keep more than a result screenshot. Record analysis method, two- versus one-sided choice, Δ, SD, α, power, allocation ratio, unit of calculated n, and software version. Usually round up, then adjust for the required block or pair structure.
Example planning statement
The primary endpoint was the difference in donor mean day-7 viability. We planned 55 donors per group for a minimum detectable difference of 5 percentage points, between-donor SD of 8, two-sided α=.05, 90% power, and 1:1 allocation. Allowing 10% to be non-analysable, we will recruit 62 donors per group. Technical well replicates will improve measurement precision but will not count toward independent n.
Takeaways
- Adequate n is a function of effect, variability, α, power, and analysis structure.
- Δ is a pre-specified effect that matters scientifically.
- Reflect uncertainty in σ through sensitivity analysis.
- Power is a curve over effects.
- Do not reinterpret results with observed-effect post hoc power.
- Add attrition, clustering, repetition, and multiplicity to practical planning.
The next units move from group differences to relationships between two continuous variables. Read the scatterplot before the correlation coefficient.
Official supplementary resources
The n in this Lab is an educational normal approximation. Review a real study plan against its design and applicable requirements.