Comparing mean yield for media A, B, C, and D with every pairwise t-test creates six tests. At ฮฑ=.05 for each, the opportunity for at least one false positive somewhere grows.
This unit asks one question.
How can we ask about mean differences among three or more groups under one shared error criterion?
ANOVA calculates a ratio of variation rather than comparing means directly
The one-way ANOVA null is H0: ฮผโ=ฮผโ=โฆ=ฮผk. Its alternative is not โevery mean differs,โ but at least one mean differs. This overall question is the omnibus test.
Partition the total variation of observations around the grand mean into:
- between-group variation: how far group means are from the grand mean
- within-group variation: how observations vary around their own group mean
SS_total = SS_between + SS_within
Sum of squares (SS) naturally grows with the number of observations. Divide each SS by its degrees of freedom (df) to form mean squares (MS), then calculate F=MS_between/MS_within. Under H0, both MS target the same error variance, so F tends to lie near 1. F grows when separation among group means exceeds within-group noise.
Each ANOVA-table column connects in one statement
With k groups and N total independent observations:
- between df =
kโ1 - within df =
Nโk MS_between=SS_between/(kโ1)MS_within=SS_within/(Nโk)F=MS_between/MS_within
The ANOVA p-value is the probability, under H0 and the model, of an F at least this large. A significant F is evidence that a difference exists somewhere; it does not identify which group, how large the difference is, or whether all groups differ.
After an omnibus result, report group means, SDs, and n; effect sizes; and planned comparisons or multiplicity-adjusted post hoc comparisons with CIs. Do not automatically label the largest mean as โbest.โ
Repeated t-tests versus post hoc comparisons
If a specific comparison was the scientific question before analysis, state it as a planned comparison. If all pairs are explored after viewing the data, a family-wise-error-controlled procedure such as Tukey HSD is an appropriate candidate. Use methods matched to the purpose, such as Dunnett for comparing several treatments with one control.
Adjustment can widen each CI. This is not a penalty: it is the cost of making one error promise across many opportunities. Selecting arbitrary pairwise p-values when ANOVA is non-significant needs justification from the pre-specified plan and method.
Assumptions are not a collection of per-group normality p-values
Classical one-way ANOVA assumes independent errors, roughly normal errors within groups, and equal variances. Balanced designs have some robustness, but problems grow with small n, strong skew, outliers, or an inverse relationship between variance and n.
Welch ANOVA is an alternative less dependent on equal variances. Rank-based KruskalโWallis is not an assumption-free version of mean ANOVA; its distributional and location interpretations differ. Repeated-measures or clustered data require a structure-matched analysis such as a mixed model, not direct insertion into a one-way ANOVA table.
In-Silico Lab: verify SS by hand
The synthetic means of three groups are 10, 10+d, and 10+2d.
- At d=0, see that F is not fixed at exactly 1.
- Increase d and inspect the tendency for SS_between and F to grow.
- Change the seed and explain why SS_within and F change.
- Use the calculation engine to verify that displayed
SS_between + SS_withinequals total SS.
์ธ ํ๊ท ์ ๋ถ๋ฆฌ์ ์ง๋จ๋ด ํ๋ค๋ฆผ์ ๋๋์ธ์
์ธ ์ง๋จ ํ๊ท ํจํด์ ๋ฐ๊พธ์ด SS_between, SS_within๊ณผ F ratio๊ฐ ์ด๋ป๊ฒ ์ฐ๊ฒฐ๋๋์ง ๋ด ๋๋ค.
์ฒ์์ด๋ผ๋ฉด: ๋ฌด์์ ๋๋ฌ์ผ ํ๋์?
- 1. ์ง๋ฌธ์ ๋จผ์ ์ฝ๊ธฐLab ์ ๋ชฉ์์ ์ด๋ฒ์ ๋น๊ตํ ํ ๊ฐ์ง๋ฅผ ํ์ธํฉ๋๋ค.
- 2. ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ๊ธฐ์ฒ์์๋ n, ํจ๊ณผ, ์ฐํฌ ๊ฐ์ ์ ๋ ฅ ์ค ํ๋๋ง ๋ฐ๊พธ์ญ์์ค.
- 3. ์ ํฉ์ฑ ํ๋ณธ ๋๋ฅด๊ธฐ์ ํฉ์ฑ ๋ฐ์ดํฐ๊ฐ ๋ง๋ค์ด์ง๋๋ค. ๊ฐ์ ์กฐ๊ฑด๋ ํ๋ณธ์ ๋ฐ๋ผ ๋ฌ๋ผ์ง ์ ์์ต๋๋ค.
- 4. ๊ทธ๋ฆผ๊ณผ ๊ณ์ฐ ๊ฒฐ๊ณผ ๋น๊ตํ๊ธฐ๋ฐ๊พธ๊ธฐ ์ ํ ๋ฌด์์ด ์์ง์ด๊ณ ๋ฌด์์ด ๊ทธ๋๋ก์ธ์ง ํ ๋ฌธ์ฅ์ผ๋ก ์ ์ด๋ณด์ญ์์ค.
๋งํ๋ฉด ์ด๊ธฐํ๋ก ๋์๊ฐ ๊ธฐ๋ณธ ๊ฒฐ๊ณผ๋ฅผ ๋ณธ ๋ค ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ์ญ์์ค. ์ด Lab์ ์ ๋ต ํ์ ๊ธฐ๊ฐ ์๋๋ผ ํจํด ๊ด์ฐฐ ๋๊ตฌ์ ๋๋ค.
๊ฐ์ ์ค์ ์ ํฉ์ฑ ๊ด์ธก
๊ณ์ฐ ๊ฒฐ๊ณผ
F๊ฐ ํฌ๋ค๋ ๊ฒ์ ์ง๋จ๊ฐ ์ ํธ๊ฐ ์ง๋จ๋ด ์ก์๋ณด๋ค ํฌ๋ค๋ ๋ป์ ๋๋ค. ์ด๋ ์์ด ๋ค๋ฅธ์ง๋ ๋ณ๋ ๋น๊ต๊ฐ ํ์ํฉ๋๋ค.
๊ต์ก์ฉ synthetic model ยท bjs-comparison-sequence-v1. ์ค์ ์ฐ๊ตฌ ํ๋จ์๋ ์คํ๋จ์, ๊ฒฐ์ธก, ๋ถํฌ, ๋ค์ค์ฑ, ์ฌ์ ๊ณํ๊ณผ ๋๋ฉ์ธ ๊ธฐ์ค์ ๋ณ๋๋ก ๋ฐ์ํด์ผ ํฉ๋๋ค.
Start JMP ANOVA interpretation from the decomposition table
SSยทdfยทMSยทF๊ฐ ๋ณ๋ ๋ถํด๋ฅผ ํ๋ก ๋ณด์ฌ์ค๋๋ค.
๋ถ์ฐ๊ณผ n์ด ํฌ๊ฒ ๋ค๋ฅธ ์ํฉ์ ๋์์ ํ์ํฉ๋๋ค.
omnibus ๋ค ์ฌ์ ๋น๊ต๋ ๋ณด์ ๋ ์ฌํ๋น๊ต๋ฅผ ํด์ํฉ๋๋ค.
Read Analysis of Variance from left to right: Source, DF, Sum of Squares, Mean Square, F Ratio, and Prob > F. Do not declare a model good from Summary of Fit Rยฒ alone. Connect it to raw group data and residuals, Welch results, and simultaneous CIs for post hoc comparisons.
Example result statement
Mean yield for four independent batch conditions was compared by one-way ANOVA (n=8 each). The condition effect was F(3,28)=5.42, p=.004. In prespecified Dunnett comparisons versus control, BโA was 3.1 percentage points (simultaneous 95% CI 0.8โ5.4), while the C and D CIs contained zero. Residual plots showed no marked curvature or increasing-variance pattern.
Takeaways
- The ANOVA alternative means at least one mean differs.
- Total SS partitions exactly into between- and within-group SS.
- F is the ratio of two mean squares.
- An omnibus p-value does not say which groups differ.
- Planned and adjusted post hoc comparisons differ in timing and scope of the question.
- Do not reduce repeated or clustered structures to independent one-way ANOVA.
The next unit addresses the error of calling a non-significant result โthe same,โ then introduces equivalence margins and TOST.
Official supplementary resources
The values, figures, and Lab in this article are educational synthetic material.