A difference test between the means of A and B gives p=.24. May we write, โThere is no significant difference, so the two conditions are the sameโ? No. With a small sample and large variation, even a large difference can remain undetected.
This unit has one question.
How do we construct positive evidence that two conditions are sufficiently similar?
Difference tests and equivalence tests have opposite null hypotheses
A difference test begins with H0: difference=0 and seeks evidence that a difference exists. Failure to reject means the data are compatible with 0; it does not prove that the difference lies within a small range.
An equivalence test first specifies a practically negligible boundary, โฮ and +ฮ. If the mean difference is ฮด=ฮผBโฮผA:
H0: ฮดโคโฮ or ฮดโฅ+ฮโ the difference lies outside the acceptable range.H1: โฮ<ฮด<+ฮโ the difference lies inside the acceptable range.
TOST tests the following two null hypotheses separately at one-sided ฮฑ.
- Reject H0 that B is at least
ฮlower than A. - Reject H0 that B is at least
ฮhigher than A.
Both must be rejected before evidence for equivalence exists. If either fails, the result may establish non-equivalence, or it may simply be inconclusive.
A TOST at ฮฑ=.05 corresponds to a 90% CI
In general, if each one-sided test uses ฮฑ, the same decision can be made by checking whether the (1โ2ฮฑ)ร100% two-sided CI for the difference lies entirely inside the equivalence margins. At ฮฑ=.05, this is a 90% CI.
- the full CI lies within
โฮand+ฮ: evidence for equivalence - the CI crosses a margin but includes 0: inconclusive for both difference and equivalence
- the CI is far from 0 and outside a margin: a meaningful difference is possible
Do not use a conventional 95% CI interchangeably with TOST p-values. A 95% CI may also be presented for estimation, but state clearly which CI corresponds to which decision.
Widening a margin until the observed CI fits changes the question after the fact. ฮ must be justified before seeing the data using scientific, clinical, or quality impact, measurement reliability, and prior evidence. Regulatory settings require the relevant guidance and expert review separately.
See the CI and margins before two small p-values
TOST output contains a lower-test and an upper-test p-value. The larger p-value must be below ฮฑ for both conditions to pass. But placing the CI and margins on the same axis shows more directly which boundary failed and in which direction the effect lies.
Equivalence does not automatically establish sameness of an entire process, specification conformity of an individual product, equality of variances, or interchangeability. State equivalence narrowly: for which endpoint, population, time window, and effect scale?
In-Silico Lab: the same true difference can lead to different conclusions at different n
- Compare n=8 and n=100 with a true difference of 0 and ฮ=1.
- Observe the region where a small n supports neither a difference test nor an equivalence test.
- Raise the true difference to 1.5 and see whether the CI center moves outside a margin.
- Explain why widening ฮ without changing the data adds no scientific evidence.
CI ์ ์ฒด๊ฐ ๋๋ฑ์ฑ ํ๊ณ ์์ ๋ค์ด์ค๋์ง ๋ณด์ธ์
ฮฑ=.05 TOST์ ๋์ํ๋ 90% CI๋ฅผ ์ฌ์ฉํด ์ฐจ์ดยท๋๋ฑยท๋ถํ์ ์ง๋ฌธ์ด ์ ๋ค๋ฅธ์ง ํ์ธํฉ๋๋ค.
์ฒ์์ด๋ผ๋ฉด: ๋ฌด์์ ๋๋ฌ์ผ ํ๋์?
- 1. ์ง๋ฌธ์ ๋จผ์ ์ฝ๊ธฐLab ์ ๋ชฉ์์ ์ด๋ฒ์ ๋น๊ตํ ํ ๊ฐ์ง๋ฅผ ํ์ธํฉ๋๋ค.
- 2. ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ๊ธฐ์ฒ์์๋ n, ํจ๊ณผ, ์ฐํฌ ๊ฐ์ ์ ๋ ฅ ์ค ํ๋๋ง ๋ฐ๊พธ์ญ์์ค.
- 3. ์ ํฉ์ฑ ํ๋ณธ ๋๋ฅด๊ธฐ์ ํฉ์ฑ ๋ฐ์ดํฐ๊ฐ ๋ง๋ค์ด์ง๋๋ค. ๊ฐ์ ์กฐ๊ฑด๋ ํ๋ณธ์ ๋ฐ๋ผ ๋ฌ๋ผ์ง ์ ์์ต๋๋ค.
- 4. ๊ทธ๋ฆผ๊ณผ ๊ณ์ฐ ๊ฒฐ๊ณผ ๋น๊ตํ๊ธฐ๋ฐ๊พธ๊ธฐ ์ ํ ๋ฌด์์ด ์์ง์ด๊ณ ๋ฌด์์ด ๊ทธ๋๋ก์ธ์ง ํ ๋ฌธ์ฅ์ผ๋ก ์ ์ด๋ณด์ญ์์ค.
๋งํ๋ฉด ์ด๊ธฐํ๋ก ๋์๊ฐ ๊ธฐ๋ณธ ๊ฒฐ๊ณผ๋ฅผ ๋ณธ ๋ค ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ์ญ์์ค. ์ด Lab์ ์ ๋ต ํ์ ๊ธฐ๊ฐ ์๋๋ผ ํจํด ๊ด์ฐฐ ๋๊ตฌ์ ๋๋ค.
๊ฐ์ ์ค์ ์ ํฉ์ฑ ๊ด์ธก
๊ณ์ฐ ๊ฒฐ๊ณผ
CI๊ฐ ํ๊ณ์ ๊ฒน์น๋ ๊ฒ๋ง์ผ๋ก๋ ๋ถ์กฑํฉ๋๋ค. ์ ์ฒด ๊ตฌ๊ฐ์ด ์์ ์์ด์ผ ๋ ๋จ์ธก ๊ท๋ฌด๊ฐ์ค์ ๋ชจ๋ ๊ธฐ๊ฐํฉ๋๋ค.
๊ต์ก์ฉ synthetic model ยท bjs-comparison-sequence-v1. ์ค์ ์ฐ๊ตฌ ํ๋จ์๋ ์คํ๋จ์, ๊ฒฐ์ธก, ๋ถํฌ, ๋ค์ค์ฑ, ์ฌ์ ๊ณํ๊ณผ ๋๋ฉ์ธ ๊ธฐ์ค์ ๋ณ๋๋ก ๋ฐ์ํด์ผ ํฉ๋๋ค.
The Lab is an educational one-sample difference model with known ฯ. Actual two-sample, paired, proportion, or log-transformed endpoints need design-appropriate SEs and CIs.
JMP displays margins and CIs but does not justify the margin
์ฐจ์ด CI์ ์ฌ์ ๋๋ฑ์ฑ ํ๊ณ๋ฅผ ๊ฐ์ ์ถ์์ ๋ด ๋๋ค.
๋ ๋จ์ธก ๊ฒ์ ์ด ๋ชจ๋ ๊ธฐ์ค์ ๋์๋์ง ํ์ธํฉ๋๋ค.
ํต๊ณ ํ๋ก๊ทธ๋จ์ด ํ๊ณ์ ๊ณผํ์ ํ๋น์ฑ์ ์ ํด์ฃผ์ง๋ ์์ต๋๋ค.
Historical feed, biosimilar examples, and regulatory limits from earlier lectures are not reused as public answers. This article uses a synthetic process difference; JMP is limited to displaying the relationship among Practical Difference, difference estimates, the two one-sided results, and the CI.
Example result statement
The prespecified equivalence margin was ยฑ2.0 percentage points for mean activity difference. The independent-batch difference BโA was 0.4, and the 90% CI corresponding to an ฮฑ=.05 TOST was โ1.1 to 1.9, entirely within the margins. Both one-sided tests were rejected, providing evidence for equivalence for this endpoint and analysis population. This result does not establish individual-batch specification conformity or equivalence of other quality attributes.
When evidence for equivalence is insufficient, do not automatically replace it with โnon-equivalent.โ Distinguish an inconclusive result from a clear margin exceedance according to the CI position.
At the end of this unit
- A nonsignificant difference test does not prove equivalence.
- The equivalence null hypothesis says the difference is outside the margins.
- TOST requires rejection of both one-sided null hypotheses.
- An ฮฑ=.05 TOST corresponds to a 90% CI lying entirely inside the margins.
- ฮ is justified before data using domain evidence.
- State conclusions with a limited endpoint and scope.
The next unit links effect, variation, ฮฑ, power, and sample size so that a study has enough precision to reach an active conclusion such as equivalence.
Official supplementary resources
This article explains equivalence principles and does not approve regulatory or clinical study designs.