Back to List

No Difference versus Equivalence: TOST and Equivalence Margins

Explains why a nonsignificant difference test does not establish equivalence, and how a prespecified margin, two one-sided tests, and a 90% CI form an equivalence decision.

Intermediate
|
28min
|
Verified (2026-08-14)
equivalenceTOSTequivalence margin90% confidence intervalnoninferiorityJMP
Progress0/19 (0%)

A difference test between the means of A and B gives p=.24. May we write, โ€œThere is no significant difference, so the two conditions are the sameโ€? No. With a small sample and large variation, even a large difference can remain undetected.

This unit has one question.


How do we construct positive evidence that two conditions are sufficiently similar?
๋™๋“ฑ์„ฑ ํ•œ๊ณ„ ฮ”์‹ค์งˆ์ ์œผ๋กœ ํ—ˆ์šฉํ•  ์ฐจ์ด์˜ ์‚ฌ์ „ ๊ฒฝ๊ณ„
TOST๋‘ ๋‹จ์ธก ๊ท€๋ฌด๊ฐ€์„ค์„ ๋ชจ๋‘ ๊ธฐ๊ฐํ•˜๋Š” ์ ˆ์ฐจ
90% CIฮฑ=.05 TOST์™€ ๋Œ€์‘ํ•˜๋Š” ์ฐจ์ด ๊ตฌ๊ฐ„
๋ถˆํ™•์ •์ฐจ์ด๋„ ๋™๋“ฑ์„ฑ๋„ ์ž…์ฆํ•˜๊ธฐ ๋ถ€์กฑํ•œ ๊ฒฐ๊ณผ

Difference tests and equivalence tests have opposite null hypotheses

A difference test begins with H0: difference=0 and seeks evidence that a difference exists. Failure to reject means the data are compatible with 0; it does not prove that the difference lies within a small range.

An equivalence test first specifies a practically negligible boundary, โˆ’ฮ” and +ฮ”. If the mean difference is ฮด=ฮผBโˆ’ฮผA:

  • H0: ฮดโ‰คโˆ’ฮ” or ฮดโ‰ฅ+ฮ” โ€” the difference lies outside the acceptable range.
  • H1: โˆ’ฮ”<ฮด<+ฮ” โ€” the difference lies inside the acceptable range.
U12 ยท Figure 01
์ฐจ์ด๊ฒ€์ •๊ณผ ๋™๋“ฑ์„ฑ๊ฒ€์ •์€ ๊ท€๋ฌด๊ฐ€์„ค์˜ ๋ฐฉํ–ฅ์ด ๋ฐ˜๋Œ€์ž…๋‹ˆ๋‹ค
โˆ’ฮ”ํ—ˆ์šฉ ์˜์—ญ+ฮ”CI ๋ฐ–CI ์ „์ฒด ์•ˆ๋™๋“ฑ์„ฑ
TOST์—์„œ๋Š” ์ฐจ์ด๊ฐ€ โˆ’ฮ” ์ดํ•˜์ด๊ฑฐ๋‚˜ +ฮ” ์ด์ƒ์ด๋ผ๋Š” ๋‘ ๊ท€๋ฌด๊ฐ€์„ค์„ ๋ชจ๋‘ ๊ธฐ๊ฐํ•ด์•ผ ํ•ฉ๋‹ˆ๋‹ค.

TOST tests the following two null hypotheses separately at one-sided ฮฑ.

  1. Reject H0 that B is at least ฮ” lower than A.
  2. Reject H0 that B is at least ฮ” higher than A.

Both must be rejected before evidence for equivalence exists. If either fails, the result may establish non-equivalence, or it may simply be inconclusive.

A TOST at ฮฑ=.05 corresponds to a 90% CI

In general, if each one-sided test uses ฮฑ, the same decision can be made by checking whether the (1โˆ’2ฮฑ)ร—100% two-sided CI for the difference lies entirely inside the equivalence margins. At ฮฑ=.05, this is a 90% CI.

  • the full CI lies within โˆ’ฮ” and +ฮ”: evidence for equivalence
  • the CI crosses a margin but includes 0: inconclusive for both difference and equivalence
  • the CI is far from 0 and outside a margin: a meaningful difference is possible

Do not use a conventional 95% CI interchangeably with TOST p-values. A 95% CI may also be presented for estimation, but state clearly which CI corresponds to which decision.

ฮ” is not a number chosen conveniently from the data

Widening a margin until the observed CI fits changes the question after the fact. ฮ” must be justified before seeing the data using scientific, clinical, or quality impact, measurement reliability, and prior evidence. Regulatory settings require the relevant guidance and expert review separately.

See the CI and margins before two small p-values

TOST output contains a lower-test and an upper-test p-value. The larger p-value must be below ฮฑ for both conditions to pass. But placing the CI and margins on the same axis shows more directly which boundary failed and in which direction the effect lies.

Equivalence does not automatically establish sameness of an entire process, specification conformity of an individual product, equality of variances, or interchangeability. State equivalence narrowly: for which endpoint, population, time window, and effect scale?

In-Silico Lab: the same true difference can lead to different conclusions at different n

  1. Compare n=8 and n=100 with a true difference of 0 and ฮ”=1.
  2. Observe the region where a small n supports neither a difference test nor an equivalence test.
  3. Raise the true difference to 1.5 and see whether the CI center moves outside a margin.
  4. Explain why widening ฮ” without changing the data adds no scientific evidence.
In-Silico Lab ยท U12

CI ์ „์ฒด๊ฐ€ ๋™๋“ฑ์„ฑ ํ•œ๊ณ„ ์•ˆ์— ๋“ค์–ด์˜ค๋Š”์ง€ ๋ณด์„ธ์š”

ฮฑ=.05 TOST์™€ ๋Œ€์‘ํ•˜๋Š” 90% CI๋ฅผ ์‚ฌ์šฉํ•ด ์ฐจ์ดยท๋™๋“ฑยท๋ถˆํ™•์ • ์งˆ๋ฌธ์ด ์™œ ๋‹ค๋ฅธ์ง€ ํ™•์ธํ•ฉ๋‹ˆ๋‹ค.

์ฒ˜์Œ์ด๋ผ๋ฉด: ๋ฌด์—‡์„ ๋ˆŒ๋Ÿฌ์•ผ ํ•˜๋‚˜์š”?
  1. 1. ์งˆ๋ฌธ์„ ๋จผ์ € ์ฝ๊ธฐLab ์ œ๋ชฉ์—์„œ ์ด๋ฒˆ์— ๋น„๊ตํ•  ํ•œ ๊ฐ€์ง€๋ฅผ ํ™•์ธํ•ฉ๋‹ˆ๋‹ค.
  2. 2. ์กฐ๊ฑด ํ•˜๋‚˜๋งŒ ๋ฐ”๊พธ๊ธฐ์ฒ˜์Œ์—๋Š” n, ํšจ๊ณผ, ์‚ฐํฌ ๊ฐ™์€ ์ž…๋ ฅ ์ค‘ ํ•˜๋‚˜๋งŒ ๋ฐ”๊พธ์‹ญ์‹œ์˜ค.
  3. 3. ์ƒˆ ํ•ฉ์„ฑ ํ‘œ๋ณธ ๋ˆ„๋ฅด๊ธฐ์ƒˆ ํ•ฉ์„ฑ ๋ฐ์ดํ„ฐ๊ฐ€ ๋งŒ๋“ค์–ด์ง‘๋‹ˆ๋‹ค. ๊ฐ™์€ ์กฐ๊ฑด๋„ ํ‘œ๋ณธ์— ๋”ฐ๋ผ ๋‹ฌ๋ผ์งˆ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
  4. 4. ๊ทธ๋ฆผ๊ณผ ๊ณ„์‚ฐ ๊ฒฐ๊ณผ ๋น„๊ตํ•˜๊ธฐ๋ฐ”๊พธ๊ธฐ ์ „ํ›„ ๋ฌด์—‡์ด ์›€์ง์ด๊ณ  ๋ฌด์—‡์ด ๊ทธ๋Œ€๋กœ์ธ์ง€ ํ•œ ๋ฌธ์žฅ์œผ๋กœ ์ ์–ด๋ณด์‹ญ์‹œ์˜ค.

๋ง‰ํžˆ๋ฉด ์ดˆ๊ธฐํ™”๋กœ ๋Œ์•„๊ฐ€ ๊ธฐ๋ณธ ๊ฒฐ๊ณผ๋ฅผ ๋ณธ ๋’ค ์กฐ๊ฑด ํ•˜๋‚˜๋งŒ ๋ฐ”๊พธ์‹ญ์‹œ์˜ค. ์ด Lab์€ ์ •๋‹ต ํŒ์ •๊ธฐ๊ฐ€ ์•„๋‹ˆ๋ผ ํŒจํ„ด ๊ด€์ฐฐ ๋„๊ตฌ์ž…๋‹ˆ๋‹ค.

๊ฐ™์€ ์„ค์ •์˜ ํ•ฉ์„ฑ ๊ด€์ธก

โˆ’ฮ” ยท 90% CI ยท +ฮ”

๊ณ„์‚ฐ ๊ฒฐ๊ณผ

์ถ”์ • ์ฐจ์ด1.67
90% CI0.94โ€“2.41
๋™๋“ฑ์„ฑ ํ•œ๊ณ„ยฑ1.0
ํŒ์ •๋™๋“ฑ์„ฑ ๊ทผ๊ฑฐ ๋ถ€์กฑ

CI๊ฐ€ ํ•œ๊ณ„์™€ ๊ฒน์น˜๋Š” ๊ฒƒ๋งŒ์œผ๋กœ๋Š” ๋ถ€์กฑํ•ฉ๋‹ˆ๋‹ค. ์ „์ฒด ๊ตฌ๊ฐ„์ด ์•ˆ์— ์žˆ์–ด์•ผ ๋‘ ๋‹จ์ธก ๊ท€๋ฌด๊ฐ€์„ค์„ ๋ชจ๋‘ ๊ธฐ๊ฐํ•ฉ๋‹ˆ๋‹ค.

๊ต์œก์šฉ synthetic model ยท bjs-comparison-sequence-v1. ์‹ค์ œ ์—ฐ๊ตฌ ํŒ๋‹จ์—๋Š” ์‹คํ—˜๋‹จ์œ„, ๊ฒฐ์ธก, ๋ถ„ํฌ, ๋‹ค์ค‘์„ฑ, ์‚ฌ์ „๊ณ„ํš๊ณผ ๋„๋ฉ”์ธ ๊ธฐ์ค€์„ ๋ณ„๋„๋กœ ๋ฐ˜์˜ํ•ด์•ผ ํ•ฉ๋‹ˆ๋‹ค.

The Lab is an educational one-sample difference model with known ฯƒ. Actual two-sample, paired, proportion, or log-transformed endpoints need design-appropriate SEs and CIs.

JMP displays margins and CIs but does not justify the margin

Equivalence Test

์ฐจ์ด CI์™€ ์‚ฌ์ „ ๋™๋“ฑ์„ฑ ํ•œ๊ณ„๋ฅผ ๊ฐ™์€ ์ถ•์—์„œ ๋ด…๋‹ˆ๋‹ค.

Lower / Upper Test

๋‘ ๋‹จ์ธก ๊ฒ€์ •์ด ๋ชจ๋‘ ๊ธฐ์ค€์„ ๋„˜์—ˆ๋Š”์ง€ ํ™•์ธํ•ฉ๋‹ˆ๋‹ค.

Practical Difference

ํ†ต๊ณ„ ํ”„๋กœ๊ทธ๋žจ์ด ํ•œ๊ณ„์˜ ๊ณผํ•™์  ํƒ€๋‹น์„ฑ์„ ์ •ํ•ด์ฃผ์ง€๋Š” ์•Š์Šต๋‹ˆ๋‹ค.

Historical feed, biosimilar examples, and regulatory limits from earlier lectures are not reused as public answers. This article uses a synthetic process difference; JMP is limited to displaying the relationship among Practical Difference, difference estimates, the two one-sided results, and the CI.

Example result statement

The prespecified equivalence margin was ยฑ2.0 percentage points for mean activity difference. The independent-batch difference Bโˆ’A was 0.4, and the 90% CI corresponding to an ฮฑ=.05 TOST was โˆ’1.1 to 1.9, entirely within the margins. Both one-sided tests were rejected, providing evidence for equivalence for this endpoint and analysis population. This result does not establish individual-batch specification conformity or equivalence of other quality attributes.

When evidence for equivalence is insufficient, do not automatically replace it with โ€œnon-equivalent.โ€ Distinguish an inconclusive result from a clear margin exceedance according to the CI position.

At the end of this unit

  • A nonsignificant difference test does not prove equivalence.
  • The equivalence null hypothesis says the difference is outside the margins.
  • TOST requires rejection of both one-sided null hypotheses.
  • An ฮฑ=.05 TOST corresponds to a 90% CI lying entirely inside the margins.
  • ฮ” is justified before data using domain evidence.
  • State conclusions with a limited endpoint and scope.
์œ ์˜ํ•˜์ง€ ์•Š์Œ์€ ๋™๋“ฑ์„ฑ์˜ ์ฆ๊ฑฐ๊ฐ€ ์•„๋‹™๋‹ˆ๋‹ค. ์‚ฌ์ „์— ์ •๋‹นํ™”ํ•œ ฮ” ์•ˆ์— ์ „์ฒด CI๊ฐ€ ๋“ค์–ด์™€์•ผ โ€œ์ถฉ๋ถ„ํžˆ ๋น„์Šทํ•˜๋‹คโ€๋Š” ์งˆ๋ฌธ์— ๋‹ตํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

The next unit links effect, variation, ฮฑ, power, and sample size so that a study has enough precision to reach an active conclusion such as equivalence.

Official supplementary resources

This article explains equivalence principles and does not approve regulatory or clinical study designs.

๐Ÿ’ฌ Questions & Comments

0 comments

You can post without signing in. Guest comments cannot be edited or deleted by their author.

0/2000

Loading...