Back to List

No Difference versus Equivalence: TOST and Equivalence Margins

Explains why a nonsignificant difference test does not establish equivalence, and how a prespecified margin, two one-sided tests, and a 90% CI form an equivalence decision.

Intermediate
|
28min
|
Verified (2026-08-14)
equivalenceTOSTequivalence margin90% confidence intervalnoninferiorityJMP
Progress0/28 (0%)

A difference test between the means of A and B gives p=.24. May we write, “There is no significant difference, so the two conditions are the same”? No. With a small sample and large variation, even a large difference can remain undetected.

This unit has one question.


How do we construct positive evidence that two conditions are sufficiently similar?
Equivalence limit ΔPrior boundaries of differences to be tolerated in practice
TOSTProcedure for rejecting both one-sided null hypotheses
90% CIα=.05 Difference interval corresponding to TOST
indeterminateResults insufficient to prove neither difference nor equivalence

Difference tests and equivalence tests have opposite null hypotheses

A difference test begins with H0: difference=0 and seeks evidence that a difference exists. Failure to reject means the data are compatible with 0; it does not prove that the difference lies within a small range.

An equivalence test first specifies a practically negligible boundary, −Δ and +Δ. If the mean difference is δ=μB−μA:

  • H0: δ≤−Δ or δ≥+Δ — the difference lies outside the acceptable range.
  • H1: −Δ<δ<+Δ — the difference lies inside the acceptable range.
U12 · Figure 01
The difference test and equivalence test have opposite directions of the null hypothesis.
−Δallowable area+ΔOutside CICI in fullequivalence
TOST requires us to reject both null hypotheses that the difference is less than or equal to −Δ or greater than or equal to +Δ.

TOST tests the following two null hypotheses separately at one-sided α.

  1. Reject H0 that B is at least Δ lower than A.
  2. Reject H0 that B is at least Δ higher than A.

Both must be rejected before evidence for equivalence exists. If either fails, the result may establish non-equivalence, or it may simply be inconclusive.

A TOST at α=.05 corresponds to a 90% CI

In general, if each one-sided test uses α, the same decision can be made by checking whether the (1−2α)×100% two-sided CI for the difference lies entirely inside the equivalence margins. At α=.05, this is a 90% CI.

  • the full CI lies within −Δ and +Δ: evidence for equivalence
  • the CI crosses a margin but includes 0: inconclusive for both difference and equivalence
  • the CI is far from 0 and outside a margin: a meaningful difference is possible

Do not use a conventional 95% CI interchangeably with TOST p-values. A 95% CI may also be presented for estimation, but state clearly which CI corresponds to which decision.

Δ is not a number chosen conveniently from the data

Widening a margin until the observed CI fits changes the question after the fact. Δ must be justified before seeing the data using scientific, clinical, or quality impact, measurement reliability, and prior evidence. Regulatory settings require the relevant guidance and expert review separately.

See the CI and margins before two small p-values

TOST output contains a lower-test and an upper-test p-value. The larger p-value must be below α for both conditions to pass. But placing the CI and margins on the same axis shows more directly which boundary failed and in which direction the effect lies.

Equivalence does not automatically establish sameness of an entire process, specification conformity of an individual product, equality of variances, or interchangeability. State equivalence narrowly: for which endpoint, population, time window, and effect scale?

In-Silico Lab: the same true difference can lead to different conclusions at different n

  1. Compare n=8 and n=100 with a true difference of 0 and Δ=1.
  2. Observe the region where a small n supports neither a difference test nor an equivalence test.
  3. Raise the true difference to 1.5 and see whether the CI center moves outside a margin.
  4. Explain why widening Δ without changing the data adds no scientific evidence.
In-Silico Lab · U12

See if the entire CI falls within the equivalence limits

α=.05 Use TOST and the corresponding 90% CI to determine why the difference, equality, and uncertainty questions are different.

If this is your first time: What should I press?
  1. 1. Read the question firstIn the Lab title, check the one thing you will compare this time.
  2. 2. Change just one conditionInitially, change only one of the inputs: n, effect, or spread.
  3. 3. new composite specimen pressureNew synthetic data is created. The same conditions may vary depending on the sample.
  4. 4. Pictures and calculation results CompareWrite in one sentence what moves and what stays the same before and after the change.

If it gets stuckresetGo back to see the default results and change just one condition. This Lab is not a correct answer tester but a pattern observation tool.

Synthetic observations of the same settings

−Δ · 90% CI · +Δ

calculation result

estimated difference1.67
90% CI0.94–2.41
equivalence limit±1.0
verdictLack of equivalence evidence

It is not enough for a CI to overlap a limit. Both one-sided null hypotheses must be rejected if the entire interval is within it.

educational synthetic modelbjs-comparison-sequence-v1. Actual research judgments must separately reflect experimental units, missingness, distribution, multiplicity, pre-planning, and domain criteria.

The Lab is an educational one-sample difference model with known σ. Actual two-sample, paired, proportion, or log-transformed endpoints need design-appropriate SEs and CIs.

JMP displays margins and CIs but does not justify the margin

Equivalence Test

View the difference CI and prior equivalence limits on the same axis.

Lower / Upper Test

Check whether both one-sided tests exceed the criterion.

Practical Difference

Statistical programs do not determine the scientific validity of limits.

Historical feed, biosimilar examples, and regulatory limits from earlier lectures are not reused as public answers. This article uses a synthetic process difference; JMP is limited to displaying the relationship among Practical Difference, difference estimates, the two one-sided results, and the CI.

Example result statement

The prespecified equivalence margin was ±2.0 percentage points for mean activity difference. The independent-batch difference B−A was 0.4, and the 90% CI corresponding to an α=.05 TOST was −1.1 to 1.9, entirely within the margins. Both one-sided tests were rejected, providing evidence for equivalence for this endpoint and analysis population. This result does not establish individual-batch specification conformity or equivalence of other quality attributes.

When evidence for equivalence is insufficient, do not automatically replace it with “non-equivalent.” Distinguish an inconclusive result from a clear margin exceedance according to the CI position.

At the end of this unit

  • A nonsignificant difference test does not prove equivalence.
  • The equivalence null hypothesis says the difference is outside the margins.
  • TOST requires rejection of both one-sided null hypotheses.
  • An α=.05 TOST corresponds to a 90% CI lying entirely inside the margins.
  • Δ is justified before data using domain evidence.
  • State conclusions with a limited endpoint and scope.
Non-significance is not evidence of equivalence. Only when the entire CI falls within the pre-justified Δ can the question “sufficiently similar” be answered.

The next unit links effect, variation, α, power, and sample size so that a study has enough precision to reach an active conclusion such as equivalence.

Official supplementary resources

This article explains equivalence principles and does not approve regulatory or clinical study designs.

💬 Questions & Comments

0 comments

You can post without signing in. Guest comments cannot be edited or deleted by their author.

0/2000

Loading...