Back to List

What Hypothesis Testing Actually Asks: Could Chance Alone Produce the Difference?

Understand whether an observed difference could arise from sampling variation through null and alternative hypotheses, p-values, significance level, and two kinds of error.

Beginner
|
48min
|
Verified (2026-08-14)
null hypothesisalternative hypothesisp-valuesignificance levelType I errorType II errorpowerJMP
Progress0/28 (0%)

Existing process A had mean yield 84.2 and improved process B had mean 85.5, a difference of 1.3. Because B is larger, can we immediately declare improvement?

Not yet. Even if the two population means are equal, different samples can have different sample means. First ask how easily a difference like 1.3 could appear in a reference world with no effect.

This unit asks one question.
Could sampling variation alone reasonably produce the observed difference?

A hypothesis test is not a machine that decides whether a cause exists. It defines a reference statistical model and measures how extreme the result is under that model, providing one part of the evidence.

Null hypothesis H0Starting model for testing for baseline difference or no effect
Alternative hypothesis H1The direction and form of difference targeted by the research question
p-valueProbability of an extreme outcome above the observed value below H0
Significance level αPrior criterion for type 1 error to be tolerated when H0 is true

At least two explanations lie behind an observed difference

When a sample shows a difference of 1.3:

  • population means may be equal and sampling variation produced 1.3;
  • population means may differ and that difference contributed to 1.3.

The first explanation becomes the reference world, the null hypothesis H0. For example, H0: μB−μA=0 states no population-mean difference.

The research alternative is H1. A non-directional question uses H1: μB−μA≠0; a prespecified improvement-only question might use H1: μB−μA>0.

Figure 01 · Two competing explanations
Observation differences can occur in H0 worlds as well as in H1 worlds.
H0 · Reference difference 0H1 · There is a real differenceobserved test statistic
Hypothesis testing does not directly calculate the probability that a world is true. First, we evaluate how unusual the current result is from the sampling distribution created by H0.

H0 and H1 are not labels attached after seeing favorable data. Connect the effect definition, direction, design, and analysis model to the research question in advance.

H0 is a calculation reference, not a preferred truth

We do not begin because we believe H0 is completely true in reality. It supplies a no-effect reference model for examining how compatible the current data are with that model.

A test statistic puts the difference on a common scale

A difference of 1.3 is difficult to judge alone. It carries different evidence when observations barely vary versus when they are widely dispersed.

A test statistic generally follows this intuition:

test statistic = difference between observed effect and H0 expectation / standard error of the effect

It expresses how far the effect lies from the H0 reference relative to sampling variation. The specific statistic may be t, F, χ², or a rank statistic depending on the question and structure. Later units choose among one, two, or many groups; continuous or categorical outcomes; and independent or paired designs.

A p-value is a tail probability under H0

Omitting its starting condition creates most p-value misunderstandings.

Assuming H0 and the test-model assumptions hold, the probability of a test statistic as extreme as or more extreme than the observed value in the direction defined by H1.

Figure 02 · Place of p-value
p-value is the area of ​​the extreme tails above the observation in the H0 distribution
Test statistic distribution under H0p/2p/2
A two-tailed test combines extremes in both directions. The p-value is not the probability that H0 is true, the probability that the result is a coincidence, or the size of the effect.

For a two-sided p=.03, under H0 and the model, about 3% of repeated results would have test statistics at least this extreme in either direction.

p=.03 does not mean:

  • H0 has a 3% probability of being true;
  • chance caused the result with 3% probability;
  • H1 has a 97% probability of being true;
  • 97% of repetitions will reproduce the result;
  • the effect is large or biologically important.
A p-value is not a property of the data alone

The same numerical table can produce different p-values under different test directions, independent versus paired structures, distributional models, missing-data rules, or test statistics. A p-value without what was tested loses its meaning.

Set significance level α before seeing the data

Significance level α sets the long-run willingness to make a Type I error: rejecting H0 when H0 is true. Choosing α=.05 uses a rule allowing about 5% false positives in repeated settings where H0 and all procedural assumptions hold.

p < αReject H0

Evidence that this is a rare result in the H0 standard world.

p ≥ αFailed to reject H0

Lack of evidence, not proof of ineffectiveness

effect sizeHow big is the difference?

Practical questions separate from statistical significance

confidence intervalWhich effects are compatible with the data?

Express the scope and direction of the estimate together

If p<α, the result is sufficiently extreme under H0 to reject H0. If p≥α, there is insufficient evidence to reject H0. We have not accepted or proved H0.

With small samples and large spread, an important effect can still produce p≥α. Therefore “not significant = no difference = equivalent” is invalid. Equivalence requires prespecified equivalence margins and an analysis designed for that question.

0.049 and 0.051 are not different universes

A threshold creates a decision rule, but evidence strength does not flip discontinuously at the boundary. Report the exact p-value with effect size, confidence interval, design quality, and practical importance.

One-sided and two-sided tests encode direction

A two-sided H1 treats both larger and smaller effects as differences. A right-sided H1 targets larger values; a left-sided H1 targets smaller values.

Figure 04 · Direction in advance
Two-sided tests two directions, and one-sided tests one direction determined by the study.
Two-sided H1: There is a differenceα/2α/2One-sided H1: largerα
If you look at the data and change it to a one-sided test in the favorable direction, the Type 1 error standard breaks down. When it comes to questions that cannot be overlooked, even significant changes in the opposite direction are natural for both sides.

A one-sided test places all α in one tail and can more readily detect an effect in the prespecified direction. It does not count a large effect in the opposite direction as evidence for that H1. Therefore:

  1. choose direction with the research question before seeing data;
  2. justify why the opposite direction is scientifically irrelevant or handled by a separate safety decision;
  3. report unexpected opposite-direction results.

Changing a two-sided test to a favorable one-sided test after seeing the mean can inflate the actual Type I error above the promised α.

The two error types are different failures

Figure 03 · Four lines of judgment
There are always two types of error possibilities in test results:
judgmentDo not reject H0Reject H0real worldH0 trueright decisionH0 trueType 1 error αH1 trueType 2 error βH1 truePower 1−β
α is the long-run ratio in which H0 is rejected when it is true. β and power can be calculated only after determining the specific real effect and are covered in detail in U13.
  • Type I error: reject a true H0; a false positive claiming an effect that is absent.
  • Type II error: fail to reject H0 when a specified H1 is true; a false negative missing a real effect.
  • Power 1−β: probability of correctly rejecting H0 under a specified H1.

α is set before analysis, but β is not one universal constant. It depends on true effect size, spread, n, test direction, and method. Larger effects, smaller spread, and more independent units are generally easier to detect. Unit 13 develops sample size and power.

Repetition changes p-values too

The Lab repeats a one-sample z test in a simple normal model with reference mean μ0=10 and known σ=3. It is an educational model connecting sampling variation, p-values, and errors—not a replacement for choosing a real analysis.

  1. At δ=0 with a two-sided test, inspect 80 p-values. Even with true H0, some may be below .05.
  2. Press New batch of 80 and watch false-positive counts vary.
  3. Increase δ through 0.5, 1.0, and 1.5; compare n=8 with n=32.
  4. With positive δ, choose a “smaller” one-sided H1 and observe the conflict between data direction and question direction.
If this is your first time: What should I press?
  1. 1. Read the question firstIn the Lab title, check the one thing you will compare this time.
  2. 2. Change just one conditionInitially, change only one of the inputs: n, effect, or spread.
  3. 3. New batch of 80 pressureNew synthetic data is created. The same conditions may vary depending on the sample.
  4. 4. p-value distribution and rejection count CompareWrite in one sentence what moves and what stays the same before and after the change.

If it gets stuckresetGo back to see the default results and change just one condition. This Lab is not a correct answer tester but a pattern observation tool.

In-Silico Lab · Hypothesis Testing

See how the p-value changes even in the same world

It is a normal model for education with a reference mean of μ0=10 and a known σ=3. Observe the true mean difference, independent n, and p-value for 80 replicates while changing the direction of the test.

80 p-values ​​repeated with the same settings

α=.05
p<.05 p≥.05

First sample and repeat results

first x̄9.857
first z-0.213
first p0.8312
first judgmentH0 cannot be rejected
80 times p<.055/80

Even if H0 is true, some samples will produce false positives at α=.05.

Educational one-sample z test: μ0=10 · known σ=3 · α=.05 · normal independent sampling · simulator_version:u07-ztest-guided-v1. In actual research, a method must be selected that reflects the unknown σ, design structure, test assumptions, multiplicity, etc.

Exactly four of 80 need not have p<.05. α=.05 is not a quota for every batch; it is a long-run error rate under many repetitions of the same procedure.

Distinguish three routes to a smaller p-value

  • A larger observed effect can increase the test statistic.
  • Smaller spread makes the same effect clearer.
  • Larger independent n reduces standard error, so even a tiny effect can yield a small p-value.

Thus p<.001 in a huge sample may accompany a trivial practical effect, while p=.08 in a small exploratory study does not rule out an important effect. Read effect size and CI in both cases.

JMP output separates parts of the evidence

JMP report names vary by analysis, but the core relationship joins the test statistic, p-value, H1, effect, and confidence interval.

Test Statistic

This is the value converted to the H0 standard scale by comparing the effect with the standard error.

Prob > |t| or p-value

Read the tail probability corresponding to the selected two-sided or one-sided H1.

Difference + Confidence Interval

View the direction and size of the difference and the compatible range along with the p-value.

Output such as Prob > |t| identifies the tails used in a two-sided t test. For a one-sided p-value, verify that the statistic sign matches the prespecified H1. Do not select whichever p-value appears smallest after analysis.

Figure 05 · black flow
α and hypothesis are determined before looking at the data, and the interpretation is determined along with the effect size.
questionEffect/DirectionhypothesisH0 · H1standardα·ModeldataStatistic pconclusionEffect · With CIDon't just leave p<α, but reconnect with questions, assumptions, and size of effect.
The p-value is one element of the overall judgment flow. Test selection, assumptions, missingness, multiplicity, and sample structure also affect conclusions.

Statistical software does not guarantee that:

  1. each row is the correct independent experimental unit;
  2. assumptions and sample selection are valid;
  3. only favorable analyses were not selected from many outcomes;
  4. the effect is biologically, clinically, or qualitatively important;
  5. a new experiment will reproduce the result.

Report more than reject or fail to reject

Bad:

The improved process works because p<.05.

Better:

In a prespecified two-sided analysis of independent batches, mean yield difference B−A was 1.3 percentage points, with 95% CI −1.8–4.4 and p=.39. This sample did not reject H0; it does not establish no improvement or equivalence of the processes.

Preserve the method, independent n, effect units, CI, exact p-value, test direction, and whether the analysis was prespecified. A p-value is part of an evidence sentence, not a one-character conclusion.

Seven statements to check before finishing

  • H0 is a no-effect reference model; H1 carries the research direction.
  • A test statistic scales the observed effect by its standard error under H0.
  • A p-value is the probability, under H0 and the model, of an equally or more extreme result.
  • A p-value is neither the probability H0 is true nor the effect size.
  • At p≥α, fail to reject H0; do not accept it.
  • α is the Type I error criterion under true H0 and is set before data.
  • A rejection decision cannot replace effect size, CI, design, and practical meaning.
The p-value is not the probability that an effect exists.How extreme are the current results in the baseline world created by H0 and the model?It shows and should be read along with effect size, CI, design, and biological significance.

The next unit asks whether spread is equal before comparing two group means. We examine equal-variance assumptions and why one variance-test result should not mechanically choose the entire analysis.

Official supplementary resources

The numbers, figures, and In-Silico Lab in this article are synthetic material for explaining statistics. They cannot support real research, clinical, quality, or regulatory decisions.

💬 Questions & Comments

0 comments

You can post without signing in. Guest comments cannot be edited or deleted by their author.

0/2000

Loading...