Suppose processes A and B both have mean yield 82. Most A values lie from 80 to 84, while B ranges broadly from 70 to 94. Their means look identical, but predictability of the next batch and the standard error of a mean difference are not the same at all.
This unit asks one question.
How does uncertainty in a mean comparison change when two groups have different spread?
Spread is background noise for a mean comparison
An observed difference between two means must be compared with natural variation within each group. Tightly clustered values can make a small mean difference clear; widely spread values can bury the same difference in sampling fluctuation. Assuming equal population variances is homoscedasticity; assuming different variances is heteroscedasticity.
Equal variance does not mean two sample SDs must match exactly. Samples from the same population variance have different SDs each time. Conversely, two accidentally similar SDs do not guarantee equal uncertainty in a mean comparison when distribution shapes or sample sizes differ.
First place dot plots, box plots, and SDs for each group side by side. Ask whether extra width reflects an overall scale change, one or two outliers, a right tail, or a mixture of subgroups. A variance test does not replace these plots.
Spread tests do not ask the same question in the same way
A common null is approximately H0: σ₁²=σ₂²=…=σk², but test statistics measure spread differently and have different distributional assumptions.
- A two-group F test compares the ratio of sample variances with an F distribution; it is sensitive to the normal model.
- The Bartlett test can be efficient across groups but is sensitive to non-normality and outliers.
- The Levene test compares absolute distances of observations from their group center as a new response.
- A Levene variant using the median rather than the mean center is often called Brown–Forsythe and can be more robust to tails and outliers.
Rather than memorize names, ask which center and distance each method uses and how sensitive it is to departures from normality.
A non-significant spread test means the current sample lacks evidence to detect a variance difference. Small n can miss even a large variance ratio. Do not use it as proof of equality or automatic permission for a pooled t-test.
Do not split the mean test by one preliminary test
An older common procedure first tested equal variances, then chose Welch t-test when significant and pooled t-test otherwise. This two-step automatic branch complicates final error rates and uncertainty, and is unstable in small samples.
Choose the analysis before looking at data, using design and scientific assumptions. For two independent group means, Welch's method separately reflects group variances and n, making it a useful robust option when equal variance is uncertain. It is not the sole answer for every design: paired, clustered, repeated-measures, or censored data need a structure-matched model.
In-Silico Lab: keep the mean fixed and change only width
In the Lab, both groups have the same true mean. Change only B's SD ratio, n, and seed.
- Raise the SD ratio from 1 to 3 and see that the sample-SD ratio is not exactly the set value every time.
- Repeat new samples at n=8 and n=40 and compare fluctuation in estimated spread.
- Find a sample where SD and median absolute deviation do not move in the same direction.
- Answer again whether equal means imply equal stability.
Fix the mean and just change the spread.
We vary the SD ratio of B in two groups with the same mean and see how the sample SD and robust distance change.
If this is your first time: What should I press?
- 1. Read the question firstIn the Lab title, check the one thing you will compare this time.
- 2. Change just one conditionInitially, change only one of the inputs: n, effect, or spread.
- 3. new composite specimen pressureNew synthetic data is created. The same conditions may vary depending on the sample.
- 4. Pictures and calculation results CompareWrite in one sentence what moves and what stays the same before and after the change.
If it gets stuckresetGo back to see the default results and change just one condition. This Lab is not a correct answer tester but a pattern observation tool.
Synthetic observations of the same settings
calculation result
Differences in sample spread are evidence created by the data. A single p value cannot prove that the variances of two populations are equal.
educational synthetic modelbjs-comparison-sequence-v1. Actual research judgments must separately reflect experimental units, missingness, distribution, multiplicity, pre-planning, and domain criteria.
The Lab does not make spread-test p-values compete. It first shows which data-generating conditions change which spread summaries. With real data, also inspect raw-data plots, sample size, causes of outliers, measurement limits, and independence.
JMP output puts several sensitivities side by side
F·Levene·Brown-Forsythe-type results look at the same spread question with different sensitivities.
First check the size and uncertainty of the sample SD graphically.
It is judged based on distribution shape, n, and outlier values rather than just the p value.
The goal is not to follow a historical JMP menu path. Read how each result examines the same H0, and how much Std Dev and CIs show the actual scale. Do not pick the most favorable p-value among several tests.
Separate spread and mean questions in reporting
Poor statement:
Levene p=.31, so the two variances are equal and we used a pooled t-test.
Better statement:
Sample SDs for independent batches were 1.8 in A and 3.1 in B. B showed wider spread and a right tail. We estimated the mean difference using the pre-specified Welch method and presented the spread comparison as sensitivity information with raw-data plots and a median-based Levene result.
If variance difference itself is the research question, report a CI for a variance ratio or scale difference. If it is an assumption check for a mean comparison, the mean effect and CI are central and the spread test is supporting diagnosis.
Takeaways
- Equal variance is a population-model assumption, not matching sample SDs.
- Spread differences change standard error and predictability for mean differences.
- F and Bartlett are sensitive to departures from normality; Levene methods use center distances for more robustness.
- A non-significant spread test does not prove equal variances.
- Do not mechanically change analysis after a spread test.
- Raw-data plots and experimental units come before every test.
The next unit compares one group mean with an external target and joins effect, standard error, CI, and p-value in one result statement.
Official supplementary resources
The values, figures, and Lab in this article are educational synthetic material. They cannot be the sole basis for real research, clinical, quality, or regulatory decisions.