Suppose experiments A and B both have mean response 10. If a report gives only the mean, they look identical. But A may cluster tightly around 10 while B spreads from 2 to 18. Did we really observe the same result?
If two datasets have the same mean, are they the same data?
No. A mean usefully summarizes a center, but does not reveal distance from that center, a long tail, or a mixture of groups. This does not make means useless. It means we must distinguish what a mean says from what it cannot say.
You need not memorize every term at once. Keep five terms below as a compass; add others when a picture makes them necessary.
Before calculation: what does one row represent?
Before calculating a mean or SD, check what a table row means. Measuring three different culture wells once each and reading one well three times can both create three rows, but they are not the same evidence.
The experimental unit is the smallest unit receiving treatment independently. Repeats from different experimental units are biological replicates; repeated readings of one unit are technical replicates. Our calculation table first follows an explicit rule for technical repeats, then represents one independent unit per row.
Two collections with exactly the same mean
Let each unit be a relative response. A and B below are synthetic teaching values, not research data.
| Set | Observations | Mean | Median | Sample variance | Sample SD |
|---|---|---|---|---|---|
| A | 8, 9, 10, 11, 12 | 10 | 10 | 2.5 | 1.58 |
| B | 2, 6, 10, 14, 18 | 10 | 10 | 40 | 6.32 |
The mean is the sum divided by count. Both sets total 50 across five values, so both means are 10. This fast center summary erases distance between points, so spread must sit beside it.
Turn distance from the mean into a number
Subtracting the mean from each observation gives a deviation. A has โ2, โ1, 0, 1, 2; B has โ8, โ4, 0, 4, 8. B is visibly wider. Adding raw deviations gives zero in both sets because left and right cancel. Square deviations to preserve distance and remove sign; points farther away get more weight.
Why divide by nโ1 rather than n?
When a sample uses its own mean as its center, spread tends to look smaller than it would around the unknown population center. The sample mean has already been fitted as the closest center to that sample. Dividing sample variance by nโ1 corrects this tendency.
A proof of nโ1 belongs to later work on degrees of freedom and expectation. For now, the useful intuition is that estimating one center leaves nโ1 independent pieces of distance information.
Aโs sample SD is about 1.58; Bโs is about 6.32. They share a mean but B has much larger typical spread. Do not conclude automatically that A is the better experiment: a small SD can arise from detection limits, dependent technical repeats, or a narrowly selected sample range.
Mean and SD still do not show every shape
Even with center and spread, the overall shape remains. Values can be balanced around a center, have a long right tail, or form two clusters. This shape is the distribution.
Two peaks do not prove two cell types were mixed. Different batches, treatment times, equipment states, or hidden subgroups are possibilities; small samples or bin choices can also make that appearance. A graph is not a verdict on cause. It is a map of questions to check next.
Use the same x-axis range and bin boundaries for both groups. With small samples, small bin changes can change appearance greatly, so inspect raw points as well as the histogram.
What do median and box plots add?
The median is the middle value after ordering values, and is less pulled by a far extreme than the mean. Q1 and Q3 are the 25% and 75% positions; their difference, IQR, is the width occupied by the middle 50%.
A box plot compresses this order information: box ends are Q1 and Q3, its internal line is the median, whiskers show ordinary-range ends, and points beyond them may be marked as outlier candidates.
First check sample ID, raw record, units, instrument log, timing, and conditions. If an error is confirmed, follow a pre-specified rule and document the reason; if it is valid, do not remove it casually. Record how conclusions differ with and without the point.
Now make the same mean yourself
So far the values were chosen by the author. Real samples vary. In the mini Lab, change Bโs spread and shape. Both group means are fixed exactly at 10 for teaching, so this is constructed comparison data, not a random sample.
ํ๊ท 10์ ๊ณ ์ ํ๊ณ , B์ ๋ชจ์์ ๋ฐ๊ฟ๋ณด์ธ์
๊ธฐ์ค ๋น๊ต์์๋ ํ๊ท ์ ๊ณ ์ ํด ์ฐํฌ์ ๋ชจ์์ ์ฐจ์ด์ ์ง์คํฉ๋๋ค. ์๋ ๋ฐ์ดํฐ๋ ์ฐ๊ตฌ ๊ฒฐ๊ณผ๊ฐ ์๋๋ผ ํต๊ณ ๊ฐ๋ ์ ๊ด์ฐฐํ๋๋ก ๊ตฌ์ฑํ ๊ต์ก์ฉ ํฉ์ฑ ๋ฐ์ดํฐ์ ๋๋ค.
์ฒ์์ด๋ผ๋ฉด: ๋ฌด์์ ๋๋ฌ์ผ ํ๋์?
- 1. ์ง๋ฌธ์ ๋จผ์ ์ฝ๊ธฐLab ์ ๋ชฉ์์ ์ด๋ฒ์ ๋น๊ตํ ํ ๊ฐ์ง๋ฅผ ํ์ธํฉ๋๋ค.
- 2. ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ๊ธฐ์ฒ์์๋ n, ํจ๊ณผ, ์ฐํฌ ๊ฐ์ ์ ๋ ฅ ์ค ํ๋๋ง ๋ฐ๊พธ์ญ์์ค.
- 3. ๋ถํฌ ๋ชจ์ ๋๋ ์ฐํฌ ์กฐ๊ฑด ๋ฐ๊พธ๊ธฐ ๋๋ฅด๊ธฐ์ ํฉ์ฑ ๋ฐ์ดํฐ๊ฐ ๋ง๋ค์ด์ง๋๋ค. ๊ฐ์ ์กฐ๊ฑด๋ ํ๋ณธ์ ๋ฐ๋ผ ๋ฌ๋ผ์ง ์ ์์ต๋๋ค.
- 4. ๋ ์ง๋จ์ ์ ๊ทธ๋ฆผ๊ณผ ์์ฝ๊ฐ ๋น๊ตํ๊ธฐ๋ฐ๊พธ๊ธฐ ์ ํ ๋ฌด์์ด ์์ง์ด๊ณ ๋ฌด์์ด ๊ทธ๋๋ก์ธ์ง ํ ๋ฌธ์ฅ์ผ๋ก ์ ์ด๋ณด์ญ์์ค.
๋งํ๋ฉด ์ด๊ธฐํ๋ก ๋์๊ฐ ๊ธฐ๋ณธ ๊ฒฐ๊ณผ๋ฅผ ๋ณธ ๋ค ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ์ญ์์ค. ์ด Lab์ ์ ๋ต ํ์ ๊ธฐ๊ฐ ์๋๋ผ ํจํด ๊ด์ฐฐ ๋๊ตฌ์ ๋๋ค.
์์๋ฃ์ ๊ณผ ํ๊ท ์
์ด๋ฒ ๋น๊ต์ ์์ฝ
์ฐ๊ตฌยท์์ยทํ์ง ํ๋จ์ ์ฌ์ฉํ์ง ๋ง์ญ์์ค. ์ ์ Lab์์๋ ์์ฑ ๋ชจ๋, seed, simulator version, ๋จ์, ํ ํ์ ์๋ฏธ, raw/analysis-ready ํ์ ๋ฐ์ดํฐ ์ฌ์ ์ ํจ๊ป ์ ๊ณตํฉ๋๋ค.
As shape changes, Bโs mean remains 10 while point width, tail, and clustering change. In the full Labโs free-exploration mode, even with target center 10, each actual sample mean will not be exactly 10. That is normal: a generating target and a summary of one chance sample differ.
How does JMP express this theory?
JMP appears here not to teach button sequences, but to connect the relationships we have seen to output. For continuous data, JMP Distribution shows three connected groups of results.
๊ฐ๋ค์ด ๋์ธ ์ ์ฒด ๋ชจ์๊ณผ ์ค์ฌยท๊ฐ์ด๋ฐ ํญยท์ด์๊ฐ ํ๋ณด๋ฅผ ํจ๊ป ๋ด ๋๋ค.
์ต์๊ฐ, Q1, ์ค์๊ฐ, Q3, ์ต๋๊ฐ์ฒ๋ผ ์์์ ๊ธฐ๋ฐํ ์์น๋ฅผ ํ์ธํฉ๋๋ค.
ํ๊ท , ํ์คํธ์ฐจ์ N์ฒ๋ผ ์ค์ฌ๊ณผ ์ฐํฌ๋ฅผ ์์น๋ก ํ์ธํฉ๋๋ค.
For A and B, first verify N equals the number of independent units. Then read mean and SD together in Summary Statistics; inspect width, tails, and clusters in the histogram and raw points; and use Quantiles and Outlier Box Plot for median, IQR, and outlier candidates.
Do not stop because both means are 10. A larger B SD and wider histogram are โsame center, different distance.โ Check the history of a whisker-out point rather than deleting it. A two-peaked plot supports a hypothesis to check batches or hidden conditions, not confirmation of a cause.
You can follow the unit with the theory and representative results alone. If you have JMP or Minitab, use a table copied from the Lab to inspect the same questions of center, spread, and distribution. The statistical relationships are the same despite different screen layouts.
Four questions to place beside every mean
- How many independent n entered this mean? Ensure technical-repeat rows were not counted as independent samples.
- How far do values spread from the center? Read SD with the width of raw points.
- What is the distribution shape? Use same-axis, same-bin histograms to inspect tails and clusters.
- What does an unusual point require? Rechecking sample and measurement history, not automatic deletion.
A mean is not wrong; it compresses too much information into one number. Good analysis restores spread, distribution, raw data, and replication structure beside it.
The next unit widens the question. How well do the mean and SD in hand describe the full population? Why do values change when an experiment is repeated, and how can that movement be expressed?
Official supplementary resources
- JMP Help ยท Distributions of Continuous Variables
- JMP Statistics Knowledge Portal ยท Box Plot
- NIST/SEMATECH e-Handbook ยท Measures of Scale
The examples and mini Lab in this article are synthetic data for explaining statistical concepts. They cannot be used as evidence for real research, clinical, quality, or regulatory decisions.