A sample has mean response 10.2 and a 95% confidence interval for the mean of 9.5โ10.9. Does the next independent sample almost certainly fall inside it? Does the interval cover 95% of population observations?
No to both. A range for a mean, a range for one new observation, and a range covering most population observations answer different questions.
Are ranges for a mean, a new observation, and most of a population really the same question?
The previous unit established that a mean confidence interval targets the fixed parameter ฮผ. This unit prevents that interval from being stretched into every possible meaning of โrange.โ
State the target before the interval name
Memorizing only CI, PI, and TI makes them easy to swap. Complete the question first.
ํ ๋ชจ์ง๋จ ๋ชจ์์ ์ถ์ ๋ถํ์ค์ฑ
์ถ์ ๋ถํ์ค์ฑ + ๊ฐ๋ณ ๊ด์ธก์ ์์ฐ ๋ณ๋
coverage p + confidence ฮณ์ ๋ ์์ค
ํต๊ณ๊ฐ ์ถ์ ํ๋ ๋ฒ์๊ฐ ์๋๋ผ ์ฌ์ ๊ฒฐ์
- โWhere is the population mean ฮผ?โ โ confidence interval, CI
- โWhere might one next independent observation from the same process fall?โ โ prediction interval, PI
- โWhat range covers a specified proportion of population observations?โ โ tolerance interval, TI
All three may begin with the same xฬ and s, but they include different sources of uncertainty and therefore produce different results.
Under a common normal-model comparison at equal confidence, a mean CI is often narrower than a PI, while a TI requiring high coverage may be wider. Do not memorize width order alone: confidence, coverage, number of future observations, sidedness, and model change the width. The target defines the interval; width follows.
A confidence interval expresses uncertainty about a parameter
A 95% CI for a mean targets ฮผ. In repeated sampling, about 95% of intervals made by the same procedure contain ฮผ.
Because it describes precision of the estimated mean, the interval tends to narrow as independent n grows. Natural spread of individual observations does not disappear. Even a very narrow mean CI can exclude the next sample value.
A CI is not designed to cover the central 95% of individual values. Several raw points outside a narrow mean CI do not by themselves indicate a calculation error.
A prediction interval includes variation of the next observation
Predicting one next independent sample includes two uncertainties:
- uncertainty from estimating the population mean with a sample;
- natural variation of an individual value around the mean even if the mean were known.
That is why a PI for one next observation is generally wider than a mean CI from the same sample. In a representative normal model with unknown mean, a two-sided PI has structure:
xฬ ยฑ t* ร s ร โ(1 + 1/n)
Compared with s/โn in the mean CI, โ(1+1/n) preserves individual-observation variation. Even with very large n and precise mean estimation, a PI does not collapse to zero because individual spread remains.
โNextโ means a randomly obtained independent observation from the same defined process and conditions. Changes in instrument, material, operator, donor, culture condition, or time range weaken the basis for applying the original PI.
A tolerance interval targets a population proportion
A TI aims to cover a specified proportion of population observations, rather than one mean or one next observation. A complete TI statement therefore needs two percentages.
For example: โa two-sided tolerance interval containing at least 95% of the population with 95% confidence.โ
- coverage
p=95%: proportion of observations in one population the interval aims to contain; - confidence
ฮณ=95%: long-run confidence that repeated TI procedures achieve that coverage target.
You may abbreviate this as a 95%/95% TI, but identify which number is confidence and which is coverage. โA 95% tolerance intervalโ is incomplete.
A two-sided TI for a normal population is often written xฬ ยฑ kยทs. The factor k is not fixed at 1.96 or 3; it depends on n, confidence, coverage, and whether the limit is one- or two-sided.
If true ฮผ and ฯ were known for a normal population, ฮผยฑ1.96ฯ would cover about 95% of observations. In a real sample, ฮผ and ฯ are unknown and xฬ and s vary. A TI must include this estimation uncertainty and therefore uses an n-dependent, typically larger k.
If a normal model is inappropriate, its intended coverage guarantee may fail. Distribution-free methods exist, but may require much larger samples or wider intervals for the same confidence and coverage.
A specification is not a statistical interval
Specification limits or acceptance criteria are external requirements that a product, process, or test value must meet. They should be set before analysis from design needs, scientific purpose, quality risk, clinical meaning, or regulatory and contractual context.
CI, PI, and TI are calculated from a sample and model, so they move from sample to sample.
Comparing a TI with specification lines can help investigate whether a modeled population may lie within the requirements. It does not create these equalities:
TI = specificationTI inside specification โ automatic approvalpost-change values inside a pre-change TI โ comparability established
Comparability is a domain judgment about what must remain comparable, which quality attributes matter, and what differences and risks are acceptable. One historical TI cannot replace that full assessment.
Widening specifications after seeing data so every observation fits removes an independent decision criterion. Record the basis, decision time, and change history of every limit.
Compare three intervals from the same sample
The Lab draws independent samples from a normal process with ฮผ=10 and ฯ=3. All three intervals begin with the same xฬ and s but target different objects.
- At n=20, compare the mean CI, the PI for one next observation, and a 95% confidence/95% coverage TI.
- Press
New sampleand watch the center and intervals move together. - Compare n=5 with n=40 to see the role of estimation uncertainty.
์ฒ์์ด๋ผ๋ฉด: ๋ฌด์์ ๋๋ฌ์ผ ํ๋์?
- 1. ์ง๋ฌธ์ ๋จผ์ ์ฝ๊ธฐLab ์ ๋ชฉ์์ ์ด๋ฒ์ ๋น๊ตํ ํ ๊ฐ์ง๋ฅผ ํ์ธํฉ๋๋ค.
- 2. ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ๊ธฐ์ฒ์์๋ n, ํจ๊ณผ, ์ฐํฌ ๊ฐ์ ์ ๋ ฅ ์ค ํ๋๋ง ๋ฐ๊พธ์ญ์์ค.
- 3. New sample ๋๋ฅด๊ธฐ์ ํฉ์ฑ ๋ฐ์ดํฐ๊ฐ ๋ง๋ค์ด์ง๋๋ค. ๊ฐ์ ์กฐ๊ฑด๋ ํ๋ณธ์ ๋ฐ๋ผ ๋ฌ๋ผ์ง ์ ์์ต๋๋ค.
- 4. targets and widths of three intervals ๋น๊ตํ๊ธฐ๋ฐ๊พธ๊ธฐ ์ ํ ๋ฌด์์ด ์์ง์ด๊ณ ๋ฌด์์ด ๊ทธ๋๋ก์ธ์ง ํ ๋ฌธ์ฅ์ผ๋ก ์ ์ด๋ณด์ญ์์ค.
๋งํ๋ฉด ์ด๊ธฐํ๋ก ๋์๊ฐ ๊ธฐ๋ณธ ๊ฒฐ๊ณผ๋ฅผ ๋ณธ ๋ค ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ์ญ์์ค. ์ด Lab์ ์ ๋ต ํ์ ๊ธฐ๊ฐ ์๋๋ผ ํจํด ๊ด์ฐฐ ๋๊ตฌ์ ๋๋ค.
๊ฐ์ ํ๋ณธ์ผ๋ก ์๋ก ๋ค๋ฅธ ์ธ ์ง๋ฌธ์ ๋ตํด๋ณด์ธ์
์ ๊ท ์์ฑ ๊ณผ์ ฮผ=10, ฯ=3์ ์ฌ์ฉํฉ๋๋ค. ํ๊ท , ๋ค์ ํ ๊ด์ธก, ๋ชจ์ง๋จ ๊ด์ธก๊ฐ์ 95%๋ผ๋ ์ธ ๋์์ ๊ฐ์ ์ถ์์ ๋น๊ตํฉ๋๋ค.
๋์ผ ํ๋ณธ์ ์ธ ๊ตฌ๊ฐ
์ด๋ฒ ํ๋ณธ
95%/95% TI๋ ํ๋ณธ ๊ธฐ๋ฐ ์์ธก ์ ๊ท ํ์ฉ๊ตฌ๊ฐ์ ๊ทผ์ฌ๊ฐ์ ๋๋ค. ์ ๊ท์ฑยท๋ ๋ฆฝ ํ์ง ์กฐ๊ฑด์ด ๋ง์ง ์์ผ๋ฉด ๋ค๋ฅธ ๋ฐฉ๋ฒ์ด ํ์ํฉ๋๋ค.
CI: ์์ธก Student t ํ๊ท ๊ตฌ๊ฐ ยท PI: ์๋ ค์ง์ง ์์ ํ๊ท ์ ๋ค์ ํ ์ ๊ท ๊ด์ธก์ ๋ํ ์์ธก t ์์ธก๊ตฌ๊ฐ ยท TI: NIST Table 9.6 ๊ธฐ๋ฐ ์์ธก 95% confidence/95% coverage ์ ๊ท ํ์ฉ๊ณ์. ๊ต์ก์ฉ ํฉ์ฑ ์๋ฃ์ด๋ฉฐ ์น์ธยท๊ท๊ฒฉยทcomparability ํ๋จ์ ์ฌ์ฉํ ์ ์์ต๋๋ค.
The Lab reveals ฮผ and ฯ only because we defined the generating process. In actual analysis they are unknown, and the model and sampling scope must be justified.
The three JMP outputs begin with different input questions
Current JMP Distribution documentation provides separate results for Confidence Intervals, Prediction Intervals, and Tolerance Intervals for continuous variables. Read their targets, not a click sequence.
Mean๊ณผ Std Dev ๊ฐ์ ๋ชจ์์ ์ ๋ขฐํ๊ณ๋ฅผ ๊ตฌ๋ถํด ํ์ํฉ๋๋ค.
๋ค์ ๋ฌด์์ ๊ด์ธก ํ๋ ๋๋ ๋ค์ ํ๋ณธ์ ์์ฝ์ ๊ฒจ๋ฅํฉ๋๋ค.
Confidence Level๊ณผ Proportion to Cover๋ฅผ ๋ณ๋ ์ ๋ ฅ์ผ๋ก ๋ค๋ฃน๋๋ค.
For a Prediction Interval, distinguish one future observation from a summary of a future sample; confidence level and future sample count affect interpretation. Tolerance Interval output separately asks for confidence level and proportion to coverโthe two percentages in this unit.
Before and after reading JMP output, ask:
- Is my target a parameter, one future observation, or a population proportion?
- What are the independent units and population scope?
- Is the distributional model appropriate?
- Is a one-sided limit or two-sided interval needed?
- For a TI, did I record both confidence and coverage?
- Was the specification or acceptance criterion defined independently of these data?
A result sentence must include the target
Bad:
The 95% interval was 6.2โ14.1.
Better:
From 20 independent samples, the two-sided 95% Student t confidence interval for the population mean was 8.8โ11.6.
Under the same normal generating process, the two-sided 95% prediction interval for one next independent observation was 3.5โ16.9.
Under a normal model, the two-sided tolerance interval calculated to contain at least 95% of population observations with 95% confidence was 1.7โ18.7.
The numbers change across samples, but the sentence structure remains. Preserve the target, level, sidedness, model, and independent n.
Five statements to check before finishing
- A CI targets a parameter such as a mean.
- A PI includes natural variation of one next independent observation.
- A TI covers a specified population proportion with specified confidence.
ฮผยฑ1.96ฯis not the same as a sample-based 95%/95% TI.- Statistical intervals do not automatically set or approve specifications or comparability criteria.
The next unit moves from comparing ranges to asking whether an observed difference could arise from chance alone. We connect null and alternative hypotheses, p-values, ฮฑ, and the errors created by a decision rule.
Official supplementary resources
- NIST/SEMATECH ยท Tolerance intervals for a normal distribution
- NIST/SEMATECH ยท Prediction and tolerance intervals
- JMP Help ยท Prediction Intervals
- JMP Help ยท Confidence Interval for One Sample Mean
The numbers, figures, and In-Silico Lab in this article are synthetic material for explaining statistics. They cannot support real research, clinical, quality, regulatory, or comparability decisions.