We filled the CCD table with response values and fitted a full quadratic model. If Rยฒ is 0.98, can we now find the optimal conditions? A high proportion of explained variance is just a starting point; it does not check for missing curvature, sequence effects, extrapolation, or selection bias.
The question for this section is:
How do you determine if a quadratic model adequately describes the data?
First, link candidate models to the design
A two-factor full quadratic model has six terms.
ลท = ฮฒโ + ฮฒโx + ฮฒโy + ฮฒโโxยฒ + ฮฒโโyยฒ + ฮฒโโxy
The list of terms is not invented after looking at the results to find the best combination of p-values. It is pre-specified based on the design objective and scientific question, and then we check whether the design can estimate those terms. The first thing to check is whether the row with the attached response values deviates from the original design rows.
Maintain the quadratic hierarchy
If xยฒ is kept in the model, then x should also be kept; if xy is kept, then the main effects of x and y should also be kept. Deleting terms simply because the p-value of the lower-order term is large can change the meaning of the higher-order term depending on the coordinate origin and coding.
Compare three candidate models on the same data
Suppose we created a composite response as follows:
Y = 82 + 5x โ 3y โ 6xยฒ โ 4.2yยฒ + 3xy + error
| Candidate Model | Included Terms | Expected Issues |
|---|---|---|
| linear | 1, x, y | Curvature and xy will remain in the residuals |
| quadratic without interaction | 1, x, y, xยฒ, yยฒ | The xy structure may remain |
| full quadratic | All six terms | Fits the current generating function, but check for overfitting and range in real data |
In a teaching fixture with zero noise, the full quadratic recovers the true coefficients. However, in real data, each coefficient estimate has an SE and CI, and the information content is not uniform across the design region.
ANOVA partitions variance but does not automatically approve a model
The Model F and p-value of ANOVA look at whether "this candidate model explains the signal better than a model with only an intercept." Each Parameter Estimate represents the coded effect with other terms fixed. The following questions remain to be answered separately:
- Is there a selection bias in choosing the terms to fit the data?
- Was pure error estimated using repeat points?
- Have blocks, run order, and measurement batches been included in the model?
- Does the variance and shape of the residuals vary depending on the conditions?
- Are the locations for which you want to make predictions actually within the design region?
Rยฒ and adjusted Rยฒ summarize the proportion of variance explained, but they do not answer these questions.
Lack of fit and residuals provide different warnings
If there are repeated conditions, the residual SS can be partitioned into pure error and lack-of-fit SS. If the LOF is larger than the pure error, it is possible that the current model has missed the structure beyond the repeated variation. Conversely, a large LOF p-value does not prove that the model is correct. In small samples, the power may be low.
In residual plots, look for the following:
- Does a U-shape or S-shape remain depending on the predicted value?
- Does the spread widen in a fan shape as the response increases?
- Is there a drift along the run order?
- Does one point excessively pull the model?
Iteratively deleting from the largest p-value does not reflect the selection process in the final p-value and prediction uncertainty. Maintain the prior model, hierarchy, and domain importance, and review the reduced model by comparing it with the full model and with new confirmation data.
Actual by Predicted reads the identity diagonal
The closer the points are to the diagonal, the closer the observed and fitted predictions are. However, fitting well at the design points is different from predicting new conditions well. If the repeated center points fall in the same direction, or if only the boundary points deviate significantly, consider the structure separately.
In-Silico Lab: Re-examine missing terms with residuals
- In the full quadratic model, examine the number of coefficients, Rยฒ, RMSE, and LOF SS.
- Change to a linear model and increase the curvature to see how the residuals and SSE change.
- In
quadratic-no-interaction, increase the true xy coefficients. - Change the noise and seed to see how much the small sample diagnostics fluctuate.
ํ๋ณด๋ชจํ์ ๋ฐ๊พธ๊ณ ์์ฐจ๊ฐ ๋จ๊ธฐ๋ ๊ตฌ์กฐ๋ฅผ ๋ณด์ธ์
๊ฐ์ CCD ํฉ์ฑ์๋ฃ์ linear, ์ ๊ณฑํญ ๋ชจํ, full quadratic๋ฅผ ์ ํฉํด SSEยทLOFยท์์ฐจ๋ฅผ ๋น๊ตํฉ๋๋ค.
์ฒ์์ด๋ผ๋ฉด: ๋ฌด์์ ๋๋ฌ์ผ ํ๋์?
- 1. ์ง๋ฌธ์ ๋จผ์ ์ฝ๊ธฐLab ์ ๋ชฉ์์ ์ด๋ฒ์ ๋น๊ตํ ํ ๊ฐ์ง๋ฅผ ํ์ธํฉ๋๋ค.
- 2. ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ๊ธฐ์ฒ์์๋ n, ํจ๊ณผ, ์ฐํฌ ๊ฐ์ ์ ๋ ฅ ์ค ํ๋๋ง ๋ฐ๊พธ์ญ์์ค.
- 3. ์ ํฉ์ฑ ํ๋ณธ ๋๋ฅด๊ธฐ์ ํฉ์ฑ ๋ฐ์ดํฐ๊ฐ ๋ง๋ค์ด์ง๋๋ค. ๊ฐ์ ์กฐ๊ฑด๋ ํ๋ณธ์ ๋ฐ๋ผ ๋ฌ๋ผ์ง ์ ์์ต๋๋ค.
- 4. ๊ทธ๋ฆผ๊ณผ ๊ณ์ฐ ๊ฒฐ๊ณผ ๋น๊ตํ๊ธฐ๋ฐ๊พธ๊ธฐ ์ ํ ๋ฌด์์ด ์์ง์ด๊ณ ๋ฌด์์ด ๊ทธ๋๋ก์ธ์ง ํ ๋ฌธ์ฅ์ผ๋ก ์ ์ด๋ณด์ญ์์ค.
๋งํ๋ฉด ์ด๊ธฐํ๋ก ๋์๊ฐ ๊ธฐ๋ณธ ๊ฒฐ๊ณผ๋ฅผ ๋ณธ ๋ค ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ์ญ์์ค. ์ด Lab์ ์ ๋ต ํ์ ๊ธฐ๊ฐ ์๋๋ผ ํจํด ๊ด์ฐฐ ๋๊ตฌ์ ๋๋ค.
๊ฐ์ ์ค์ ์ ํฉ์ฑ ๊ด์ธก
๊ณ์ฐ ๊ฒฐ๊ณผ
Rยฒ๊ฐ ๋์๋ ์์ฐจ์ ๊ณก์ ยท์์ ํจํด๊ณผ lack of fit์ ํ์ธํฉ๋๋ค. ๋ฐ๋ณต์ ์ด ์์ผ๋ฉด pure error์ LOF๋ฅผ ๋ถ๋ฆฌํ ์ ์์ต๋๋ค.
๊ต์ก์ฉ synthetic model ยท bjs-response-surface-sequence-v1. ํ ํ์ ๋ณ๋ ํ์๊ฐ ์๋ ํ ํ๋์ ๋ ๋ฆฝ simulation ๋๋ ์ค๊ณ run์ ๋๋ค. ์ค์ ์ฐ๊ตฌยทํ์งยท๊ท์ ํ๋จ์๋ ์ฌ์ฉํ ์ ์์ต๋๋ค.
The Lab's 13 rows comprise 8 CCF coordinates and 5 center points. Model fitting is calculated using least squares and is not a substitute for using the menu or actual JMP values.
Read JMP output as a set
์ฌ์ 2์ฐจ๋ชจํ์ ํญ๊ณผ ์ ์ฒด ๋ถํ์ค์ฑ์ ํจ๊ป ๋ด ๋๋ค.
๋ฐ๋ณต์ pure error๋ฅผ ๊ธฐ์ค์ผ๋ก ๋จ์ ๊ตฌ์กฐ๋ฅผ ์ ๊ฒํฉ๋๋ค.
๊ณก๋ฅ ๋๋ฝยท์ด๋ถ์ฐยท์์ยท์ด์์ ์ ํ ํ๋ฉด์ ์์ถํ์ง ์์ต๋๋ค.
ANOVA and Parameter Estimates show the signal and coefficients, Lack of Fit shows the discrepancy compared to the replicate points, and Residual and Actual by Predicted show the shape. None of these are a single gate. If you have made a model change, leave a receipt with the initial model, the reason for the change, and the diagnostics before and after.
Example of a result statement
A pre-specified full quadratic model was fitted to the 13-run CCF data. The hierarchy was maintained, and pure error was estimated using the replicate center points. There were no obvious curvature or fan patterns in the Actual by Predicted and residual-by-predicted plots, and LOF was not significant in the current sample. This supports the adequacy of the candidate model within the study range, but does not replace a confirmation experiment under new conditions.
End of the section
- The second-order model connects the design to the prior questions.
- It adheres to the quadratic hierarchy.
- Neither Rยฒ, p-value, nor lack of fit alone is a criterion for acceptance.
- Residuals serve as a map for identifying curvature, heteroscedasticity, order, and points of influence.
- Record the reduction and selection process, as well as the confirmation data.
In the next section, we will read the candidate model that has passed the diagnosis using contour plots, Profiler, and desirability, and then proceed to the confirmation experiment.
Official Supplementary Materials
- NIST/SEMATECH ยท Assessing the Model
- NIST/SEMATECH ยท Response Surface Model Example
- JMP Help ยท Fitting Linear Models
This article and the Lab are synthetic materials for educational purposes and should not be used as evidence for actual research, process, quality, or regulatory decisions.