Back to List

Creating and Questioning a Quadratic Model: Diagnostics Beyond Just Rยฒ

Construct a full quadratic model for the response surface according to the hierarchy and diagnose the model fit using ANOVA, lack of fit, residual, and Actual by Predicted.

Advanced
|
43min
|
Verified (2026-08-14)
quadratic modelmodel hierarchyANOVAparameter estimateslack of fitresidualActual by PredictedJMP
Progress0/28 (0%)

We filled the CCD table with response values and fitted a full quadratic model. If Rยฒ is 0.98, can we now find the optimal conditions? A high proportion of explained variance is just a starting point; it does not check for missing curvature, sequence effects, extrapolation, or selection bias.

The question for this section is:


How do you determine if a quadratic model adequately describes the data?
2์ฐจํ•ญxยฒ์ฒ˜๋Ÿผ ํ•œ ์š”์ธ์˜ ํœ˜์–ด์ง์„ ํ‘œํ˜„ํ•˜๋Š” ํ•ญ
quadratic hierarchy2์ฐจยท๊ตํ˜ธ์ž‘์šฉ์— ๊ด€๋ จ๋œ ๋‚ฎ์€ ์ฐจ์ˆ˜ ํ•ญ์„ ํ•จ๊ป˜ ์œ ์ง€
์ž”์ฐจ๊ด€์ธก ๋ฐ˜์‘์—์„œ ๋ชจํ˜• ์˜ˆ์ธก์„ ๋บ€ ๊ฐ’
๋ชจํ˜• ์ ํ•ฉ์„ฑ๋ชฉ์  ๋ฒ”์œ„์—์„œ ํŽธํ–ฅ ์—†์ด ์“ธ ๋งŒํ•œ์ง€์— ๋Œ€ํ•œ ์ข…ํ•ฉ ํŒ๋‹จ

First, link candidate models to the design

A two-factor full quadratic model has six terms.

ลท = ฮฒโ‚€ + ฮฒโ‚x + ฮฒโ‚‚y + ฮฒโ‚โ‚xยฒ + ฮฒโ‚‚โ‚‚yยฒ + ฮฒโ‚โ‚‚xy

The list of terms is not invented after looking at the results to find the best combination of p-values. It is pre-specified based on the design objective and scientific question, and then we check whether the design can estimate those terms. The first thing to check is whether the row with the attached response values deviates from the original design rows.

Maintain the quadratic hierarchy

If xยฒ is kept in the model, then x should also be kept; if xy is kept, then the main effects of x and y should also be kept. Deleting terms simply because the p-value of the lower-order term is large can change the meaning of the higher-order term depending on the coordinate origin and coding.

U24 ยท Figure 01
๋ˆ„๋ฝ๋œ ๊ณก๋ฅ ์€ ์ž”์ฐจ์— ๊ณก์„ ์œผ๋กœ ๋Œ์•„์˜ต๋‹ˆ๋‹ค
Actual by PredictedResidual by Predicted
Actual by Predicted์™€ residual์„ ๊ฐ™์€ ์ž๋ฃŒ์— ๋‚˜๋ž€ํžˆ ๋‘ก๋‹ˆ๋‹ค. ๋†’์€ ์„ค๋ช…๋น„์œจ๋„ ๊ตฌ์กฐ์  ์ž”์ฐจ๋ฅผ ์ง€์šฐ์ง€ ๋ชปํ•ฉ๋‹ˆ๋‹ค.

Compare three candidate models on the same data

Suppose we created a composite response as follows:

Y = 82 + 5x โˆ’ 3y โˆ’ 6xยฒ โˆ’ 4.2yยฒ + 3xy + error

Candidate ModelIncluded TermsExpected Issues
linear1, x, yCurvature and xy will remain in the residuals
quadratic without interaction1, x, y, xยฒ, yยฒThe xy structure may remain
full quadraticAll six termsFits the current generating function, but check for overfitting and range in real data

In a teaching fixture with zero noise, the full quadratic recovers the true coefficients. However, in real data, each coefficient estimate has an SE and CI, and the information content is not uniform across the design region.

ANOVA partitions variance but does not automatically approve a model

The Model F and p-value of ANOVA look at whether "this candidate model explains the signal better than a model with only an intercept." Each Parameter Estimate represents the coded effect with other terms fixed. The following questions remain to be answered separately:

  • Is there a selection bias in choosing the terms to fit the data?
  • Was pure error estimated using repeat points?
  • Have blocks, run order, and measurement batches been included in the model?
  • Does the variance and shape of the residuals vary depending on the conditions?
  • Are the locations for which you want to make predictions actually within the design region?

Rยฒ and adjusted Rยฒ summarize the proportion of variance explained, but they do not answer these questions.

Lack of fit and residuals provide different warnings

If there are repeated conditions, the residual SS can be partitioned into pure error and lack-of-fit SS. If the LOF is larger than the pure error, it is possible that the current model has missed the structure beyond the repeated variation. Conversely, a large LOF p-value does not prove that the model is correct. In small samples, the power may be low.

In residual plots, look for the following:

  • Does a U-shape or S-shape remain depending on the predicted value?
  • Does the spread widen in a fan shape as the response increases?
  • Is there a drift along the run order?
  • Does one point excessively pull the model?
Do not automate term deletion all at once

Iteratively deleting from the largest p-value does not reflect the selection process in the final p-value and prediction uncertainty. Maintain the prior model, hierarchy, and domain importance, and review the reduced model by comparing it with the full model and with new confirmation data.

Actual by Predicted reads the identity diagonal

The closer the points are to the diagonal, the closer the observed and fitted predictions are. However, fitting well at the design points is different from predicting new conditions well. If the repeated center points fall in the same direction, or if only the boundary points deviate significantly, consider the structure separately.

In-Silico Lab: Re-examine missing terms with residuals

  1. In the full quadratic model, examine the number of coefficients, Rยฒ, RMSE, and LOF SS.
  2. Change to a linear model and increase the curvature to see how the residuals and SSE change.
  3. In quadratic-no-interaction, increase the true xy coefficients.
  4. Change the noise and seed to see how much the small sample diagnostics fluctuate.
In-Silico Lab ยท U24

ํ›„๋ณด๋ชจํ˜•์„ ๋ฐ”๊พธ๊ณ  ์ž”์ฐจ๊ฐ€ ๋‚จ๊ธฐ๋Š” ๊ตฌ์กฐ๋ฅผ ๋ณด์„ธ์š”

๊ฐ™์€ CCD ํ•ฉ์„ฑ์ž๋ฃŒ์— linear, ์ œ๊ณฑํ•ญ ๋ชจํ˜•, full quadratic๋ฅผ ์ ํ•ฉํ•ด SSEยทLOFยท์ž”์ฐจ๋ฅผ ๋น„๊ตํ•ฉ๋‹ˆ๋‹ค.

์ฒ˜์Œ์ด๋ผ๋ฉด: ๋ฌด์—‡์„ ๋ˆŒ๋Ÿฌ์•ผ ํ•˜๋‚˜์š”?
  1. 1. ์งˆ๋ฌธ์„ ๋จผ์ € ์ฝ๊ธฐLab ์ œ๋ชฉ์—์„œ ์ด๋ฒˆ์— ๋น„๊ตํ•  ํ•œ ๊ฐ€์ง€๋ฅผ ํ™•์ธํ•ฉ๋‹ˆ๋‹ค.
  2. 2. ์กฐ๊ฑด ํ•˜๋‚˜๋งŒ ๋ฐ”๊พธ๊ธฐ์ฒ˜์Œ์—๋Š” n, ํšจ๊ณผ, ์‚ฐํฌ ๊ฐ™์€ ์ž…๋ ฅ ์ค‘ ํ•˜๋‚˜๋งŒ ๋ฐ”๊พธ์‹ญ์‹œ์˜ค.
  3. 3. ์ƒˆ ํ•ฉ์„ฑ ํ‘œ๋ณธ ๋ˆ„๋ฅด๊ธฐ์ƒˆ ํ•ฉ์„ฑ ๋ฐ์ดํ„ฐ๊ฐ€ ๋งŒ๋“ค์–ด์ง‘๋‹ˆ๋‹ค. ๊ฐ™์€ ์กฐ๊ฑด๋„ ํ‘œ๋ณธ์— ๋”ฐ๋ผ ๋‹ฌ๋ผ์งˆ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
  4. 4. ๊ทธ๋ฆผ๊ณผ ๊ณ„์‚ฐ ๊ฒฐ๊ณผ ๋น„๊ตํ•˜๊ธฐ๋ฐ”๊พธ๊ธฐ ์ „ํ›„ ๋ฌด์—‡์ด ์›€์ง์ด๊ณ  ๋ฌด์—‡์ด ๊ทธ๋Œ€๋กœ์ธ์ง€ ํ•œ ๋ฌธ์žฅ์œผ๋กœ ์ ์–ด๋ณด์‹ญ์‹œ์˜ค.

๋ง‰ํžˆ๋ฉด ์ดˆ๊ธฐํ™”๋กœ ๋Œ์•„๊ฐ€ ๊ธฐ๋ณธ ๊ฒฐ๊ณผ๋ฅผ ๋ณธ ๋’ค ์กฐ๊ฑด ํ•˜๋‚˜๋งŒ ๋ฐ”๊พธ์‹ญ์‹œ์˜ค. ์ด Lab์€ ์ •๋‹ต ํŒ์ •๊ธฐ๊ฐ€ ์•„๋‹ˆ๋ผ ํŒจํ„ด ๊ด€์ฐฐ ๋„๊ตฌ์ž…๋‹ˆ๋‹ค.

๊ฐ™์€ ์„ค์ •์˜ ํ•ฉ์„ฑ ๊ด€์ธก

run 10.11run 20.11run 3-0.12run 4-0.12run 5-0.22run 60.24

๊ณ„์‚ฐ ๊ฒฐ๊ณผ

๋ชจํ˜•ํ•ญ ์ˆ˜6
Rยฒ0.9984
RMSE0.280
LOF SS0.163

Rยฒ๊ฐ€ ๋†’์•„๋„ ์ž”์ฐจ์˜ ๊ณก์„ ยท์ˆœ์„œ ํŒจํ„ด๊ณผ lack of fit์„ ํ™•์ธํ•ฉ๋‹ˆ๋‹ค. ๋ฐ˜๋ณต์ ์ด ์—†์œผ๋ฉด pure error์™€ LOF๋ฅผ ๋ถ„๋ฆฌํ•  ์ˆ˜ ์—†์Šต๋‹ˆ๋‹ค.

๊ต์œก์šฉ synthetic model ยท bjs-response-surface-sequence-v1. ํ•œ ํ–‰์€ ๋ณ„๋„ ํ‘œ์‹œ๊ฐ€ ์—†๋Š” ํ•œ ํ•˜๋‚˜์˜ ๋…๋ฆฝ simulation ๋˜๋Š” ์„ค๊ณ„ run์ž…๋‹ˆ๋‹ค. ์‹ค์ œ ์—ฐ๊ตฌยทํ’ˆ์งˆยท๊ทœ์ œ ํŒ๋‹จ์—๋Š” ์‚ฌ์šฉํ•  ์ˆ˜ ์—†์Šต๋‹ˆ๋‹ค.

The Lab's 13 rows comprise 8 CCF coordinates and 5 center points. Model fitting is calculated using least squares and is not a substitute for using the menu or actual JMP values.

Read JMP output as a set

ANOVA / Parameter Estimates

์‚ฌ์ „ 2์ฐจ๋ชจํ˜•์˜ ํ•ญ๊ณผ ์ „์ฒด ๋ถˆํ™•์‹ค์„ฑ์„ ํ•จ๊ป˜ ๋ด…๋‹ˆ๋‹ค.

Lack of Fit

๋ฐ˜๋ณต์  pure error๋ฅผ ๊ธฐ์ค€์œผ๋กœ ๋‚จ์€ ๊ตฌ์กฐ๋ฅผ ์ ๊ฒ€ํ•ฉ๋‹ˆ๋‹ค.

Residual / Actual by Predicted

๊ณก๋ฅ  ๋ˆ„๋ฝยท์ด๋ถ„์‚ฐยท์ˆœ์„œยท์ด์ƒ์ ์„ ํ•œ ํ™”๋ฉด์— ์••์ถ•ํ•˜์ง€ ์•Š์Šต๋‹ˆ๋‹ค.

ANOVA and Parameter Estimates show the signal and coefficients, Lack of Fit shows the discrepancy compared to the replicate points, and Residual and Actual by Predicted show the shape. None of these are a single gate. If you have made a model change, leave a receipt with the initial model, the reason for the change, and the diagnostics before and after.

Example of a result statement

A pre-specified full quadratic model was fitted to the 13-run CCF data. The hierarchy was maintained, and pure error was estimated using the replicate center points. There were no obvious curvature or fan patterns in the Actual by Predicted and residual-by-predicted plots, and LOF was not significant in the current sample. This supports the adequacy of the candidate model within the study range, but does not replace a confirmation experiment under new conditions.

End of the section

  • The second-order model connects the design to the prior questions.
  • It adheres to the quadratic hierarchy.
  • Neither Rยฒ, p-value, nor lack of fit alone is a criterion for acceptance.
  • Residuals serve as a map for identifying curvature, heteroscedasticity, order, and points of influence.
  • Record the reduction and selection process, as well as the confirmation data.
2์ฐจ๋ชจํ˜•์€ ๋†’์€ Rยฒ ํ•˜๋‚˜๋กœ ์Šน์ธํ•˜์ง€ ์•Š์Šต๋‹ˆ๋‹ค. hierarchy, lack of fit, residual, ๋ฐ˜๋ณต์ ๊ณผ ์˜ˆ์ธก ๋ชฉ์ ์„ ํ•จ๊ป˜ ๋ณด๊ณ  ๋ถ€์กฑํ•˜๋ฉด ๋ชจํ˜•์ด๋‚˜ ์„ค๊ณ„๋ฅผ ๋ณด๊ฐ•ํ•ฉ๋‹ˆ๋‹ค.

In the next section, we will read the candidate model that has passed the diagnosis using contour plots, Profiler, and desirability, and then proceed to the confirmation experiment.

Official Supplementary Materials

This article and the Lab are synthetic materials for educational purposes and should not be used as evidence for actual research, process, quality, or regulatory decisions.

๐Ÿ’ฌ Questions & Comments

0 comments

You can post without signing in. Guest comments cannot be edited or deleted by their author.

0/2000

Loading...