If you explore all six continuous factors with a full quadratic model, you will have 28 coefficients. If you don't have enough runs to precisely estimate all of those terms from the beginning, how can you find the most important factors and curvature candidates first?
The question for this section is:
Can we screen many continuous factors with a small number of runs while also exploring curvature?
DSD looks at each continuous factor at three levels
In the Definitive Screening Design (DSD), the continuous factors are coded as โ1, 0, and +1. Unlike a typical two-level screening design, each factor has a 0 level, which allows you to explore the curvature information of the individual quadratic terms.
The six-factor teaching design consists of 13 runs. It can be created by stacking the rows of the 6ร6 conference matrix C, the rows of the sign-inverted matrix โC, and the all-center row. Each factor column is balanced, and the main effect columns are orthogonal to each other.
This geometry provides an important advantage, but it does not mean that "all 28 full quadratic coefficients can be freely estimated with 13 runs." There may still be correlations between several two-factor interactions and quadratic term candidates.
Read the Correlation Map before looking at the response values
The correlation map of the design shows the correlations between the candidate model columns. In DSD, the main effects are designed to be separated from the two-factor interactions, but not all higher-order candidates are independent of each other.
Before obtaining the responses, check the following:
- Are the main effects separated from the nuisance/block terms?
- How much are the quadratic terms related to each other or to the 2FI?
- How does the basic structure change when categorical factors are mixed in?
- Which terms are reinforced by the center/extra runs?
We evaluate the actual model matrix, not just the name of the design.
Stepwise analysis uses sparsity and heredity
Since we cannot put all candidate terms into the 13 data points at once, we ask questions step by step.
- Stage 1 โ Main Effects: We look at the linear signal candidates for each continuous factor.
- Stage 2 โ Curvature: We review the promising factors and quadratic term candidates.
- Stage 3 โ Interaction: We review the selected main effects and scientifically plausible 2FI in a hierarchy.
- Stage 4 โ Augmentation: We determine additional runs to separate ambiguous terms or optimize.
This flow often relies on the assumptions that most effects are small (sparsity) and that important interactions are likely to involve related main effects (heredity). Both are useful analytical principles, but they are not facts guaranteed by the data.
Selecting candidates from the same data and then testing them again creates selection bias. Keep a record of the steps, thresholds, and candidate terms, and verify the results with shrinkage, sensitivity analysis, additional runs, and independent confirmation.
Even if the main effects are small, the crossover interaction can be large.
Two factors may have opposing interactions, resulting in a mean main effect close to zero. If only strong heredity is assumed, such terms may be overlooked. Mechanistically important combinations can be included in the initial candidate set regardless of their main effect rank.
In small samples, if half-normal or stepwise-selection results vary with the seed or noise, report the stability of the selection rather than only one selected list.
If curvature candidates exist, move to RSM or augmentation
If the set of promising factors is reduced to 2โ3 and there is evidence of curvature, evaluate whether an RSM can be fitted using those factors. If information is lacking, augment with axial points, center points, or selected candidate runs. If many factors remain promising or several interactions remain unresolved, do not force optimization with only 13 runs.
DSD is a bridge between screening and optimization, not a final design for all cases.
In-Silico Lab: Observe the fluctuation of stepwise candidates
- Change the true main effect factors between A and F, and check the top factors in Stage 1.
- Change the true curvature factors and see if the Stage 2 candidates follow.
- Increase the CรD interaction to see if the stepwise ranking fluctuates.
- Change the seed and noise to observe the selection stability.
DSD์ ์ฃผํจ๊ณผ ๋จ๊ณ์ ๊ณก๋ฅ ๋จ๊ณ๋ฅผ ๋ถ๋ฆฌํ์ธ์
6์์ธ 13-run ํฉ์ฑ DSD์์ ์ฐธ ์ฃผํจ๊ณผยท๊ณก๋ฅ ยท๊ตํธ์์ฉยทnoise๋ฅผ ๋ฐ๊พธ๋ฉฐ ๋จ๊ณ๋ณ ์์ ํ๋ณด๊ฐ ์ผ๋ง๋ ์์ ์ ์ธ์ง ๋ด ๋๋ค.
์ฒ์์ด๋ผ๋ฉด: ๋ฌด์์ ๋๋ฌ์ผ ํ๋์?
- 1. ์ง๋ฌธ์ ๋จผ์ ์ฝ๊ธฐLab ์ ๋ชฉ์์ ์ด๋ฒ์ ๋น๊ตํ ํ ๊ฐ์ง๋ฅผ ํ์ธํฉ๋๋ค.
- 2. ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ๊ธฐ์ฒ์์๋ n, ํจ๊ณผ, ์ฐํฌ ๊ฐ์ ์ ๋ ฅ ์ค ํ๋๋ง ๋ฐ๊พธ์ญ์์ค.
- 3. ์ ํฉ์ฑ ํ๋ณธ ๋๋ฅด๊ธฐ์ ํฉ์ฑ ๋ฐ์ดํฐ๊ฐ ๋ง๋ค์ด์ง๋๋ค. ๊ฐ์ ์กฐ๊ฑด๋ ํ๋ณธ์ ๋ฐ๋ผ ๋ฌ๋ผ์ง ์ ์์ต๋๋ค.
- 4. ๊ทธ๋ฆผ๊ณผ ๊ณ์ฐ ๊ฒฐ๊ณผ ๋น๊ตํ๊ธฐ๋ฐ๊พธ๊ธฐ ์ ํ ๋ฌด์์ด ์์ง์ด๊ณ ๋ฌด์์ด ๊ทธ๋๋ก์ธ์ง ํ ๋ฌธ์ฅ์ผ๋ก ์ ์ด๋ณด์ญ์์ค.
๋งํ๋ฉด ์ด๊ธฐํ๋ก ๋์๊ฐ ๊ธฐ๋ณธ ๊ฒฐ๊ณผ๋ฅผ ๋ณธ ๋ค ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ์ญ์์ค. ์ด Lab์ ์ ๋ต ํ์ ๊ธฐ๊ฐ ์๋๋ผ ํจํด ๊ด์ฐฐ ๋๊ตฌ์ ๋๋ค.
๊ฐ์ ์ค์ ์ ํฉ์ฑ ๊ด์ธก
๊ณ์ฐ ๊ฒฐ๊ณผ
๋จ๊ณ ๊ฒฐ๊ณผ๋ ์ ํํธํฅ ์๋ ์ต์ข ๋ชจํ์ด ์๋๋๋ค. sparsityยทheredity๊ฐ ์ฝํ๊ฑฐ๋ ๊ตํธ์์ฉ์ด ๊ฐํ๋ฉด ์ถ๊ฐ run๊ณผ ๋ ๋ฆฝ ํ์ธ์ด ํ์ํฉ๋๋ค.
๊ต์ก์ฉ synthetic model ยท bjs-advanced-design-sequence-v1. ํ ํ์ ๋ณ๋ ํ์๊ฐ ์๋ ํ ํ๋์ ๋ ๋ฆฝ simulation ๋๋ ์ค๊ณ run์ ๋๋ค. ์ค์ ์ฐ๊ตฌยทํ์งยท๊ท์ ํ๋จ์๋ ์ฌ์ฉํ ์ ์์ต๋๋ค.
The Lab is a fixed 6-factor, 13-run teaching DSD created with a conference matrix. It is a simple stepwise example that only shows the main effect score and centered-square score, and does not reproduce the selection method or inference of JMP Fit Definitive Screening.
In JMP, leave both the design evaluation and analysis steps together
3์์ค ์ขํ์ ์ฃผํจ๊ณผยท2์ฐจํญ์ ์๊ด ๊ตฌ์กฐ๋ฅผ ๋ด ๋๋ค.
์ด๋ค ํ๋ณดํญ๋ผ๋ฆฌ ๊ตฌ๋ถ์ด ์ฝํ์ง ๋จ๊ณ๋ถ์ ์ ์ ํ์ธํฉ๋๋ค.
์ ํ ๋จ๊ณ์ hierarchy ๊ฐ์ ์ด ๊ฒฐ๊ณผ์ ๋ฏธ์น ์ํฅ์ ๊ธฐ๋กํฉ๋๋ค.
In Definitive Screening evaluation, we first examine the three-level table and correlation map. The step-by-step results of Fit Definitive Screening record the threshold, selected main effects, quadratic terms, and interaction terms, as well as the hierarchy. Before calling the automatic selection results the final model, we review the full/alternative model and the possibility of augmentation.
Example of Result Statement
A 13-run DSD was used with 6 continuous factors. In the correlation map before the response, the main effect columns were orthogonal to each other, and the correlations of quadratic terms and 2FI candidates were recorded. As a pre-stage strategy, the main effects of A and D, and Dยฒ, were selected as subsequent candidates, but since this depends on sparsity, heredity, and selection bias, we planned 4 augmentation runs and independent confirmation.
Concluding the Section
- DSD continuous factors have a three-level structure of -1, 0, and +1.
- Main effects and curvature exploration are linked in a small number of runs.
- The correlation map shows which higher-order candidates are not separable.
- Staged analysis depends on sparsity, heredity, and the selection process.
- Promising effects and curvature are candidates until augmentation and confirmation.
In the final section, we package the design, analysis, confirmation, and constraints that have been carried out so far into a verifiable package that others can retrace.
Official Supplementary Materials
- JMP Knowledge Portal ยท Definitive Screening Designs
- JMP Help ยท Fit Definitive Screening Platform Options
- JMP Knowledge Portal ยท Screening Designs
This article and the Lab are educational synthetic materials and should not be used as evidence for actual research, process, quality, or regulatory decisions.