Back to List

Express and Diagnose a Relationship: Regression and Residuals

Explains regression slope, intercept, fitted values, least squares, and R², then diagnoses curvature, unequal variance, drift, and outliers with residual plots.

Intermediate
|
30min
|
Verified (2026-08-14)
regressionslopeinterceptleast squaresresidualR-squaredlack of fitJMP
Progress0/28 (0%)

A correlation summarizes how two variables move together, but does not directly say how much average Y changes when X changes by one unit. Regression expresses that question as an equation.

This unit asks one question.


How well can one line explain a relationship, and where does it fail?
inclinationAverage change in Y for one unit change in X
predicted value ŷThe value suggested by the fitting equation for each
residual eDifference of observed value y minus predicted value ŷ
R²What proportion of observed Y variation is explained by the model?

Simple linear regression compresses an average relationship into two coefficients

Yᵢ=β₀+β₁Xᵢ+εᵢ

  • β₀, intercept: mean Y when X=0
  • β₁, slope: change in mean Y for a one-unit increase in X
  • εᵢ, error: individual variation the equation does not explain

In a sample, fit ŷᵢ=b₀+b₁xᵢ. The slope has units of Y units/X units. If X is concentration in mg/mL and Y is viability percentage points, retain those units for scientific meaning.

If X=0 lies outside the observed range, the intercept may be necessary for calculation but have no practical interpretation. Centering X at a scientifically meaningful value makes the intercept the mean response at that observed reference.

Least squares minimizes the sum of squared vertical residuals

A residual is eᵢ=yᵢ−ŷᵢ. Least squares chooses the line with the smallest Σeᵢ². Squaring gives large residuals substantial influence, so always investigate the measurement causes and leverage of outliers.

U15 · Figure 01
Structures missed by straight lines reappear in the residual plot
Observations and fitted lineResidual = y − ŷ
The straight line on the left summarizes the average trend, but the U-shaped pattern in the residuals on the right indicates that the linear equation is missing a curvature term.

A positive residual means the observation is above prediction; a negative one is below. Residuals from a suitable linear model should scatter around zero without special structure against fitted values or run order.

R² is an explained proportion, not a quality certificate

R²=1−SSE/SST

R² is the proportion of total observed Y variation explained by the fitted equation within this sample. In simple regression it equals r². A high R² does not guarantee that:

  • the linear functional form is right;
  • residual variance is constant;
  • errors are independent;
  • important variables are included;
  • the relationship is causal; or
  • prediction on new data is good.

Widening the X range can raise R², and narrowing it can lower R². Where response variation is inherently small, even a low R² can accompany a useful average effect.

A high R² cannot rescue a U-shaped residual pattern

Inspect curvature, funnels, time trends, clusters, and extreme points in residual-versus-fitted plots. One numeric summary can hide structural model failure.

Four residual patterns point to different problems

  1. U or S shape: the linear equation may need curvature or transformation.
  2. Funnel: variance changes with the mean level.
  3. Run-order trend: drift, autocorrelation, or batch change.
  4. Isolated large residual: measurement error, a different population, or an influential point.

Do not select a revision merely by adding increasingly favorable polynomial degrees after viewing residuals. Choose terms using mechanism, design range, and pre-specified candidate models, then check with new data or cross-validation.

Repeated X levels permit separation of pure error from lack of fit. A lack-of-fit test can help assess whether a candidate function has systematic failure beyond repeat error, but non-significance does not make the model true.

Extrapolation is not merely extending the line farther

If observed X ranges from 1–5, predicting at X=20 may encounter different response mechanisms, saturation, toxicity, or physical boundaries. The equation has an algebraic value, but evidence exists inside the design range. A prediction interval is uncertainty conditional on the model being right; it does not justify extrapolation.

In-Silico Lab: make the model fail deliberately

  1. At curvature 0, inspect the fitted line and R².
  2. Raise curvature to 1.2 and 2, then compare R² and residual structure.
  3. Add one outlier and inspect changes in slope and SSE.
  4. Change the seed and see how a sample fit estimates the generating model with variation.
In-Silico Lab · U15

Diagnose straight lines and residuals at the same time

After fitting a straight line, check whether R² and residuals reveal curvature and outliers differently.

If this is your first time: What should I press?
  1. 1. Read the question firstIn the Lab title, check the one thing you will compare this time.
  2. 2. Change just one conditionInitially, change only one of the inputs: n, effect, or spread.
  3. 3. new composite specimen pressureNew synthetic data is created. The same conditions may vary depending on the sample.
  4. 4. Pictures and calculation results CompareWrite in one sentence what moves and what stays the same before and after the change.

If it gets stuckresetGo back to see the default results and change just one condition. This Lab is not a correct answer tester but a pattern observation tool.

Synthetic observations of the same settings

calculation result

correlation r0.679
inclination1.726
R²0.461
SSE276.26

Even high R² does not erase the curved pattern in the residuals. Extrapolation outside the measurement range is specifically prohibited.

educational synthetic modelbjs-relationship-sequence-v1. Actual research judgments must separately reflect experimental units, missingness, distribution, multiplicity, pre-planning, and domain criteria.

JMP interpretation moves from coefficient table to residuals

Parameter Estimates

Intercept/slope estimates and SE/CI are read in actual units.

Summary of Fit

R² is an explanation ratio and not a sole proof of model fit.

Residual Plot / Lack of Fit

Find curvature, heteroscedasticity, time drift, and outlier structures.

Read coefficients, SEs, and CIs in Parameter Estimates; R² and RMSE in Summary of Fit; and the model test in Analysis of Variance, then always inspect residual plots and actual by predicted. Add polynomial terms because of scientific candidates and diagnostics—not because a menu makes them easy.

Example result statement

A linear model was fitted to 38 independent batches over 1–5 mg/mL. The slope was 2.4 percentage points/(mg/mL), 95% CI 1.6–3.2, with R²=.48. Residuals versus fitted values showed no marked curvature, but variance increased at high concentration. The slope direction remained in a weighted-model sensitivity analysis. We did not extrapolate this equation beyond the measured range.

Takeaways

  • The slope is the average Y change for a one-unit X change.
  • Least squares minimizes squared vertical residuals.
  • R² is an explained proportion, not a stand-alone fit criterion.
  • Residuals reveal curvature, unequal variance, drift, and outliers.
  • Choose polynomial terms by mechanism and validation, not post hoc p-value competition.
  • Predictive evidence exists within the observed design range.
A regression equation is a candidate model that compresses the relationship. Residual structure and extrapolation range must be checked before R² to ensure that the predictions will withstand scientific questions.

The next units move beyond observed X–Y relationships to experimental design, where several inputs are deliberately arranged.

Official supplementary resources

The model and Lab in this article are educational synthetic material.

💬 Questions & Comments

0 comments

You can post without signing in. Guest comments cannot be edited or deleted by their author.

0/2000

Loading...