Back to List

Two Variables Moving Together: Look at the Scatterplot Before the Correlation Coefficient

Explains why scatterplots should be examined before Pearson r when two variables move together, and how to distinguish nonlinearity, outliers, range restriction, clusters, and confounding.

Intermediate
|
27min
|
Verified (2026-08-14)
correlationcorrelation coefficientscatterplotlinearityconfoundingrange restrictionJMP
Progress0/28 (0%)

As metabolite concentration in culture medium increased, cell viability also increased. If the correlation coefficient between the two variables is r=.78, can we conclude that increasing the metabolite will improve viability? Correlation is a starting point, not a causal conclusion.

This unit has one question.


How much can we say from the fact that two variables move together?
correlation coefficient rDirection and normalized magnitude of linear coexistence
scatterplotActual shape of two continuous variables
Range restrictionsA phenomenon in which the relationship appears weak due to a narrow observation
derangementA variable that is entangled in both X and Y and distorts the relationship

A correlation coefficient compresses linear co-movement into −1 to +1

Pearson's correlation coefficient r is a standardized summary of whether two variables tend to depart from their own means in the same direction.

  • r>0: observations with larger X also tend to have larger Y in a linear pattern
  • r < 0: observations with larger X tend to have smaller Y in a linear pattern
  • |r| close to 1: points lie close to a straight line
  • r≈0: linear co-movement is weak

r=0 does not mean that no relationship exists. In a perfect U-shaped relationship, the slopes on the two sides can cancel and leave r near zero. Draw a scatterplot before calculating one number.

U14 · Figure 01
The shape of a scatter plot has more information than a single r of the same
Observations and fitted lineResidual = y − ŷ
One outlier, observation range and curvature can change r significantly. Look at the points first and read the coefficients.

Correlation removes the units of X and Y, which makes comparisons easier, but it does not state an actual amount of change. “How much does Y change when X increases by one?” is a question about a regression slope and is covered in the next unit.

The same r can conceal very different data

Check the following features in a scatterplot.

  1. Linearity: do points lie around a line, or are there curves, thresholds, or saturation?
  2. Outliers: do one or two points control the direction and size of the result?
  3. Range restriction: was only a narrow part of X observed?
  4. Clusters: are there groups by plate, donor, or date?
  5. Heteroscedasticity: does the spread of Y increase as X increases?
  6. Measurement error: does repeat precision in X and Y attenuate the relationship?

Restricting a sample to healthy adults can weaken an age–biomarker relationship that is strong in the full population. Conversely, combining two batch clusters can make overall r large even when no relationship exists within either cluster.

Check row independence again

If 20 cells from one donor are plotted as 20 independent points, cluster structure can make a p-value far too small. Donor-level summaries, repeated-measures correlation, or a hierarchical model are needed. The number of points is not the same as the amount of independent information.

Correlation has the same value when direction is reversed and does not assign a cause

cor(X,Y)=cor(Y,X). Correlation does not determine which variable is explanatory and which is the response. A high r has several possible explanations.

  • X affects Y.
  • Y affects X.
  • A third variable Z affects both.
  • Selection, measurement, or a time trend created an artificial relationship.
  • Chance and multiple exploration selected a large r.

Causal judgment requires temporal ordering, an intervenable design, randomization, control of confounding, mechanism, and replication. A small correlation p-value does not automatically satisfy those conditions.

p-values and CIs depend on n

A test of H0: population linear correlation ρ=0 uses r and n. With a large n, even a small r can be significant; with a small n, the CI for a large r can be wide. Report r, its CI, the independent n, and the scatterplot together.

If thousands of correlations are tested in a matrix of genes and endpoints, many small p-values arise by chance. Restrict candidates before analysis, or plan multiplicity control such as FDR and independent validation.

In-Silico Lab: create the fragility of r yourself

  1. Change the slope to −2, 0, and +2 to see the direction of r.
  2. Increase curvature to create a situation where the relationship is strong but r is weak.
  3. Add one outlier and see how much r moves.
  4. Check whether the sample r changes under the same generating conditions with a new seed.
In-Silico Lab · U14

Change the scatterplot and see the fragility of r

Change the slope, noise, curvature, and outliers to see if different shapes are hidden behind the same coefficients.

If this is your first time: What should I press?
  1. 1. Read the question firstIn the Lab title, check the one thing you will compare this time.
  2. 2. Change just one conditionInitially, change only one of the inputs: n, effect, or spread.
  3. 3. new composite specimen pressureNew synthetic data is created. The same conditions may vary depending on the sample.
  4. 4. Pictures and calculation results CompareWrite in one sentence what moves and what stays the same before and after the change.

If it gets stuckresetGo back to see the default results and change just one condition. This Lab is not a correct answer tester but a pattern observation tool.

Synthetic observations of the same settings

calculation result

correlation r0.941
inclination1.726
R²0.886
SSE30.40

r is a summary of the linear shape. Causality, reproducibility, and population independence are separate questions.

educational synthetic modelbjs-relationship-sequence-v1. Actual research judgments must separately reflect experimental units, missingness, distribution, multiplicity, pre-planning, and domain criteria.

The line in the Lab is a simple regression line used to support explanation. Seeing a line does not establish causation or justify extrapolation.

In JMP, inspect each scatterplot before the matrix

Scatterplot Matrix

Look at shape, range, clusters, and outliers before the correlation coefficient.

Correlation Matrix

Read the sign and magnitude of r and the n and p values ​​together.

Density Ellipse

Ellipses are linear relationship representations and do not guarantee causal structure.

A Correlation Matrix compresses many pairs but hides nonlinearity and outliers. Inspect the shape of every pair in a Scatterplot Matrix, and read a density ellipse as an auxiliary summary of a linear-normal structure. The slope of an ellipse does not indicate causal direction.

Example result statement

Across 42 independent batches, the Pearson correlation between metabolite X and viability was .61 (95% CI .38–.77). The scatterplot showed a positive linear trend, but two high-concentration points were influential, and coloring by batch date suggested that clustering explained part of the pattern. This observational correlation was not interpreted as a causal effect of increasing X.

At the end of this unit

  • r gives the direction and standardized size of linear co-movement.
  • A nonlinear relationship can exist even when r≈0.
  • Outliers, range restriction, clusters, and confounding can change r substantially.
  • Correlation does not establish a causal direction between X and Y.
  • Consider r, its CI, independent n, and the scatterplot together rather than a p-value alone.
  • Large-scale correlation searches require multiplicity control and validation.
A single correlation coefficient only compresses the shape of the scatterplot. Summarize only the linear motion in r, even after looking at nonlinearity, outliers, ranges, clusters, and disturbances.

In the next unit, we express a relationship as Y=β₀+β₁X+ε and use residuals to find what the straight line missed.

Official supplementary resources

The examples and Lab in this article use synthetic educational data.

💬 Questions & Comments

0 comments

You can post without signing in. Guest comments cannot be edited or deleted by their author.

0/2000

Loading...