As metabolite concentration in culture medium increased, cell viability also increased. If the correlation coefficient between the two variables is r=.78, can we conclude that increasing the metabolite will improve viability? Correlation is a starting point, not a causal conclusion.
This unit has one question.
How much can we say from the fact that two variables move together?
A correlation coefficient compresses linear co-movement into −1 to +1
Pearson's correlation coefficient r is a standardized summary of whether two variables tend to depart from their own means in the same direction.
- r>0: observations with larger X also tend to have larger Y in a linear pattern
r < 0: observations with larger X tend to have smaller Y in a linear pattern- |r| close to 1: points lie close to a straight line
- r≈0: linear co-movement is weak
r=0 does not mean that no relationship exists. In a perfect U-shaped relationship, the slopes on the two sides can cancel and leave r near zero. Draw a scatterplot before calculating one number.
Correlation removes the units of X and Y, which makes comparisons easier, but it does not state an actual amount of change. “How much does Y change when X increases by one?” is a question about a regression slope and is covered in the next unit.
The same r can conceal very different data
Check the following features in a scatterplot.
- Linearity: do points lie around a line, or are there curves, thresholds, or saturation?
- Outliers: do one or two points control the direction and size of the result?
- Range restriction: was only a narrow part of X observed?
- Clusters: are there groups by plate, donor, or date?
- Heteroscedasticity: does the spread of Y increase as X increases?
- Measurement error: does repeat precision in X and Y attenuate the relationship?
Restricting a sample to healthy adults can weaken an age–biomarker relationship that is strong in the full population. Conversely, combining two batch clusters can make overall r large even when no relationship exists within either cluster.
If 20 cells from one donor are plotted as 20 independent points, cluster structure can make a p-value far too small. Donor-level summaries, repeated-measures correlation, or a hierarchical model are needed. The number of points is not the same as the amount of independent information.
Correlation has the same value when direction is reversed and does not assign a cause
cor(X,Y)=cor(Y,X). Correlation does not determine which variable is explanatory and which is the response. A high r has several possible explanations.
- X affects Y.
- Y affects X.
- A third variable Z affects both.
- Selection, measurement, or a time trend created an artificial relationship.
- Chance and multiple exploration selected a large r.
Causal judgment requires temporal ordering, an intervenable design, randomization, control of confounding, mechanism, and replication. A small correlation p-value does not automatically satisfy those conditions.
p-values and CIs depend on n
A test of H0: population linear correlation ρ=0 uses r and n. With a large n, even a small r can be significant; with a small n, the CI for a large r can be wide. Report r, its CI, the independent n, and the scatterplot together.
If thousands of correlations are tested in a matrix of genes and endpoints, many small p-values arise by chance. Restrict candidates before analysis, or plan multiplicity control such as FDR and independent validation.
In-Silico Lab: create the fragility of r yourself
- Change the slope to −2, 0, and +2 to see the direction of r.
- Increase curvature to create a situation where the relationship is strong but r is weak.
- Add one outlier and see how much r moves.
- Check whether the sample r changes under the same generating conditions with a new seed.
Change the scatterplot and see the fragility of r
Change the slope, noise, curvature, and outliers to see if different shapes are hidden behind the same coefficients.
If this is your first time: What should I press?
- 1. Read the question firstIn the Lab title, check the one thing you will compare this time.
- 2. Change just one conditionInitially, change only one of the inputs: n, effect, or spread.
- 3. new composite specimen pressureNew synthetic data is created. The same conditions may vary depending on the sample.
- 4. Pictures and calculation results CompareWrite in one sentence what moves and what stays the same before and after the change.
If it gets stuckresetGo back to see the default results and change just one condition. This Lab is not a correct answer tester but a pattern observation tool.
Synthetic observations of the same settings
calculation result
r is a summary of the linear shape. Causality, reproducibility, and population independence are separate questions.
educational synthetic modelbjs-relationship-sequence-v1. Actual research judgments must separately reflect experimental units, missingness, distribution, multiplicity, pre-planning, and domain criteria.
The line in the Lab is a simple regression line used to support explanation. Seeing a line does not establish causation or justify extrapolation.
In JMP, inspect each scatterplot before the matrix
Look at shape, range, clusters, and outliers before the correlation coefficient.
Read the sign and magnitude of r and the n and p values together.
Ellipses are linear relationship representations and do not guarantee causal structure.
A Correlation Matrix compresses many pairs but hides nonlinearity and outliers. Inspect the shape of every pair in a Scatterplot Matrix, and read a density ellipse as an auxiliary summary of a linear-normal structure. The slope of an ellipse does not indicate causal direction.
Example result statement
Across 42 independent batches, the Pearson correlation between metabolite X and viability was .61 (95% CI .38–.77). The scatterplot showed a positive linear trend, but two high-concentration points were influential, and coloring by batch date suggested that clustering explained part of the pattern. This observational correlation was not interpreted as a causal effect of increasing X.
At the end of this unit
- r gives the direction and standardized size of linear co-movement.
- A nonlinear relationship can exist even when r≈0.
- Outliers, range restriction, clusters, and confounding can change r substantially.
- Correlation does not establish a causal direction between X and Y.
- Consider r, its CI, independent n, and the scatterplot together rather than a p-value alone.
- Large-scale correlation searches require multiplicity control and validation.
In the next unit, we express a relationship as Y=β₀+β₁X+ε and use residuals to find what the straight line missed.
Official supplementary resources
The examples and Lab in this article use synthetic educational data.