As metabolite concentration in culture medium increased, cell viability also increased. If the correlation coefficient between the two variables is r=.78, can we conclude that increasing the metabolite will improve viability? Correlation is a starting point, not a causal conclusion.
This unit has one question.
How much can we say from the fact that two variables move together?
A correlation coefficient compresses linear co-movement into โ1 to +1
Pearson's correlation coefficient r is a standardized summary of whether two variables tend to depart from their own means in the same direction.
- r>0: observations with larger X also tend to have larger Y in a linear pattern
r < 0: observations with larger X tend to have smaller Y in a linear pattern- |r| close to 1: points lie close to a straight line
- rโ0: linear co-movement is weak
r=0 does not mean that no relationship exists. In a perfect U-shaped relationship, the slopes on the two sides can cancel and leave r near zero. Draw a scatterplot before calculating one number.
Correlation removes the units of X and Y, which makes comparisons easier, but it does not state an actual amount of change. โHow much does Y change when X increases by one?โ is a question about a regression slope and is covered in the next unit.
The same r can conceal very different data
Check the following features in a scatterplot.
- Linearity: do points lie around a line, or are there curves, thresholds, or saturation?
- Outliers: do one or two points control the direction and size of the result?
- Range restriction: was only a narrow part of X observed?
- Clusters: are there groups by plate, donor, or date?
- Heteroscedasticity: does the spread of Y increase as X increases?
- Measurement error: does repeat precision in X and Y attenuate the relationship?
Restricting a sample to healthy adults can weaken an ageโbiomarker relationship that is strong in the full population. Conversely, combining two batch clusters can make overall r large even when no relationship exists within either cluster.
If 20 cells from one donor are plotted as 20 independent points, cluster structure can make a p-value far too small. Donor-level summaries, repeated-measures correlation, or a hierarchical model are needed. The number of points is not the same as the amount of independent information.
Correlation has the same value when direction is reversed and does not assign a cause
cor(X,Y)=cor(Y,X). Correlation does not determine which variable is explanatory and which is the response. A high r has several possible explanations.
- X affects Y.
- Y affects X.
- A third variable Z affects both.
- Selection, measurement, or a time trend created an artificial relationship.
- Chance and multiple exploration selected a large r.
Causal judgment requires temporal ordering, an intervenable design, randomization, control of confounding, mechanism, and replication. A small correlation p-value does not automatically satisfy those conditions.
p-values and CIs depend on n
A test of H0: population linear correlation ฯ=0 uses r and n. With a large n, even a small r can be significant; with a small n, the CI for a large r can be wide. Report r, its CI, the independent n, and the scatterplot together.
If thousands of correlations are tested in a matrix of genes and endpoints, many small p-values arise by chance. Restrict candidates before analysis, or plan multiplicity control such as FDR and independent validation.
In-Silico Lab: create the fragility of r yourself
- Change the slope to โ2, 0, and +2 to see the direction of r.
- Increase curvature to create a situation where the relationship is strong but r is weak.
- Add one outlier and see how much r moves.
- Check whether the sample r changes under the same generating conditions with a new seed.
์ฐ์ ๋๋ฅผ ๋ฐ๊พธ๊ณ r์ ์ทจ์ฝ์ฑ์ ๋ณด์ธ์
๊ธฐ์ธ๊ธฐยทnoiseยท๊ณก๋ฅ ยท์ด์๊ฐ์ ๋ฐ๊พธ๋ฉฐ ๊ฐ์ ๊ณ์ ๋ค์ ๋ค๋ฅธ ๋ชจ์์ด ์จ๋์ง ํ์ธํฉ๋๋ค.
์ฒ์์ด๋ผ๋ฉด: ๋ฌด์์ ๋๋ฌ์ผ ํ๋์?
- 1. ์ง๋ฌธ์ ๋จผ์ ์ฝ๊ธฐLab ์ ๋ชฉ์์ ์ด๋ฒ์ ๋น๊ตํ ํ ๊ฐ์ง๋ฅผ ํ์ธํฉ๋๋ค.
- 2. ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ๊ธฐ์ฒ์์๋ n, ํจ๊ณผ, ์ฐํฌ ๊ฐ์ ์ ๋ ฅ ์ค ํ๋๋ง ๋ฐ๊พธ์ญ์์ค.
- 3. ์ ํฉ์ฑ ํ๋ณธ ๋๋ฅด๊ธฐ์ ํฉ์ฑ ๋ฐ์ดํฐ๊ฐ ๋ง๋ค์ด์ง๋๋ค. ๊ฐ์ ์กฐ๊ฑด๋ ํ๋ณธ์ ๋ฐ๋ผ ๋ฌ๋ผ์ง ์ ์์ต๋๋ค.
- 4. ๊ทธ๋ฆผ๊ณผ ๊ณ์ฐ ๊ฒฐ๊ณผ ๋น๊ตํ๊ธฐ๋ฐ๊พธ๊ธฐ ์ ํ ๋ฌด์์ด ์์ง์ด๊ณ ๋ฌด์์ด ๊ทธ๋๋ก์ธ์ง ํ ๋ฌธ์ฅ์ผ๋ก ์ ์ด๋ณด์ญ์์ค.
๋งํ๋ฉด ์ด๊ธฐํ๋ก ๋์๊ฐ ๊ธฐ๋ณธ ๊ฒฐ๊ณผ๋ฅผ ๋ณธ ๋ค ์กฐ๊ฑด ํ๋๋ง ๋ฐ๊พธ์ญ์์ค. ์ด Lab์ ์ ๋ต ํ์ ๊ธฐ๊ฐ ์๋๋ผ ํจํด ๊ด์ฐฐ ๋๊ตฌ์ ๋๋ค.
๊ฐ์ ์ค์ ์ ํฉ์ฑ ๊ด์ธก
๊ณ์ฐ ๊ฒฐ๊ณผ
r์ ์ ํ ๋ชจ์์ ์์ฝ์ ๋๋ค. ์ธ๊ณผยท์ฌํ์ฑยท๊ตฐ์ง ๋ ๋ฆฝ์ฑ์ ๋ณ๋ ์ง๋ฌธ์ ๋๋ค.
๊ต์ก์ฉ synthetic model ยท bjs-relationship-sequence-v1. ์ค์ ์ฐ๊ตฌ ํ๋จ์๋ ์คํ๋จ์, ๊ฒฐ์ธก, ๋ถํฌ, ๋ค์ค์ฑ, ์ฌ์ ๊ณํ๊ณผ ๋๋ฉ์ธ ๊ธฐ์ค์ ๋ณ๋๋ก ๋ฐ์ํด์ผ ํฉ๋๋ค.
The line in the Lab is a simple regression line used to support explanation. Seeing a line does not establish causation or justify extrapolation.
In JMP, inspect each scatterplot before the matrix
์๊ด๊ณ์๋ณด๋ค ๋จผ์ ๋ชจ์ยท๋ฒ์ยท๊ตฐ์งยท์ด์๊ฐ์ ๋ด ๋๋ค.
r์ ๋ถํธ์ ํฌ๊ธฐ, n๊ณผ p๊ฐ์ ํจ๊ป ์ฝ์ต๋๋ค.
ํ์์ ์ ํ ๊ด๊ณ ํํ์ด๋ฉฐ ์ธ๊ณผ ๊ตฌ์กฐ๋ฅผ ๋ณด์ฆํ์ง ์์ต๋๋ค.
A Correlation Matrix compresses many pairs but hides nonlinearity and outliers. Inspect the shape of every pair in a Scatterplot Matrix, and read a density ellipse as an auxiliary summary of a linear-normal structure. The slope of an ellipse does not indicate causal direction.
Example result statement
Across 42 independent batches, the Pearson correlation between metabolite X and viability was .61 (95% CI .38โ.77). The scatterplot showed a positive linear trend, but two high-concentration points were influential, and coloring by batch date suggested that clustering explained part of the pattern. This observational correlation was not interpreted as a causal effect of increasing X.
At the end of this unit
- r gives the direction and standardized size of linear co-movement.
- A nonlinear relationship can exist even when rโ0.
- Outliers, range restriction, clusters, and confounding can change r substantially.
- Correlation does not establish a causal direction between X and Y.
- Consider r, its CI, independent n, and the scatterplot together rather than a p-value alone.
- Large-scale correlation searches require multiplicity control and validation.
In the next unit, we express a relationship as Y=ฮฒโ+ฮฒโX+ฮต and use residuals to find what the straight line missed.
Official supplementary resources
The examples and Lab in this article use synthetic educational data.