ETC1000 Chap.4 Models of Relationships Between Data
Models of Relationships Between Data
The scatter is where the tools are licensed
Before any coefficient is computed, the scatter plot settles four things that nothing else shows: the direction of the pattern, its form, whether a straight line is even the right shape, its strength, and whether one point sitting away from the rest is going to pull the line towards itself.
Covariance carries direction, correlation carries size
Covariance averages the product of the two variables' deviations, so it is positive when the deviations usually share a sign.
Its units are the product of the two variables' units, which makes its magnitude uninterpretable and sensitive to rescaling either variable. Dividing by both standard deviations cancels the units and leaves the correlation coefficient, which always lies between minus one and plus one.
One line, chosen by a rule
Least squares picks the line that makes the sum of the squared vertical misses smallest.
Its slope is the change in the outcome per one unit change in the explanation, and it is where nearly all the marks in a regression question sit. The intercept is interpretable only when zero is a value the explanatory variable actually takes in the data.
The leftovers are the diagnostic
Residuals average to zero by construction, so their pattern rather than their average is what to read.
Curvature says a straight line was forced onto a bend; a fan says the outcome is more variable in some conditions than others; one large residual marks a case that may be carrying the whole conclusion.
What this chapter covers
- 01
Reading a scatter for direction, form, strength and exceptions
- 02
Covariance, its sign rule, and why its magnitude cannot be compared
- 03
The correlation coefficient and what a given value looks like
- 04
Fitting the least squares line and interpreting both coefficients
- 05
Residual patterns and what each of the three shapes warns about
- 06
The coefficient of determination and the three things it does not certify
Fit a line and refuse the extrapolation it offers
- 1Compute the slope and the intercept.
- 1Interpret the slope in the units of the problem.
- 1Say what the line may not be used for.
Key terms
- Scatter plot
- A display of paired observations that shows direction, form, strength and exceptions at once.
- Covariance
- The average product of two variables' deviations, carrying direction but not comparable size.
- Correlation coefficient
- Covariance divided by both standard deviations, always between minus one and plus one.
- Least squares line
- The line minimising the sum of the squared vertical distances to the observed points.
- Slope coefficient
- The change in the outcome associated with a one unit change in the explanatory variable.
- Residual
- An observed value minus the value the fitted line predicts for it.
- Coefficient of determination
- The squared correlation, giving the share of outcome variation the line accounts for.
- Extrapolation
- Using a fitted line beyond the range of explanatory values the data contained.
Models of Relationships Between Data FAQ
Does a strong correlation show that one variable causes the other?
No, and nothing in the calculation could. The coefficient is symmetric in the two variables, so it cannot distinguish advertising driving sales from busy stores being given larger budgets, and a third variable such as store size can produce both. The only defence available in this unit is to name the plausible third variable out loud whenever the coefficient is reported.
Can a correlation of zero coexist with an obvious relationship?
Yes. The coefficient measures how tightly the points sit around a straight line, so a perfect arch of points, rising then falling, has a correlation near zero and a completely clear pattern. This is the strongest single argument for drawing the scatter before computing anything.
Which variable goes on which axis?
The one you treat as the explanation goes horizontally and the outcome vertically, and the choice is a modelling decision rather than a presentational one. Swapping them produces a genuinely different fitted line, because least squares measures error vertically. Correlation is unaffected by the swap, which is one of the practical differences between the two tools.
What does a coefficient of determination of 0.97 certify?
That the fitted line accounts for about 97 per cent of the variation in the outcome across the cases in the sample. It is silent on whether a straight line was the right shape, whether the relationship is causal, and whether the model holds outside the range observed. A curved relationship can post a high value while being systematically wrong at both ends.
Exam move
Work one dataset the whole way down, from scatter to covariance to correlation to fitted line to residual, so the numbers can be checked against each other rather than memorised separately. Then practise the two refusals this topic exists to teach: no causation from observed data, and no prediction outside the range the data covered.
Working through Models of Relationships Between Data in ETC1000? Sia is AskSia’s AI Statistics tutor — ask any ETC1000 Models of Relationships Between Data question and get a clear, step-by-step explanation grounded in how ETC1000 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.