Monash University · FACULTY OF STATISTICS

ETC1000 Chap.4 Models of Relationships Between Data

- one subject, every graph, every model, every mark
6 Chapters6-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 4 of 11 · ETC1000

Models of Relationships Between Data

The scatter is where the tools are licensed

Before any coefficient is computed, the scatter plot settles four things that nothing else shows: the direction of the pattern, its form, whether a straight line is even the right shape, its strength, and whether one point sitting away from the rest is going to pull the line towards itself.

Covariance carries direction, correlation carries size

Covariance averages the product of the two variables' deviations, so it is positive when the deviations usually share a sign.

Its units are the product of the two variables' units, which makes its magnitude uninterpretable and sensitive to rescaling either variable. Dividing by both standard deviations cancels the units and leaves the correlation coefficient, which always lies between minus one and plus one.

One line, chosen by a rule

Least squares picks the line that makes the sum of the squared vertical misses smallest.

Its slope is the change in the outcome per one unit change in the explanation, and it is where nearly all the marks in a regression question sit. The intercept is interpretable only when zero is a value the explanatory variable actually takes in the data.

The leftovers are the diagnostic

Residuals average to zero by construction, so their pattern rather than their average is what to read.

Curvature says a straight line was forced onto a bend; a fan says the outcome is more variable in some conditions than others; one large residual marks a case that may be carrying the whole conclusion.

In this chapter

What this chapter covers

  • 01

    Reading a scatter for direction, form, strength and exceptions

  • 02

    Covariance, its sign rule, and why its magnitude cannot be compared

  • 03

    The correlation coefficient and what a given value looks like

  • 04

    Fitting the least squares line and interpreting both coefficients

  • 05

    Residual patterns and what each of the three shapes warns about

  • 06

    The coefficient of determination and the three things it does not certify

Worked example · free

Fit a line and refuse the extrapolation it offers

Q [3 marks]. AskSia assigns three practice points to this independent exercise; they are not a University marking scheme. Seven authored stores spend 1 to 7 thousand dollars a week on advertising and record weekly sales of 3.2, 4.1, 4.4, 6.0, 6.3, 7.9 and 8.1 thousand dollars.
  • 1Compute the slope and the intercept.
  • 1Interpret the slope in the units of the problem.
  • 1Say what the line may not be used for.
The spends average 4 and the sales average about 5.71 thousand dollars. The cross products of the deviations sum to 24.2 and the squared spend deviations to 28, so the slope is about 0.864 and the intercept is about 2.26. Each extra thousand dollars of advertising is associated with about 864 dollars more weekly sales across these seven stores, which on this evidence does not quite pay for itself. The line may not be used to predict sales at a spend of 30 thousand dollars, because the data covers 1 to 7 thousand and nothing in it says the relationship stays straight beyond that; diminishing returns are exactly the kind of bend a seven point sample inside a narrow range could not detect.
Sia tip — State the slope as a change per one unit with both units named. A slope reported as a bare number has not answered the question.
Glossary

Key terms

Scatter plot
A display of paired observations that shows direction, form, strength and exceptions at once.
Covariance
The average product of two variables' deviations, carrying direction but not comparable size.
Correlation coefficient
Covariance divided by both standard deviations, always between minus one and plus one.
Least squares line
The line minimising the sum of the squared vertical distances to the observed points.
Slope coefficient
The change in the outcome associated with a one unit change in the explanatory variable.
Residual
An observed value minus the value the fitted line predicts for it.
Coefficient of determination
The squared correlation, giving the share of outcome variation the line accounts for.
Extrapolation
Using a fitted line beyond the range of explanatory values the data contained.
FAQ

Models of Relationships Between Data FAQ

Does a strong correlation show that one variable causes the other?

No, and nothing in the calculation could. The coefficient is symmetric in the two variables, so it cannot distinguish advertising driving sales from busy stores being given larger budgets, and a third variable such as store size can produce both. The only defence available in this unit is to name the plausible third variable out loud whenever the coefficient is reported.

Can a correlation of zero coexist with an obvious relationship?

Yes. The coefficient measures how tightly the points sit around a straight line, so a perfect arch of points, rising then falling, has a correlation near zero and a completely clear pattern. This is the strongest single argument for drawing the scatter before computing anything.

Which variable goes on which axis?

The one you treat as the explanation goes horizontally and the outcome vertically, and the choice is a modelling decision rather than a presentational one. Swapping them produces a genuinely different fitted line, because least squares measures error vertically. Correlation is unaffected by the swap, which is one of the practical differences between the two tools.

What does a coefficient of determination of 0.97 certify?

That the fitted line accounts for about 97 per cent of the variation in the outcome across the cases in the sample. It is silent on whether a straight line was the right shape, whether the relationship is causal, and whether the model holds outside the range observed. A curved relationship can post a high value while being systematically wrong at both ends.

Study strategy

Exam move

Work one dataset the whole way down, from scatter to covariance to correlation to fitted line to residual, so the numbers can be checked against each other rather than memorised separately. Then practise the two refusals this topic exists to teach: no causation from observed data, and no prediction outside the range the data covered.

Working through Models of Relationships Between Data in ETC1000? Sia is AskSia’s AI Statistics tutor — ask any ETC1000 Models of Relationships Between Data question and get a clear, step-by-step explanation grounded in how ETC1000 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 92 of your Monash University subjects - and 1,000+ Bibles across every Australian university.
Sia - your ETC1000 tutor, unlimited, worked the way the exam marks it
The full 6-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works
ETC1000 · Business and Economic Statistics - independent study guide on the AskSia Library. More Monash University subjects · Microeconomics across all universities
Unlock the full ETC1000 Bible + 92 Monash University subjects
$0.99 Trial