ETC1000 Chap.2 Analysing Numerical Data
Analysing Numerical Data
Three answers to the question of centre
Mean, median and mode agree on symmetric data and diverge exactly where the data is interesting. The mean is the balance point and uses every value, so every value can move it. The median is the middle rank and responds to ordering rather than size.
The mode is the most common value and is the only one of the three that also works on categorical data.
Spread is the half that carries the risk
Centre says where the data sits; spread says how far it wanders. The range is built from the two most unusual observations and is the least stable number available. The interquartile range covers the middle half and ignores both tails.
Variance averages the squared deviations from the mean, and the standard deviation is its square root, which returns the answer to the original units so it can be quoted alongside the mean.
Shape explains the other two
Because the mean chases a long tail and the median does not, the gap between them is a free diagnostic.
A mean above the median means a long upper tail, the usual shape for prices, incomes and waiting times, which are bounded below and unbounded above. A mean below the median means a long lower tail, the shape of variables with a ceiling.
Standardisation makes two groups comparable
A value means nothing until it is read against the spread it came from.
Rewriting an observation as the number of standard deviations it sits from its own group mean strips out the units and the scale, and lets a result from one cohort be ranked against a result from another. The unit names standardisation explicitly among the techniques it teaches.
What this chapter covers
- 01
Mean, median and mode, and which question each of them answers
- 02
Range, interquartile range, variance and standard deviation
- 03
Reading skew from the gap between the mean and the median
- 04
Choosing a histogram interval width and reporting it
- 05
Standardised scores and comparing across groups with different spreads
Rank two results measured on different scales
- 1Standardise each branch's result inside its own history.
- 1Say which week was the more unusual.
- 1Give each manager a different instruction, and justify it.
Key terms
- Mean
- The balance point of the data, computed by dividing the total by the count.
- Median
- The middle value in rank order, insensitive to how extreme the extremes are.
- Range
- The largest value minus the smallest, built entirely from the two least typical cases.
- Interquartile range
- The distance from the first quartile to the third, covering the middle half.
- Variance
- The average squared deviation from the mean, in squared units.
- Standard deviation
- The square root of the variance, in the original units of the variable.
- Right skew
- A shape with a long upper tail, which pulls the mean above the median.
- Standardised score
- An observation expressed as the number of standard deviations from its own group mean.
Analysing Numerical Data FAQ
When should I report the median rather than the mean?
Whenever a few extreme values are pulling the mean away from the bulk of the data, which is the normal situation for prices, incomes, waiting times and order values. The test is to compute both: if they differ noticeably the distribution is skewed and the median is the honest summary of a typical case. The mean is still correct if the question is about a total, because the total is what it is built from.
Why does the sample variance divide by one less than the sample size?
The deviations are measured from the sample mean, and the sample mean is by construction the value that makes those deviations as small as possible. Dividing by the sample size would therefore understate the population variance systematically. Dividing by one fewer corrects it, which in a spreadsheet is the difference between the population and sample variance functions.
What does a standardised score of plus two actually mean?
That the value sits two standard deviations above the mean of the group it came from. It is a statement of position inside that group and nothing else, so it does not translate directly into rarity unless the distribution is roughly symmetric. On strongly skewed data a score of plus two is common in the tail and says much less than the same score on symmetric data.
How do I choose the width of a histogram interval?
By trying two or three and keeping the one where the shape is stable. Intervals that are too wide flatten the distribution into a single block and hide the skew you were looking for; intervals that are too narrow put one or two cases in each and report sampling noise as structure. State the width you used in the caption so a reader can judge.
Exam move
Compute both centres on every dataset you meet, even when only one is asked for, because their difference is a free reading of the shape. Then practise the sentence that names which one you would publish and why, since that judgement is what the assessments reward rather than the arithmetic.
Working through Analysing Numerical Data in ETC1000? Sia is AskSia’s AI Statistics tutor — ask any ETC1000 Analysing Numerical Data question and get a clear, step-by-step explanation grounded in how ETC1000 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.