ETC1000 Chap.3 Using Data in an Uncertain World
Using Data in an Uncertain World
Two vocabularies that must not be mixed
From week 3 the data is treated as a sample drawn from a larger population. A population summary is a parameter, fixed and almost never known. A sample summary is a statistic, always known once the data exists, and different for every sample.
That variability is not a defect: a quantity that varies in a describable way can be used to bracket one that does not vary at all.
Bias does not shrink with the sample
The unit places the choice of sampling design on the nature of the population and the resources available, with an unbiased sample as the goal. Sampling variability shrinks as the sample grows, at a knowable rate.
Selection bias, survivor bias and non-response bias do not move at all, so a larger sample from the wrong frame is more precisely wrong. Ask how a case could have entered the dataset before you ask how many cases there are.
The distribution nobody constructs
The sampling distribution of the mean describes how the sample mean behaves across every sample of a given size.
It is centred on the population mean and its spread is the standard error, the population standard deviation divided by the square root of the sample size.
That square root is the economics of sampling: quadrupling the sample halves the standard error, so precision gets expensive quickly.
What a confidence level is a statement about
An interval is the estimate plus and minus a critical value times the standard error.
The confidence level describes the procedure rather than the interval in front of you: building intervals this way captures the parameter in the stated share of repetitions. Your particular interval either contains the parameter or does not, and nothing in the data reveals which.
What this chapter covers
- 01
Parameter against statistic, and why only one of them varies
- 02
Selection, survivor and non-response bias, and why size cannot repair them
- 03
The law of large numbers and the single thing it promises
- 04
The sampling distribution of the mean and the standard error
- 05
Confidence intervals, and stating the level as a property of the procedure
- 06
Why the unit works with the t-distribution rather than the normal
Build an interval and then say what it does not cover
- 1Compute the standard error.
- 1Give an approximate 95 per cent interval for the mean.
- 1Write the reporting sentence with its two qualifications.
Key terms
- Population
- Every case the answer is meant to cover, whether or not any of them were observed.
- Parameter
- A fixed summary of the population, almost never known directly.
- Statistic
- A summary computed from the sample, known exactly and different for every sample.
- Selection bias
- A failure in which the way cases entered the sample is related to what is being measured.
- Survivor bias
- A failure in which the cases that did not last are absent from the frame entirely.
- Sampling distribution
- The distribution a statistic would follow across repeated samples of one size.
- Standard error
- The spread of a statistic across samples, falling with the square root of the sample size.
- Confidence interval
- A range built by a procedure that captures the parameter at a stated long run rate.
Using Data in an Uncertain World FAQ
What does 95 per cent confidence actually mean?
It describes the procedure, not the interval you are holding. Building 95 per cent intervals from repeated samples produces intervals that capture the parameter in about 95 per cent of those repetitions. Your particular interval either contains the parameter or it does not, and no calculation tells you which, which is why the defensible sentence is written about the method rather than about the number.
Why does the unit use the t-distribution instead of the normal?
Because the normal version of the formula assumes the population standard deviation is known, and in practice it almost never is. When it is estimated from the same sample, the extra uncertainty has to be paid for, and the t-distribution does that by having heavier tails and a critical value that depends on the sample size. The penalty is largest on small samples and fades as they grow.
Does a bigger sample fix a biased one?
No, and this is the distinction the topic turns on. Sampling variability shrinks with the sample and bias does not move at all. A survey answered by 900 self selected respondents produces a narrow interval that is a precise estimate of the mean among people who chose to answer, and no amount of extra precision converts it into an estimate for everyone.
Does a confidence interval tell me where individual cases fall?
No. An interval for a mean of 41 to 47 dollars does not say that most customers spend between 41 and 47. Individual values scatter far more widely than their mean does, because the interval was built from a standard error that had already been divided by the square root of the sample size. Confusing the two turns a narrow interval into a false claim about the spread of the population.
Exam move
Write the standard error before anything else in every question in this topic, and check it against the sample size: if it did not fall when the sample grew, the denominator is wrong and everything after it will be too. Then rehearse the three sentences this topic marks: what the interval covers, what the confidence level describes, and which population the frame supports.
Working through Using Data in an Uncertain World in ETC1000? Sia is AskSia’s AI Statistics tutor — ask any ETC1000 Using Data in an Uncertain World question and get a clear, step-by-step explanation grounded in how ETC1000 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.