DATA2002 Chap.7 Confidence intervals, power and sample size
Confidence intervals, power and sample size
A test can be wrong in two directions and the two are not treated symmetrically. The significance level fixes one of them by decree: when the null hypothesis is true, the chance of rejecting is exactly what you chose. Nothing guarantees the other.
The probability of missing a real effect depends on how large that effect is, how variable the measurements are, how many observations you took and what level you picked, and three of those four are decisions. That asymmetry is why the second half of this chapter is about designing a study before it runs. Before that comes the machinery.
The unit defines its critical value as a lower tail quantity, which catches people who learned the other convention, and from that definition three decision rules follow for the three alternatives.
The same decision can be drawn on three scales, and exam questions arrive on all of them: comparing a probability with the level, comparing a statistic with a critical value, or comparing the sample mean itself with a cut off in the original units. The third of these makes the confidence interval obvious, because a two sided test rejects a hypothesised value exactly when it lies outside the interval centred on the sample mean.
That correspondence is an identity rather than an analogy, and it means the two procedures can never disagree. Coverage is then a property of the procedure across repeated experiments, not a probability about the interval you happen to have. Power is the probability of rejecting when a specified alternative is true, and it is a curve rather than a number, because it depends on how false the null is.
Four levers move it, three of them trade something away, and only sample size improves matters without a concession.
What this chapter covers
- 01
The error table, and why only one of its two rates is guaranteed by the procedure
- 02
Why a very small significance level destroys the ability to detect anything
- 03
The lower tail definition of the critical value, and the three decision rules that follow
- 04
Splitting the level between two tails, and where the half comes from
- 05
One decision drawn on three scales: p-value, statistic, and the measurement itself
- 06
The interval as the set of hypothesised values a two sided test would not reject
- 07
Coverage as a long run property of the procedure, and the two readings it rules out
- 08
One sided intervals, and the finite endpoint read as the largest consistent value
- 09
The significance level chosen before, the p-value computed after
- 10
Power as a function of the true value, passing through the level at the null
- 11
The four levers and their directions, including the one that is often inverted
- 12
Standardising an effect, and sizing a study with the round up rule
Use an interval to answer two tests, then correct a misreading
- +1A test of 19 rejects, because 19 lies outside the interval.
- +1A test of 13 does not reject, because 13 lies inside it.
- +1.5No calculation is needed because the interval is the set of hypothesised values that a two sided test at that level would not reject. The two procedures use the same critical value and the same standard error, so they cannot disagree.
- +1.5The quoted sentence treats the parameter as random. The parameter is a fixed unknown number and this interval is already computed, so it either contains the parameter or it does not. The 95 per cent describes the procedure: intervals built this way capture the parameter in about 95 per cent of repeated studies.
Key terms
- Type I error
- Rejecting the null hypothesis when it is true. Its probability is the significance level, which is chosen in advance and guaranteed by the procedure.
- Type II error
- Failing to reject the null hypothesis when it is false. Its probability is not fixed by the procedure and depends on the effect size, the spread, the sample size and the level.
- Power
- The probability of rejecting the null hypothesis when a specified alternative is true. It is one minus the Type II error rate and is a function of how false the null is.
- Critical value
- The cut off the test statistic is compared with. The unit defines it as a lower tail quantity, so at a small level it is a negative number.
- Rejection region
- The set of values of the statistic, or equivalently of the sample mean, for which the null hypothesis is rejected. Drawing it on the measurement scale is what makes the interval visible.
- Coverage probability
- The long run proportion of intervals built by a given procedure that contain the parameter. It is a property of the procedure across repeated studies, not of one computed interval.
- Observed significance level
- Another name for the p-value: the level at which this particular sample would sit exactly on the boundary between rejecting and not rejecting.
- Effect size
- A difference expressed relative to the spread of the measurements, which makes it comparable across studies and units. Conventional benchmarks are small, medium and large.
- Minimum detectable effect
- The smallest effect a study of a given size can detect with a stated power. Fixing it is a judgement about what matters practically rather than a statistical calculation.
- Retrospective power
- Power computed after a study using the effect that study observed. It adds nothing, since a test that failed to reject will always report low power against its own observed effect.
Confidence intervals, power and sample size FAQ
Why does making the significance level smaller reduce power?
Because a smaller level moves the critical value further from zero, and the test then rejects less often whatever the truth is. Fewer false alarms and fewer genuine detections come together, so the Type II error rate rises as the Type I rate falls. The intuition that a stricter test is a better test is misleading here: strictness is precisely what a Type II error is made of.
The only lever that reduces both error rates at once is collecting more data.
What is wrong with saying there is a 95 per cent chance the parameter is in my interval?
It puts the randomness in the wrong place. The parameter is a fixed unknown constant and your interval is already computed, so the statement it makes is either certainly true or certainly false and there is no probability left in it. The randomness belongs to the endpoints, which would have come out differently in another sample.
The correct statement is about the procedure: if the study were repeated many times, about 95 per cent of the intervals built this way would contain the parameter.
How do I know whether to use a one sided or a two sided alternative?
From the question rather than from the data. A two sided alternative is the default and is appropriate whenever a departure in either direction would matter, which is most of the time. A one sided alternative needs a reason that exists before the data are seen, such as a physical impossibility on one side or a decision that only one direction is actionable.
Choosing the direction after seeing which way the sample went is a way of doubling your effective significance level without saying so.
Why do sample size answers always round up?
Because rounding down leaves the study with less power than was required, and the requirement was the point of the calculation. The rule holds whatever the fractional part: a calculation returning 42.1 gives an answer of 43, not 42. It is worth stating the reason alongside the number in an answer, because the mark is usually for the reason rather than for the rounding itself.
Are power calculations reliable?
They are estimates resting on a guess, and the unit says so directly. Every power figure needs a value for the population spread, and the only source for that is usually a pilot sample or a published study. If the real spread turns out larger than assumed, the study will have less power than promised.
That does not make the calculation useless: it makes it a plan whose assumption should be stated and, where possible, padded, rather than a guarantee.
Exam move
Memorise the four levers as directions rather than as a formula, because that is how they are examined: larger true effect raises power, larger sample raises power, larger spread lowers it, and smaller significance level lowers it. Write them on one line and check yourself by asking what each one costs.
Second, practise reading a confidence interval as a set of hypotheses, since a question asking what a test of a given value would conclude is answerable in five seconds once you see that the interval and the test are the same decision.
Third, learn the three scales as three drawings of one picture, and when a question arrives decide which scale it is speaking on before answering, because a question phrased on the measurement scale is often mistaken for one about the statistic.
Fourth, do one full sample size calculation by hand and write down the caveat with it, since the marks in that question type sit in identifying which inputs are decisions and which is an estimate, and in the round up rule with its reason. Finally, rehearse the two forbidden interpretations of an interval until you can state the correct one without hesitating, because it appears on nearly every paper in some form.
Working through Confidence intervals, power and sample size in DATA2002? Sia is AskSia’s AI Statistics tutor — ask any DATA2002 Confidence intervals, power and sample size question and get a clear, step-by-step explanation grounded in how DATA2002 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.