ETC1000 Chap.5 Hypothesis Testing and Statistical Evidence
Hypothesis Testing and Statistical Evidence
A rule that can be stated before the data is seen
Hypothesis testing assumes for the sake of argument that nothing is going on, works out how surprising the observed result would be under that assumption, and rejects the assumption when the result is surprising enough.
Starting from the position you hope to establish would let any dataset confirm it; starting from no effect means the data has to do work, and how much work is fixed in advance.
Both hypotheses are about the population
The null is written as an equality, because it is the position you can compute from: the mean equals the claimed value, the difference equals zero, the slope equals zero.
The alternative comes from the wording of the question. Differs or changed gives a two sided alternative; exceeds, improved or faster gives a one sided one. A hypothesis written about the sample statistic is a category error, because that value is already known.
The test statistic and the p-value
The test statistic counts how far the result sits from the null in standard errors.
The p-value converts that distance into the probability of seeing something at least this extreme if the null were true. It is the probability of the data given the hypothesis, and is routinely misreported as the probability of the hypothesis given the data.
Two ways to be wrong, and only one lever improves both
A Type I error rejects a true null and a Type II error misses a real effect.
The significance level is exactly the rate at which you accept the first, so tightening it trades one error for the other. Only a larger sample improves both at once, by shrinking the standard error without touching the false alarm budget.
What this chapter covers
- 01
Where a test sits in the four step inference framework the unit teaches
- 02
Writing the null and the alternative about the population, not the sample
- 03
One sided against two sided, and why the question decides it
- 04
The test statistic, the p-value, and the sentence each one licenses
- 05
Type I and Type II errors, and what changing the level actually trades
- 06
Testing a regression slope against a null of zero
Test a regression slope and bound the conclusion
- 1State the hypotheses about the population slope.
- 1Compute the test statistic and decide at the 5 per cent level.
- 1Say what the result licenses and what it does not.
Key terms
- Null hypothesis
- The position of no effect, written as an equality about a population parameter.
- Alternative hypothesis
- What you conclude if the null is rejected, taking its direction from the question.
- Test statistic
- The distance between the observed result and the null value, counted in standard errors.
- P-value
- The probability of a result at least this extreme if the null hypothesis were true.
- Significance level
- The false alarm rate you accept, fixed before the data is examined.
- Type I error
- Rejecting a null hypothesis that was in fact true.
- Type II error
- Failing to reject a null hypothesis that was in fact false.
Hypothesis Testing and Statistical Evidence FAQ
Does a p-value of 0.03 mean the null has a 3 per cent chance of being true?
No. It means that if the null were true, results at least this extreme would occur in about 3 per cent of samples. It is the probability of the data given the hypothesis, and the reversed reading is a different quantity that is not close to it. Writing the correct version takes one extra clause and is worth a mark on most interpretive questions.
What can I conclude from a large p-value?
That the data is consistent with the null, which is not the same as showing the null is true. A large p-value is equally consistent with a real but small effect that this sample was too small to detect. The defensible wording is that there is insufficient evidence of a difference at the stated level, and reporting the confidence interval alongside shows which effect sizes remain plausible.
When is a one sided alternative appropriate?
When the question itself points in a direction, using words such as exceeds, improved or faster. A one sided test spends the whole significance budget in one tail, which makes rejection easier in that direction and impossible in the other. Choosing one sided after seeing which way the data went silently converts a 5 per cent test into a 10 per cent one.
Is a significant result an important one?
Not necessarily. Significance says the result would be surprising under the null, and with a large enough sample a commercially trivial effect produces a very small p-value. Importance is read from the size of the estimate and its confidence interval in the units of the problem, which is why both should be reported together.
Exam move
Practise writing the four part conclusion until it is automatic: the decision, the level, the effect in the units of the problem, and the limit. Then rehearse correcting someone else's conclusion, because that is the form these questions most often take and the three standard errors travel together.
Working through Hypothesis Testing and Statistical Evidence in ETC1000? Sia is AskSia’s AI Statistics tutor — ask any ETC1000 Hypothesis Testing and Statistical Evidence question and get a clear, step-by-step explanation grounded in how ETC1000 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.