The University of Sydney · FACULTY OF STATISTICS

DATA2002 Chap.3 Hypothesis testing and chi-squared goodness of fit

- one subject, every graph, every model, every mark
12 Chapters8-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 3 of 16 · DATA2002

Hypothesis testing and chi-squared goodness of fit

This is where the unit stops describing data and starts making statements about the population behind it. The machinery is a single idea repeated: propose a model, work out what the data should look like if it holds, measure how far the data are from that, and ask how often a distance that large would arise by chance when the model is true. Everything later is a variation on which distance to measure.

The unit reports every test through the same six slots, and they are worth treating as an answer template rather than a checklist: significance level, hypotheses, assumptions, test statistic with its null distribution, p-value, and a decision written in the words of the problem. Two consequences follow.

The first three slots depend only on the question and the design, so under time pressure they can be written before any arithmetic. And a slot you cannot fill is diagnostic, because if you cannot state the assumptions you have not yet decided which test this is.

The first statistic the unit builds is the chi-squared goodness of fit statistic, and it is worth learning as three design choices rather than as a formula: square the gaps between observed and expected counts, because they sum to zero by construction; divide each by its expected count, because a gap of ten matters more where five were expected than where five hundred were; and add them, which makes the test one sided in the statistic even though the alternative is two sided in meaning.

Two further facts carry most of the examinable weight. The degrees of freedom are the number of categories minus one for the constraint that counts sum to the sample size, minus one more for every parameter estimated from the same data. And the chi-squared distribution is a large sample limit, which is why every expected count must be at least five and why adjacent categories are pooled when one is not.

In this chapter

What this chapter covers

  • 01

    The six reporting slots, and why the first three can be written before any number

  • 02

    A hypothesis is a statement about population parameters, never about the sample

  • 03

    Two sentences the unit repeats: a large p-value is not evidence of the null, and the level is chosen in advance

  • 04

    Where the statistic comes from: square, divide by the expected count, add

  • 05

    Why only the upper tail is used, and why there is no lower tail p-value here

  • 06

    Degrees of freedom as two separate subtractions with two separate reasons

  • 07

    Estimating a parameter costs one degree of freedom; being told it costs nothing

  • 08

    The by hand calculation table, and the two columns that check themselves

  • 09

    The expected count condition, and pooling adjacent categories before counting again

  • 10

    Goodness of fit to a named count distribution, with the rate estimated by the sample mean

  • 11

    Reading the printed result, and the degrees of freedom it does not know to subtract

  • 12

    The direction of the correction: a smaller degrees of freedom gives a smaller p-value

Worked example · free

Repair a printed goodness of fit result

Q [5 marks]. A printed goodness of fit result reads: statistic 11.34, degrees of freedom 5, p-value 0.045. The analyst estimated one parameter from the data in order to compute the expected counts. State how many categories there were, whether the printed p-value is usable, and what the corrected test concludes at the 5 per cent level. (5 marks. The mark allocation is ours, not the University's.)
  • +1The goodness of fit printout always uses the number of categories minus one for its degrees of freedom, so five degrees of freedom means six categories.
  • +2The printed p-value is not usable, because it was computed on five degrees of freedom when the correct number is categories minus one minus estimated parameters, which is six minus one minus one, giving four.
  • +1Recomputing the upper tail of a chi-squared distribution on four degrees of freedom at 11.34 gives a smaller p-value than 0.045, because removing a degree of freedom pushes the same statistic further into the tail.
  • +1The test already rejected at the 5 per cent level on the printed figure, so the corrected test rejects as well and the conclusion is unchanged. The mark is for showing that the figure needed repair and in which direction it moves.
Six categories; the printed p-value is computed on the wrong degrees of freedom and must be recomputed on four; the corrected p-value is smaller, and the decision to reject is unchanged.
Sia tip — The asymmetry is what makes this examinable. The printout is conservative here, so it errs in the direction that hides a finding: a borderline test can cross the threshold once corrected, and a test that already rejects will still reject.
Glossary

Key terms

Null hypothesis
A statement about population parameters, usually of no difference or no departure from a specified model, whose truth is assumed while the p-value is computed. Writing it in terms of observed counts is a definition error.
Significance level
A threshold for the p-value chosen before the data are seen, conventionally 0.05. It fixes the rate of false alarms the procedure will tolerate and is not a discovery about the data.
Expected count
The count a category would contain on average under the null model, computed as the sample size times the hypothesised probability. It is the quantity the validity condition is checked on.
Goodness of fit
A test of whether one categorical variable's observed counts are consistent with a stated set of category probabilities or with a named distribution.
Degrees of freedom
The number of independent pieces of information left after the constraints. For this test it is the number of categories minus one, minus one more for each parameter estimated from the sample.
Pooling
Combining adjacent categories so that every expected count reaches the required minimum. It changes the number of categories and therefore the degrees of freedom, so it must happen before they are counted.
Upper tail
The region of the null distribution used for the p-value here. A small statistic means the data agree with the model, which is not evidence against it, so no lower tail is used.
Poisson distribution
A model for counts of events in fixed intervals whose mean and variance are equal. Its single parameter is estimated by the sample mean, which costs one degree of freedom.
Observed statistic
The value the test statistic takes on this particular sample, written with a lower case letter to distinguish it from the random variable whose distribution is being used.
Contribution
One term of the goodness of fit sum, belonging to one category. The largest contribution names where the model failed hardest, and it is a description rather than a separate test.
FAQ

Hypothesis testing and chi-squared goodness of fit FAQ

Why is the p-value always the upper tail here?

Because the statistic measures distance from the model in any direction and turns all of it into a positive number. A large value means the observed counts sit far from what the model predicted, which is evidence against it. A small value means they sit close, which is agreement rather than evidence, so there is nothing for a lower tail to detect.

That is different from the t tests later in the unit, where the statistic is signed and the direction of the alternative decides which tail or tails are used.

How do I know whether a parameter counts as estimated?

Ask whether you had to look at the data to find it. If the model specifies every category probability outright, as a stated theoretical split does, then nothing was estimated. If the model is a named distribution whose rate or mean you computed from the same sample, then one parameter was estimated and the degrees of freedom fall by one.

The question is not how many parameters the distribution has in general, but how many of them this analysis took from these data.

What do I do if an expected count is below five?

Pool adjacent categories until every expected count clears the threshold, then recount the categories, because pooling changes the degrees of freedom. This is straightforward when the categories have a natural neighbour, which they usually do for an ordered variable.

If pooling is unavailable or would destroy the comparison the question asks about, move to an exact or simulated method rather than reporting an approximation you have just shown to be unreliable.

Is a non significant result evidence that the model is correct?

No, and the unit is explicit about it. Failing to reject means the data provide insufficient evidence against the model at the chosen level; it does not mean the model is true, and with a small sample a badly wrong model can easily survive. The correct wording is that there is no evidence against the stated model in these data.

A conclusion written as the model is correct claims something the test cannot deliver and reliably costs the final mark.

What can this be asked about in a closed book paper with no calculator?

Everything except a tail probability. You can be asked which degrees of freedom apply, which is arithmetic on small integers; to assemble the expression for the observed statistic without evaluating it; to identify which category contributed most and in which direction; to say what must be done when an expected count is small; and to write the decision given a stated p-value.

Practise writing the expression rather than finishing the arithmetic.

Study strategy

Exam move

Learn the six slots as a template and write them out for every worked example you meet, including the ones where you already know the answer, because in the exam the first three slots are free marks that depend on reading rather than computing.

Then drill the degrees of freedom until they are automatic, and drill them as two separate subtractions with two separate reasons rather than as one formula, since the second subtraction is what the printout does not know about. Build a small table of your own with three columns, the number of categories, the number of parameters estimated, and the resulting degrees of freedom, and fill it in for six invented scenarios.

Practise the by hand calculation table once with real arithmetic so that the two self checks become instinct: the gaps sum to zero and the expected counts sum to the sample size. After that, work with printed outputs rather than with data, because the paper hands you results far more often than it hands you numbers to add up.

For each printout, ask the same four questions: which test does the heading name, is the statistic plausible, were any parameters estimated, and does the p-value need recomputing.

Working through Hypothesis testing and chi-squared goodness of fit in DATA2002? Sia is AskSia’s AI Statistics tutor — ask any DATA2002 Hypothesis testing and chi-squared goodness of fit question and get a clear, step-by-step explanation grounded in how DATA2002 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 64 of your The University of Sydney subjects - and 1,000+ Bibles across every Australian university.
Sia - your DATA2002 tutor, unlimited, worked the way the exam marks it
The full 8-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works
Unlock the full DATA2002 Bible + 64 The University of Sydney subjects
$0.99 Trial