The University of Sydney · FACULTY OF STATISTICS

DATA2002 Chap.12 Two-way and two-factor analysis of variance

- one subject, every graph, every model, every mark
12 Chapters6-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 12 of 16 · DATA2002

Two-way and two-factor analysis of variance

Real experiments rarely vary one thing at a time. Material arrives in batches, subjects differ from one another, machines drift between days. This chapter handles a second grouping variable in two different roles, and keeping the roles apart matters because they lead to different tables and different conclusions.

In the first role the second variable is a nuisance you can record but do not care about, and the design gives one observation per treatment in each block. Recording it lets the analysis subtract it, which is worth a great deal: the same four treatments measured across five batches can be significant when the batch effect is removed and unremarkable when it is left in the residual, from identical data.

What blocking cannot do is detect an interaction, and the reason is structural rather than computational, since one observation per cell leaves nothing to estimate within cell variation from. In the second role the extra variable is a factor of interest and every combination is observed more than once.

That replication creates a genuine within cell residual, and that in turn allows a fourth source of variation to be separated out. The interaction for a cell is what is left of its mean once the overall mean and both main effects have been accounted for, and it is zero for every cell exactly when the two factors act independently, which is what parallel lines on an interaction plot mean.

When the interaction is significant the main effect rows become misleading, because each is an average over conditions in which the factor behaves differently, so the reportable finding is the effect of one factor stated separately at each level of the other.

In this chapter

What this chapter covers

  • 01

    Two roles for a second grouping variable, and the different table each produces

  • 02

    Blocking as the removal of a nuisance source, and the signature of a nuisance factor

  • 03

    What blocking buys, shown by analysing the same data with and without it

  • 04

    Why a block design cannot detect an interaction, as a property of the design

  • 05

    The three row table, and the degrees of freedom that must add

  • 06

    The block row as a housekeeping note rather than a finding

  • 07

    Replication, and the fourth row it creates

  • 08

    The interaction defined by subtraction from the two margins

  • 09

    Reading an interaction plot: parallel lines against a gap that changes

  • 10

    Why a main effect becomes misleading once an interaction is present

  • 11

    Multiple comparison multipliers carried over, and how the gap widens with more levels

  • 12

    Two rank based procedures, one ranking globally and one ranking within blocks

Worked example · free

Count degrees of freedom and say what replication is for

Q [5 marks]. An experiment crosses a factor with 3 levels against a factor with 4 levels, with 2 observations in every combination. Give the degrees of freedom for both main effects, the interaction, the residual and the total, and state what would be lost if only one observation had been taken per combination. (5 marks. The mark allocation is ours, not the University's.)
  • +1There are three times four times two, that is twenty four observations. The first factor contributes two degrees of freedom, the second contributes three.
  • +1.5The interaction contributes the product of those two, which is six. The residual is the observations minus the number of cells, twenty four minus twelve, which is twelve.
  • +1The total is twenty three, and the check is that two plus three plus six plus twelve equals twenty three.
  • +1.5With one observation per combination there would be twelve observations, eleven total degrees of freedom, and the model would consume all eleven, leaving zero for the residual. With no estimate of the error variance, no F ratio could be formed at all.
Two, three, six, twelve and twenty three. Without replication the residual degrees of freedom vanish and no test is possible, which is why a block design assumes the interaction away and uses its degrees of freedom as the residual.
Sia tip — Replication means repetition within a cell, not more cells. Twenty four observations spread one per cell across twenty four combinations give no estimate of within cell variation whatever, and adding more cells cannot rescue that.
Glossary

Key terms

Block
A grouping variable recorded because it is a known nuisance source of variation, rather than because its effect is of interest. Removing it shrinks the residual term and sharpens the comparison that is of interest.
Randomised block design
A design in which every treatment appears once within each block. It has no replication inside a cell, so it cannot separate an interaction from the residual.
Main effect
The effect of one factor averaged over the levels of the other. It describes something real only when the two factors act independently.
Interaction
The part of a cell mean that the two margins do not predict. It is zero for every cell when the factors act independently, and its presence makes each main effect an average over dissimilar conditions.
Interaction plot
Cell means plotted against one factor with one line per level of the other. Parallel lines indicate no interaction; lines that converge, diverge or cross indicate one.
Replication
More than one observation in each combination of the two factors. It is what creates the within cell residual and therefore what makes an interaction estimable.
Sum to zero constraint
A restriction on the effect parameters that makes them identifiable, since otherwise a constant could be moved between the overall mean and the effects without changing any fitted value.
Fitted value
The value the model predicts for a cell, formed from the overall mean and the estimated effects. Residuals are computed against it and are the basis of the diagnostic plots.
Nuisance factor
A variable that contributes real variation to the response but is not the subject of the experiment. Recording it converts that variation from residual noise into an explained term.
Within block ranking
Replacing each observation by its rank among the observations in its own block, which is what the two way rank based procedure does and what distinguishes it from the one way version.
FAQ

Two-way and two-factor analysis of variance FAQ

What does blocking actually buy me?

It moves a real source of variation out of the residual term and into a row of its own. Since every F ratio has the residual mean square in its denominator, shrinking that denominator makes every comparison sharper. The effect can be decisive: the same treatments can be strongly significant when the block effect is removed and unremarkable when it is not, from identical data.

That is why recording a known nuisance variable at the design stage is worth the trouble even though its own row is rarely of interest.

Why can a block design not test for an interaction?

Because with one observation per treatment per block there is exactly one number in each cell, so nothing is left over to estimate variation within a cell. The interaction and the residual would be the same quantity and the model cannot separate them.

It is a property of the design rather than a limitation of the software, and the only fix is replication, which means more than one observation in each combination rather than more combinations.

In what order should I read a two way table?

The interaction row first, then the main effects. If the interaction is significant, the main effect rows are averages over conditions in which the factor behaves differently, so a statement such as this factor raises the response describes neither condition.

If the interaction is not significant, the main effects describe something real and can be reported on their own, and some analysts then drop the interaction term and refit to return its degrees of freedom to the residual.

Should I report the block row?

Report it, and do not build a conclusion on it. A large block F ratio usually confirms what the design already assumed, that the blocks differ, which is the reason the experiment was blocked in the first place. It is also worth remembering that the blocks were not randomly assigned to anything, so their row describes an observed difference between batches or subjects rather than an effect of anything.

How do I tell the two rank based procedures apart?

What they rank. The one way procedure replaces each observation by its rank in the whole combined sample. The two way procedure, for a design with blocks, replaces each observation by its rank within its own block. The distinction follows from what the design is trying to remove: in the blocked case the block effect is exactly what you do not want the ranking to see, so ranking globally would put it straight back in.

Study strategy

Exam move

Degrees of freedom first, because they are the fastest way to tell which design a question describes and because they are examined directly. Write the four counts for a two factor design with replication and check that they add to the observations minus one; then write the three counts for a block design and notice that the interaction degrees of freedom have become the residual.

Seeing that substitution explains the whole relationship between the two designs. Next comes the reading order, which is where the constructed traps live: interaction row first, always, since a question is often built precisely so that reading the main effects first produces a plausible and wrong answer.

Support that with two interaction plots drawn by hand, one with parallel lines and one with crossing lines, and the reportable sentence written under each, because the sentence is what the marks are for and it is different in the two cases.

Keep the two rank based procedures together as a pair, distinguished only by whether the ranking is global or within block, because that is the single highest value distinction in the assumption failure material.

Working through Two-way and two-factor analysis of variance in DATA2002? Sia is AskSia’s AI Statistics tutor — ask any DATA2002 Two-way and two-factor analysis of variance question and get a clear, step-by-step explanation grounded in how DATA2002 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 64 of your The University of Sydney subjects - and 1,000+ Bibles across every Australian university.
Sia - your DATA2002 tutor, unlimited, worked the way the exam marks it
The full 6-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works
Unlock the full DATA2002 Bible + 64 The University of Sydney subjects
$0.99 Trial