DATA2002 Chap.12 Two-way and two-factor analysis of variance
Two-way and two-factor analysis of variance
Real experiments rarely vary one thing at a time. Material arrives in batches, subjects differ from one another, machines drift between days. This chapter handles a second grouping variable in two different roles, and keeping the roles apart matters because they lead to different tables and different conclusions.
In the first role the second variable is a nuisance you can record but do not care about, and the design gives one observation per treatment in each block. Recording it lets the analysis subtract it, which is worth a great deal: the same four treatments measured across five batches can be significant when the batch effect is removed and unremarkable when it is left in the residual, from identical data.
What blocking cannot do is detect an interaction, and the reason is structural rather than computational, since one observation per cell leaves nothing to estimate within cell variation from. In the second role the extra variable is a factor of interest and every combination is observed more than once.
That replication creates a genuine within cell residual, and that in turn allows a fourth source of variation to be separated out. The interaction for a cell is what is left of its mean once the overall mean and both main effects have been accounted for, and it is zero for every cell exactly when the two factors act independently, which is what parallel lines on an interaction plot mean.
When the interaction is significant the main effect rows become misleading, because each is an average over conditions in which the factor behaves differently, so the reportable finding is the effect of one factor stated separately at each level of the other.
What this chapter covers
- 01
Two roles for a second grouping variable, and the different table each produces
- 02
Blocking as the removal of a nuisance source, and the signature of a nuisance factor
- 03
What blocking buys, shown by analysing the same data with and without it
- 04
Why a block design cannot detect an interaction, as a property of the design
- 05
The three row table, and the degrees of freedom that must add
- 06
The block row as a housekeeping note rather than a finding
- 07
Replication, and the fourth row it creates
- 08
The interaction defined by subtraction from the two margins
- 09
Reading an interaction plot: parallel lines against a gap that changes
- 10
Why a main effect becomes misleading once an interaction is present
- 11
Multiple comparison multipliers carried over, and how the gap widens with more levels
- 12
Two rank based procedures, one ranking globally and one ranking within blocks
Count degrees of freedom and say what replication is for
- +1There are three times four times two, that is twenty four observations. The first factor contributes two degrees of freedom, the second contributes three.
- +1.5The interaction contributes the product of those two, which is six. The residual is the observations minus the number of cells, twenty four minus twelve, which is twelve.
- +1The total is twenty three, and the check is that two plus three plus six plus twelve equals twenty three.
- +1.5With one observation per combination there would be twelve observations, eleven total degrees of freedom, and the model would consume all eleven, leaving zero for the residual. With no estimate of the error variance, no F ratio could be formed at all.
Key terms
- Block
- A grouping variable recorded because it is a known nuisance source of variation, rather than because its effect is of interest. Removing it shrinks the residual term and sharpens the comparison that is of interest.
- Randomised block design
- A design in which every treatment appears once within each block. It has no replication inside a cell, so it cannot separate an interaction from the residual.
- Main effect
- The effect of one factor averaged over the levels of the other. It describes something real only when the two factors act independently.
- Interaction
- The part of a cell mean that the two margins do not predict. It is zero for every cell when the factors act independently, and its presence makes each main effect an average over dissimilar conditions.
- Interaction plot
- Cell means plotted against one factor with one line per level of the other. Parallel lines indicate no interaction; lines that converge, diverge or cross indicate one.
- Replication
- More than one observation in each combination of the two factors. It is what creates the within cell residual and therefore what makes an interaction estimable.
- Sum to zero constraint
- A restriction on the effect parameters that makes them identifiable, since otherwise a constant could be moved between the overall mean and the effects without changing any fitted value.
- Fitted value
- The value the model predicts for a cell, formed from the overall mean and the estimated effects. Residuals are computed against it and are the basis of the diagnostic plots.
- Nuisance factor
- A variable that contributes real variation to the response but is not the subject of the experiment. Recording it converts that variation from residual noise into an explained term.
- Within block ranking
- Replacing each observation by its rank among the observations in its own block, which is what the two way rank based procedure does and what distinguishes it from the one way version.
Two-way and two-factor analysis of variance FAQ
What does blocking actually buy me?
It moves a real source of variation out of the residual term and into a row of its own. Since every F ratio has the residual mean square in its denominator, shrinking that denominator makes every comparison sharper. The effect can be decisive: the same treatments can be strongly significant when the block effect is removed and unremarkable when it is not, from identical data.
That is why recording a known nuisance variable at the design stage is worth the trouble even though its own row is rarely of interest.
Why can a block design not test for an interaction?
Because with one observation per treatment per block there is exactly one number in each cell, so nothing is left over to estimate variation within a cell. The interaction and the residual would be the same quantity and the model cannot separate them.
It is a property of the design rather than a limitation of the software, and the only fix is replication, which means more than one observation in each combination rather than more combinations.
In what order should I read a two way table?
The interaction row first, then the main effects. If the interaction is significant, the main effect rows are averages over conditions in which the factor behaves differently, so a statement such as this factor raises the response describes neither condition.
If the interaction is not significant, the main effects describe something real and can be reported on their own, and some analysts then drop the interaction term and refit to return its degrees of freedom to the residual.
Should I report the block row?
Report it, and do not build a conclusion on it. A large block F ratio usually confirms what the design already assumed, that the blocks differ, which is the reason the experiment was blocked in the first place. It is also worth remembering that the blocks were not randomly assigned to anything, so their row describes an observed difference between batches or subjects rather than an effect of anything.
How do I tell the two rank based procedures apart?
What they rank. The one way procedure replaces each observation by its rank in the whole combined sample. The two way procedure, for a design with blocks, replaces each observation by its rank within its own block. The distinction follows from what the design is trying to remove: in the blocked case the block effect is exactly what you do not want the ranking to see, so ranking globally would put it straight back in.
Exam move
Degrees of freedom first, because they are the fastest way to tell which design a question describes and because they are examined directly. Write the four counts for a two factor design with replication and check that they add to the observations minus one; then write the three counts for a block design and notice that the interaction degrees of freedom have become the residual.
Seeing that substitution explains the whole relationship between the two designs. Next comes the reading order, which is where the constructed traps live: interaction row first, always, since a question is often built precisely so that reading the main effects first produces a plausible and wrong answer.
Support that with two interaction plots drawn by hand, one with parallel lines and one with crossing lines, and the reportable sentence written under each, because the sentence is what the marks are for and it is different in the two cases.
Keep the two rank based procedures together as a pair, distinguished only by whether the ranking is global or within block, because that is the single highest value distinction in the assumption failure material.
Working through Two-way and two-factor analysis of variance in DATA2002? Sia is AskSia’s AI Statistics tutor — ask any DATA2002 Two-way and two-factor analysis of variance question and get a clear, step-by-step explanation grounded in how DATA2002 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.