The University of Sydney · FACULTY OF STATISTICS

DATA2002 Chap.2 Study design: surveys, experiments and observation

- one subject, every graph, every model, every mark
12 Chapters8-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 2 of 16 · DATA2002

Study design: surveys, experiments and observation

This chapter contains no test statistic and it decides the marks on the questions that do. Every method later in the unit ends in a sentence, and what that sentence is allowed to say is fixed here, by how the data came to exist. Two studies can produce an identical table of counts and an identical p-value while only one of them licenses the word because.

The unit opens with a famous prediction failure for exactly this reason: a 1936 mail survey went to ten million people, drew 2.4 million replies and called the election wrongly, while a competing forecast built on far fewer people called it correctly.

The rule drawn from it is worth memorising in its exact form, that when a selection procedure is biased a larger sample does not help, because it repeats the same mistake at a larger scale and attaches a narrower interval to it.

Three named bias families follow: selection bias, where the frame does not match the population; non response bias, where those who answer differ systematically from those who do not; and measurement bias, where the instrument itself moves the answer, as in a study where the reported answer to a trust question changed from 35 per cent to 7 per cent depending on who was asking.

Against that background the unit sets the randomised controlled double blind design as the standard, and treats it as an ordered procedure with five steps, because questions typically remove exactly one feature and ask what protection is lost.

The property randomisation supplies is worth stating precisely: it does not make the two groups identical, it makes them similar in expectation on every variable at once, including variables nobody measured. Most interesting questions cannot be randomised, so the chapter closes on observational studies, confounding and the reversal that a confounder can produce when subgroups are pooled.

In this chapter

What this chapter covers

  • 01

    Parameter against statistic: what the researcher wants to know against what the researcher knows

  • 02

    Why sample at all, and the sampling procedure questions that follow

  • 03

    Selection bias, non response bias and measurement bias, each with what goes wrong

  • 04

    The two stage failure in the 1936 survey, and why a larger frame would not have fixed either stage

  • 05

    Precision against accuracy, and which of the two sample size buys

  • 06

    The randomised controlled double blind design as five ordered steps

  • 07

    What randomisation guarantees, stated as similarity in expectation on unmeasured variables

  • 08

    Where a good trial leaks its guarantee: dropout and differential adherence

  • 09

    Association against causation, and the verb each design licenses

  • 10

    Confounding and its two conditions, and why checking only one is the usual error

  • 11

    Controlling for a variable by splitting on it, and the three limitations of that repair

  • 12

    The reversal paradox, also called Simpson's paradox: a trend inside every subgroup that vanishes or flips when they are pooled

Worked example · free

Diagnose a design and write a conclusion that does not overreach

Q [8 marks]. A recruitment platform offers an optional interview coaching module. Among users who completed it, 46 per cent received an offer within three months; among users who did not, 29 per cent did. The platform announces that the module raises the offer rate by 17 percentage points. Classify the study, identify the principal threat to the causal claim, propose a design that would settle it, and write the strongest conclusion the existing data support. (8 marks. The mark allocation is ours, not the University's.)
  • +2Classify the design. Users chose for themselves whether to complete the module, so nobody allocated the exposure. This is an observational comparison of two self selected groups, and the ceiling on its conclusions is association.
  • +2Name the confounder with both of its arrows. Motivation and job readiness plausibly drive completion of an optional module, and the same qualities plausibly drive receiving an offer. That is related to the exposure and related to the outcome without being a step on the path from one to the other, which is the definition.
  • +1Say what the reported difference is mixing. The seventeen point gap contains the effect of the module and the effect of being the kind of candidate who completes optional modules, and the data give no way to separate them.
  • +2Propose the design that settles it. Randomly allocate access to the module among users who want it, keep the others on the standard product, and compare offer rates. Randomisation balances motivation and every other unmeasured trait in expectation, which subgroup analysis cannot do.
  • +1Write the conclusion the current data support: completing the module is associated with a higher offer rate in these users, and because participation was self selected the difference cannot be attributed to the module.
An observational study of self selected groups, threatened principally by confounding with motivation and readiness, settled by randomising access among users who want the module, and reportable only as an association until that is done.
Sia tip — The marks are not for scepticism in general. They are for classifying the design, naming a confounder and showing both of its relationships, and closing with a sentence whose verb matches the design. Writing only that correlation is not causation makes none of those moves.
Glossary

Key terms

Sampling frame
The list or mechanism from which a sample is actually drawn, as distinct from the population the conclusion is about. A mismatch between the two is selection bias and it is not repaired by drawing more names from the same list.
Non response bias
The distortion that arises when those who choose not to participate differ systematically from those who do. A low response rate is a warning sign rather than a defect in itself, since the question is always whether the missing group would have answered differently.
Measurement bias
A distortion introduced by the instrument itself, through question wording, question order, or the identity of the person asking. It changes the recorded answer without changing the underlying opinion.
Randomised controlled trial
A design in which the investigator allocates subjects to treatment and control at random, uses a placebo, and keeps subjects and investigators blind to the allocation. It is the only design in this chapter that can support a causal claim.
Placebo
An inert comparison given to the control group so that the effect of the treatment is separated from the effect of being treated at all.
Double blind
Blinding both the subjects and the investigators to the allocation. Subject blinding protects the response and investigator blinding protects the measurement and the decisions made during the trial.
Observational study
A study in which the exposure is observed rather than assigned. It can establish association and may suggest causation, and it cannot prove it.
Confounder
A variable related to the exposure and related to the outcome, and not on the causal path between them. Both relationships are required, and checking only one is the standard incomplete answer.
Controlling for
Dividing the sample into subgroups defined by a suspected confounder and comparing within each. It narrows the gap between an observational study and an experiment without closing it.
Reversal paradox (Simpson's paradox)
Also called Simpson's paradox: the effect in which an association between two variables changes sign once a third variable is conditioned on, whatever value that third variable takes. It is the standard consequence of a confounder in an observational comparison.
FAQ

Study design: surveys, experiments and observation FAQ

If a survey gets millions of responses, is it not reliable?

Not if the procedure that produced them is biased. The unit's opening example is the case in point: a survey with 2.4 million replies called an election wrongly while a much smaller well drawn sample called it correctly. Sample size buys precision, meaning a narrower interval around whatever the procedure is estimating. It does not buy accuracy, meaning that the thing being estimated is the quantity you wanted.

A biased procedure applied at scale produces a confident wrong answer, which is worse than an uncertain one.

Both selection bias and non response bias leave people out. Which is which?

Selection bias happens before anyone is asked: the list you drew from does not represent the population you want to describe. Non response bias happens afterwards: among those asked, the ones who answer differ systematically from the ones who do not. They can compound, and the unit's own example does exactly that, since the addresses came from sources skewed towards wealthier households and then only a quarter of them replied.

A complete answer names both stages when both are present.

Why does an observational study never establish causation?

Because the groups being compared assembled themselves, so any difference between them includes everything else that differs between them. A confounder is a specific version of this: a variable related to both the exposure and the outcome will move the comparison in a way the data cannot separate from the exposure's own effect.

Randomisation solves the problem at a stroke by balancing every variable in expectation, including variables nobody measured, which is exactly what subgroup analysis cannot do.

Is controlling for a confounder enough to fix an observational study?

It helps and it does not close the gap, for three reasons that a question asking about limitations wants all of. You can only control for variables you measured, so unmeasured confounders remain untouched. Each split reduces the number of observations per cell, so precision falls. And within a subgroup the groups are still not randomised, so they may still differ on something else.

The strategy narrows the distance to an experiment; it never removes it.

How do I spot the reversal paradox (Simpson's paradox) in a table?

Compute the rates within every subgroup and then pooled, and compare the direction. A reversal requires two things at once: the subgroup variable must be related to the outcome, and the groups being compared must have very different mixes across those subgroups. Both are readable straight off the counts.

Note also that an unequal mix is necessary but not sufficient: if one group leads or ties in every subgroup, no weighting of those subgroups can put the other in front.

Study strategy

Exam move

This chapter rewards drilling one move rather than memorising a list. Take any claim you meet, in the unit or outside it, and answer four questions: who allocated the exposure, what population is the claim about and what was the frame, who is missing and would they have answered differently, and is there a variable related to both the exposure and the outcome.

Then write the conclusion with the verb the design permits, describes, is associated with, or causes. Ten repetitions of that will do more for you than rereading the bias definitions, because the definitions are easy and applying them under time pressure is not.

Alongside the drill, learn the five step randomised design as an ordered procedure, since questions typically remove one feature and ask what is lost: no control arm and improvement cannot be separated from natural recovery, no blinding of the assessor and the measurement becomes an outcome of knowing the allocation, no randomisation and every unmeasured difference is back in play.

And practise the confounder test with both arrows out loud, because a variable that differs between the groups but is unrelated to the outcome is an imbalance rather than a confounder, and calling it one is a definition error that costs the mark.

Working through Study design: surveys, experiments and observation in DATA2002? Sia is AskSia’s AI Statistics tutor — ask any DATA2002 Study design: surveys, experiments and observation question and get a clear, step-by-step explanation grounded in how DATA2002 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 64 of your The University of Sydney subjects - and 1,000+ Bibles across every Australian university.
Sia - your DATA2002 tutor, unlimited, worked the way the exam marks it
The full 8-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works
Unlock the full DATA2002 Bible + 64 The University of Sydney subjects
$0.99 Trial