PMGM7023 Chap.10 Data Traps and Defensible Descriptive Claims
Data Traps and Defensible Descriptive Claims
Who is actually in the data
The fourth session is a clinic with four checks to run before defending any claim built on data. The first is selection, where sample size is least protective. The course uses a 1936 magazine poll that collected about 2.4 million replies and predicted the wrong winner: the result gave Roosevelt 27,750,866 votes at 60.8% against Landon's 16,679,683 at 36.5%.
The lists used to reach people covered some voters far better than others, and those who chose to reply differed systematically from those who did not. Neither problem is repaired by collecting more replies through the same channel.
Two selection problems recur in business data.
Survivorship arises when only the cases that persisted can be observed, so customers who left are absent from the satisfaction survey and failed projects are absent from the lessons repository.
Range restriction arises when a group has already been selected on something close to the variable under study, so a relationship measured among admitted students, hired candidates or approved borrowers can look weak inside that group while being strong in the wider population.
Are you measuring what you mean
The second check is measurement validity, taught through three mismatches.
A proxy mismatch uses a stand-in that rewards the wrong behaviour, such as counting words added to judge a contribution. An outcome mismatch captures an intermediate step rather than the result, such as counting matches to judge whether an application helps people form relationships.
A definition mismatch compares two figures that were never built the same way.
The test that works on your own material is to write the concept in one sentence without using the measure, then write what the measure would record if the concept were entirely absent.
A team member who contributes by deleting weak analysis adds no words, so a word count records zero for the most valuable work done.
Summaries that mislead while remaining correct
The third check concerns descriptive statistics, and every example is arithmetically right, which is what makes the family dangerous.
The mean adds every value and divides by the count, so a few very large values pull it upward; the median is the middle value and describes the typical case; the mode is the most frequent value and can sit far from both.
For earnings, waiting times and most operational quantities the distribution has a long tail, so quoting the mean alone describes someone who may not exist.
Spread matters as much as position, since two groups can share an average and behave completely differently, and a mean reported without a measure of spread is half a description.
Denominators decide meaning: a percentage on a tiny base can describe one person, and a reassuring overall rate can conceal a subgroup whose rate is very different.
Count, rate and individual risk are three separate claims.
Repairing a claim in three moves
The session closes with published claims and one instruction: name the worry, say what evidence would settle it, then restate the conclusion so the evidence covers it.
That structure is also the shape of a good examination answer, because it separates what was observed from what was inferred and states the limit alongside the claim rather than after it.
What this chapter covers
- 01
Undercoverage, nonresponse, survivorship and range restriction
- 02
Proxy, outcome and definition mismatches in measurement
- 03
Mean, median, mode, spread and the denominator
- 04
Concern, missing information, rewritten conclusion
Repairing a shift-safety claim
- 3Name the concern precisely.
- 3Name the information that would settle it.
- 3Rewrite the conclusion to fit the evidence.
Key terms
- Undercoverage
- Undercoverage occurs when the method of reaching people leaves part of the target population systematically unreachable, so the sample cannot represent it.
- Nonresponse
- Nonresponse occurs when those who answer differ systematically from those who do not, so the responses describe a self-selected group rather than the population.
- Range Restriction
- Range restriction is the narrowing of a variable's spread caused by prior selection, which can hide a relationship that exists in the wider population.
- Base Rate
- A base rate is the underlying frequency against which a reported figure should be read. Without it a percentage can describe a very small number of cases.
Data Traps and Defensible Descriptive Claims FAQ
Why can a very large sample still give the wrong answer?
Because size does not repair selection. If the method of reaching people leaves part of the population unreachable, or if those who respond differ systematically from those who do not, then collecting more responses through the same channel reproduces the same bias at greater volume. The question to ask is how the sample was chosen rather than how many observations it holds.
When should I report the median rather than the mean?
Report both whenever they differ materially, and say why they differ. The median describes the typical case and the mean carries the total exposure, so for quantities with a long tail such as claim sizes or resolution times the gap between them is itself the finding. Add the number of observations and a measure of spread, since a mean without spread is half a description.
Exam move
Collect one published claim a week and run the three-move repair on it in writing. Rehearse naming the denominator before the comparison until it becomes automatic. For any measure you use in the project, write the sentence describing what it would record if the concept were absent.
Working through Data Traps and Defensible Descriptive Claims in PMGM7023? Sia is AskSia’s AI Management tutor — ask any PMGM7023 Data Traps and Defensible Descriptive Claims question and get a clear, step-by-step explanation grounded in how PMGM7023 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.