The University of Sydney · FACULTY OF STATISTICS

DATA2002 Chap.4 Diagnostic accuracy and measures of risk

- one subject, every graph, every model, every mark
12 Chapters7-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 4 of 16 · DATA2002

Diagnostic accuracy and measures of risk

A two by two table of counts is the smallest interesting object in this unit, and almost every quantity here is a ratio of two of its cells. What makes it hard is not the arithmetic but that the same table supports several different questions, each conditioning on a different margin, and swapping two of them produces an answer about a different population rather than merely an imprecise one.

The unit uses two different table layouts in the same week and they do not put the same letters in the same cells: in the diagnostic layout the rows are the test result and the columns are the true status, while in the risk layout the rows are the risk factor and the columns are the outcome. So the first thing to write when a table appears is which layout it is.

Nine measures fall out of the diagnostic layout and they group by denominator: whole table measures such as prevalence and accuracy, measures conditioned on the truth such as sensitivity and specificity, and measures conditioned on the test result, the predictive values.

The pair most often swapped is sensitivity and positive predictive value, and the difference matters because sensitivity is a property of the instrument while predictive value is a property of the instrument and the population it is used in. That is the base rate effect, and it is best worked in whole people rather than in probabilities. Turning to the risk layout, two ratios summarise an association.

Relative risk compares the probability of the outcome with and without exposure and is only available when the design lets you observe those probabilities. The odds ratio compares odds instead, reduces to the cross product of the table, and is the same number whichever margin is conditioned on, which is precisely why it survives a retrospective design.

Its interval is built on the log scale, where the neutral value moves from one to zero.

In this chapter

What this chapter covers

  • 01

    Two table layouts in one week, and why the first move is naming which one you have

  • 02

    The four cells, and why the first word grades the test rather than the news

  • 03

    Nine measures grouped by denominator: whole table, conditioned on truth, conditioned on result

  • 04

    Sensitivity against positive predictive value, and why only one is a property of the test

  • 05

    Conditional probability run in the direction a patient cares about

  • 06

    Working in natural frequencies rather than probabilities, which matters without a calculator

  • 07

    Why accuracy is close to meaningless on a rare condition

  • 08

    Relative risk, its interpretation ladder and its unbounded behaviour

  • 09

    Odds, the odds ratio, and the invariance that makes it design independent

  • 10

    Which measure each study design permits, and why a retrospective study fixes the outcome margin

  • 11

    How the uncertainty in a log odds ratio is estimated, and why every cell must be reasonably large

  • 12

    Reading an interval for a ratio against one rather than against zero

Worked example · free

Say which measure a design permits, and correct a conclusion

Q [6 marks]. Investigators recruit 200 people who have already developed a condition and 200 matched people who have not, then ask each about a past exposure. They report that 30 per cent of cases and 15 per cent of controls were exposed, and conclude that the exposure doubles the risk of the condition. Identify the design, state which measure of association they were entitled to compute, and correct the conclusion. (6 marks. The mark allocation is ours, not the University's.)
  • +1The design is retrospective: participants were selected on the outcome and the exposure was recalled afterwards. That single fact decides everything that follows.
  • +2The 30 per cent and 15 per cent are proportions exposed within groups whose sizes the investigators chose, so no risk of the condition can be computed from them and relative risk is not available.
  • +2What is available is the odds ratio. The odds of exposure are 60 to 140 among cases and 30 to 170 among controls, so the odds ratio is 60 times 170 divided by 140 times 30, which is 2.43.
  • +1The claim that risk doubles is unsupported twice over: no risk was measured, and recall of a past exposure by people who already have the condition is subject to recall bias, which can inflate the reported difference.
A retrospective design, so only the odds ratio may be computed, and it is 2.43 rather than a doubling of risk. The conclusion should say that the odds of exposure were higher among cases, and should name recall as a threat to the comparison.
Sia tip — The giveaway is always a sentence saying how many participants came from where. Once the investigator has chosen how many cases and how many controls to recruit, the outcome margin is a design constant and no probability of the outcome can be read off the table.
Glossary

Key terms

Sensitivity
The proportion of those who have the condition whom the test correctly flags. It conditions on the truth and is a property of the instrument rather than of the population it is used in.
Specificity
The proportion of those without the condition whom the test correctly clears. Like sensitivity it conditions on the truth, so it does not change when prevalence changes.
Positive predictive value
The proportion of positive results that are correct. It conditions on the test result and therefore depends on prevalence as well as on the instrument.
Prevalence
The proportion of the population that has the condition. It is the ingredient that converts sensitivity and specificity into predictive values.
Base rate effect
The phenomenon by which a test with high sensitivity and specificity still produces mostly false positives when the condition is rare, because the small false positive rate is applied to a very large group.
Relative risk
The probability of the outcome under exposure divided by the probability without it. It requires a design in which those probabilities can be observed, so it is unavailable in a retrospective study.
Odds
The probability of an event divided by the probability of its complement. Odds run from zero to infinity with a neutral value of one, where a probability runs from zero to one.
Odds ratio
The odds of the outcome under exposure divided by the odds without it, equal to the cross product of the table. It is identical whichever margin is conditioned on, which is why it can be estimated from a retrospective study.
Prospective study
A design that defines exposure groups and follows them forward to the outcome, so that the outcome proportions are observed rather than chosen.
Retrospective study
A design that recruits participants on the outcome and looks back at exposure. It fixes the outcome margin, which is what removes relative risk from the available measures.
FAQ

Diagnostic accuracy and measures of risk FAQ

Why does a very accurate test still mostly raise false alarms?

Because accuracy conditions on the truth and a positive result does not. Take a test that is 90 per cent sensitive and 95 per cent specific used where 2 per cent of people have the condition. In 20,000 people, 400 have it and 360 are flagged; 19,600 do not and 5 per cent of them, 980, are also flagged. So 360 of 1,340 positives are correct, about 27 per cent.

Nothing is contradictory: sensitivity is computed among people who have the condition, predictive value among people who tested positive, and when the condition is rare the second group is dominated by healthy people.

What is the fastest way to do these conversions without a calculator?

Work in whole people. Choose a convenient population size, split it by the base rate, apply the two rates to the two groups, and read the answer off the bottom row as a ratio of counts. The algebra and the natural frequency tree are the same calculation, and the tree is far harder to get wrong under pressure. It also produces an answer you can explain to a non specialist, which is often what the question is really asking for.

When can I compute relative risk and when can I not?

You can compute it when the design lets you observe the probability of the outcome in each exposure group, which means a design that follows exposure groups forward, or a set of records in which every outcome is already known. You cannot compute it when the investigator chose how many people with and without the outcome to recruit, because then the outcome proportions are design constants rather than observations.

The odds ratio is available in both cases, and that is the whole reason the measure exists.

Why is the interval for an odds ratio built on the log scale?

Because the odds ratio lives on the positive half line with its neutral value at one, so its sampling distribution is skewed and a symmetric interval built directly on it would be awkward and could extend below zero. Taking logarithms moves the neutral value to zero and makes the distribution roughly symmetric, which is the shape the standard interval assumes.

The interval is constructed on that scale and then exponentiated, which is why the result is not symmetric about the estimate.

How do I read an interval for a ratio?

Against one, not against zero. The interval is compatible with no association exactly when it contains one, which is the ratio equivalent of a difference interval containing zero. Comparing an odds ratio or a relative risk against zero is a common slip and it makes every interval look significant, since both quantities are positive by construction.

The condition on the interval is also worth remembering: it needs all four cells to be reasonably large, because the standard error is a sum of reciprocals of the cell counts.

Study strategy

Exam move

Nine formulas is the wrong thing to carry into a paper and one method is the right thing. Write the four cells with both labels on the margins and ask which group the question restricts you to, because the restriction fixes the denominator and the denominator is the measure. The phrase among people who have the condition puts you in one margin; the phrase among people who tested positive puts you in the other.

That method survives a table drawn in the opposite orientation, which memorised letters do not. Three habits sit around it. Do the base rate conversion in whole people, at least five times on invented numbers, choosing populations of ten thousand or a hundred thousand so the arithmetic stays mental and the tree becomes automatic.

Make a two row table of your own with prospective and retrospective down the side and relative risk and odds ratio across the top, and write in each cell whether it is available and why, since a large share of the examinable content in this chapter is that single distinction.

Check every direction against a limiting case: a perfect test has sensitivity one and a false negative rate of zero, so if your two numbers do not sum to one, one of them is the wrong measure.

Working through Diagnostic accuracy and measures of risk in DATA2002? Sia is AskSia’s AI Statistics tutor — ask any DATA2002 Diagnostic accuracy and measures of risk question and get a clear, step-by-step explanation grounded in how DATA2002 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 64 of your The University of Sydney subjects - and 1,000+ Bibles across every Australian university.
Sia - your DATA2002 tutor, unlimited, worked the way the exam marks it
The full 7-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works
Unlock the full DATA2002 Bible + 64 The University of Sydney subjects
$0.99 Trial