ETC1000 Chap.1 Analysing Categorical Data
Analysing Categorical Data
Name the variable type first
The unit opens on a decision that looks trivial and governs everything after it. A categorical variable records which group a case belongs to; a numerical variable records how much or how many.
The display, the summary statistic and the probability you are allowed to compute all follow from that answer, which is why a mean on payment method and a pie chart on order value are both errors of the same kind.
Four kinds, split by order and by countability
Categorical variables divide into nominal, where the categories carry no order, and ordinal, where they do.
Numerical variables divide into discrete, counted in whole units, and continuous, measurable to any refinement.
Numbers stored in a variable do not settle the question: a postcode is written with digits and is nominal, because adding two postcodes produces nothing.
One table, three probabilities
Crossing two categorical variables gives a contingency table whose interior holds counts and whose edges hold the margins.
A marginal probability divides a row or column total by the grand total, a joint probability divides one interior cell by the grand total, and a conditional probability divides that same cell by its own row or column total.
The counts never move; only the denominator does, and choosing it correctly is where the marks are.
Independence, as this unit defines it
Two variables are independent when knowing the category of one tells you nothing about the probability of the other. The unit's own test is to compute the same conditional probability inside each row and compare them.
Equal conditionals mean independence; unequal ones mean the table carries a relationship, and the size of the gap is the finding worth reporting.
What this chapter covers
- 01
Nominal, ordinal, discrete and continuous, and the test that separates them
- 02
Frequency, relative frequency and percentage as one table written three ways
- 03
Why a sorted bar chart beats a pie chart beyond about four categories
- 04
Reading marginal, joint and conditional probabilities off one contingency table
- 05
Deciding independence by comparing conditional probabilities across rows
Choose the denominator from the wording
- 1Give the probability that a return was refunded.
- 1Give the probability that a return was both in store and refunded.
- 1Decide whether channel and outcome are independent.
Key terms
- Nominal variable
- A categorical variable whose categories carry no order, such as payment method.
- Ordinal variable
- A categorical variable whose categories have a direction but no measured gaps.
- Relative frequency
- A count divided by the total, so the column of them sums to one.
- Contingency table
- A cross of two categorical variables in which every case falls in one cell.
- Marginal probability
- A row or column total divided by the grand total, ignoring the other variable.
- Joint probability
- One interior cell divided by the grand total.
- Conditional probability
- A cell divided by its own row or column total rather than by the grand total.
Analysing Categorical Data FAQ
Can a variable stored as numbers still be categorical?
Yes, and it often is. A postcode, a customer identifier and a store number are all written with digits and none of them supports arithmetic, so they are nominal. The working test is to ask what an average of the values would mean; if the answer is nothing, the variable is categorical however it is stored.
Why does the unit prefer a bar chart to a pie chart?
A bar chart asks readers to compare lengths, which they do accurately, and it can be sorted so that the ranking is visible without reading a single label. A pie chart asks them to compare angles, which they do poorly, and beyond about four slices the ordering of the smaller ones becomes guesswork. The pie makes exactly one point well, that the parts sum to a whole.
What is the difference between a joint and a conditional probability?
Only the denominator. Both use the same interior cell of the table. A joint probability divides it by the grand total and answers how often the two categories occur together across everything. A conditional probability divides it by the total of its own row or column, and answers how often one occurs among the cases that already have the other.
How do I decide independence from a table?
Compute the same conditional probability inside each row and compare them. If they are equal, knowing the row tells you nothing about the column and the two variables are independent. If they differ, the table carries a relationship, and the size of the difference in percentage points is the number worth reporting alongside the verdict.
Exam move
Practise reading the denominator out of the sentence before touching the numbers, because the arithmetic in this topic is never the difficulty. Take any two way table and write out all three probabilities for the same cell, saying which question each one answers, until the choice becomes automatic.
Working through Analysing Categorical Data in ETC1000? Sia is AskSia’s AI Statistics tutor — ask any ETC1000 Analysing Categorical Data question and get a clear, step-by-step explanation grounded in how ETC1000 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.