ENVM7003 Chap.6 Response Scales, Coding and Clean Data
Response Scales, Coding and Clean Data
Choosing an anchor family is a measurement decision, not a formatting one. The course supplies a reference set of families drawn from a published compilation: acceptability, appropriateness, importance, agreement, frequency of truth, self-description, belief and priority.
Asking how important coastal protection is measures something different from asking how acceptable a specific coastal restriction is, and treating the two as interchangeable is an argument problem rather than a wording problem.
Beyond the anchors sit two decisions that shape everything downstream: whether to offer a genuine midpoint, and how the responses will be coded.
A published household instrument in the course resources carries its codes inside the instrument itself, with numeric values on every option and indicator tags showing which monitoring programme each variable feeds. Writing the codebook before anything is collected forces you to notice, while it is still cheap, that two options mean the same thing or that one cannot be coded at all.
The chapter closes on the five checks that run before any analysis: shape, range, missingness, consistency, and derivations last.
What this chapter covers
- 01
6.1 Anchor families and what each one actually measures
- 02
6.2 Mirrored ladders and the midpoint decision
- 03
6.3 Writing the codebook before the first response arrives
- 04
6.4 Missing values, other-text and the rules that protect an average
- 05
6.5 Five checks in order: shape, range, missingness, consistency, derivations
- 06
6.6 Matching your categories to published bands before fielding
A codebook and the checks that follow it
- +3One column per variable. Household size, dwelling type, watering days, acceptability, first value selected, and the open reason retained as text.
- +3Give every option a code and write the range down. Watering days is coded zero to three; acceptability one to five with nine reserved for an explicit unsure.
- +2Write the missing rule beside each variable. Blank means not answered, and never zero, because a zero that means both none and unanswered deflates every average silently.
- +2Run the range check. List the distinct values present in each coded column. A seven in a column coded zero to three is an entry error, not a finding.
- +2Handle the legitimate oddities explicitly. An unsure code is excluded from the mean by a written rule; an 'other' response keeps its text in a separate column.
Key terms
- Anchor family
- The set of labels attached to the points of a response ladder, such as acceptability or importance. Choosing the family decides what the resulting number means.
- Midpoint
- The central point of a symmetric response ladder. Labelling it as neutral rather than as no opinion matters, because the two absorb different respondents.
- Codebook
- The written map from every response option to the value stored for it, including the rule for missing data. Written before collection it is design; written after, it is reconstruction.
- Missing-data rule
- The stated convention distinguishing a blank from a real zero and from an explicit unsure. Without it, averages are computed over incompatible things.
- Range check
- Listing the distinct values present in each coded column and comparing them with the codebook. Any value outside the codebook is an error rather than an observation.
- Derived variable
- A value computed from other columns, such as a total from its components. It is built last, after the structural checks, and the rule used is recorded.
- Indicator tag
- A marker attached to a variable showing which monitoring programme or framework it reports into. It is what lets one instrument serve several reporting obligations.
- Verbatim column
- A field preserving an open response exactly as written. It is where the category you should have offered turns up, and it is the source of quotable material.
Response Scales, Coding and Clean Data FAQ
Should a Likert scale have a neutral midpoint?
It is a design decision with consequences either way. A midpoint lets genuinely undecided people say so and also absorbs people who cannot be bothered deciding. If you keep it, label it neutral rather than no opinion and consider a separate explicit unsure option so the two groups do not merge. Whichever you choose, apply it consistently across every scaled item.
How many points should a scale have?
The published anchor families the course points to are mostly seven-point ladders, and five-point versions are common in short instruments. What matters more than the count is symmetry and consistent labelling: every point labelled, the same number either side of the midpoint, and the same direction throughout the instrument.
When should I write the codebook?
Before the first response arrives. Writing it early forces you to discover that two options mean the same thing, or that an option cannot be coded, at a point where the instrument can still be changed. Written afterwards it becomes an exercise in rationalising whatever you collected.
What checks should I run before analysing?
Five, in this order: shape, comparing rows against responses and columns against the codebook; range, listing distinct values per column; missingness, counting blanks per column; consistency, checking pairs that constrain each other; and only then derived variables. Running them out of order means running some of them twice.
Why do my categories have to match published statistics?
Because you cannot realign them afterwards. Education bands split at one point cannot be compared with a published series split at another without collapsing both to a coarser grouping and losing most of the comparison. Look up the reference bands before fielding; it is a ten-minute job that has no equivalent afterwards.
Assessment move
Build the codebook as a table the whole group can see, and have two people independently code the same five responses before anyone codes the rest. Disagreement between them is information about your codes rather than about your coders.
Keep a written note of every rule you invent during cleaning, because those notes become the methods paragraph and the limitations paragraph almost verbatim, and reconstructing them from memory at the end is where accuracy is lost.
Organisational Behaviour · Identifying and Assisting Students at Risk · Digital Marketing and Social Media