The Education University of Hong Kong · FACULTY OF EDUCATION

PES6252 Chap.10 Observer Error, Reliability and Coach Development

- one subject, every graph, every model, every mark
6 Chapters3-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 10 of 11 · PES6252

Observer Error, Reliability and Coach Development

A corrupted record looks exactly like a sound one

Nothing in the method chapters works if the record is wrong, and the dangerous property of a wrong record is that it is indistinguishable from a right one by inspection. A drifting coder produces a complete sheet. A biased coder produces plausible totals. A coach who knows he is being filmed produces a session.

The only way to know is to have built the checks in beforehand, which is why this chapter belongs before data collection rather than after it.

Five classic threats are named, and none of them is a mark of carelessness: published studies report guarding against all five, which is an admission that all five happen to competent people working carefully.

The five, grouped by where the error enters

Drift and complexity are properties of a coding system meeting a human being.

Drift is the observer gradually changing the coding rules or interpreting them differently over time, with definitions loosening or tightening across a long collection period and two observers possibly drifting in opposite directions. Complexity is a category set larger than working memory: more categories means more decisions per second and accuracy falls as load rises, with live coding far more error prone than video.

Bias and cheating are properties of the coder relationship to the study: what a coder expects of the coach, or of the study, tips the borderline decisions, and completing a sheet later from memory is the everyday version, far commoner than fabrication. Reactivity is the only one that is not about the observer at all, and the only one where the data are faithfully recorded and still do not describe ordinary coaching.

In this chapter

What this chapter covers

  • 01

    Why a corrupted record cannot be identified by looking at it

  • 02

    Drift: definitions moving, and invisible to the person moving them

  • 03

    Complexity: category load, live coding, and the cost of simplifying

  • 04

    Bias on borderline calls, and why it is worse than random error

  • 05

    Reactivity, habituation, and the sessions you plan to discard

  • 06

    Cheating in its ordinary form: sheets completed from memory

Worked example · free

Diagnose four observation failures and triage them

Q [12 marks]. AskSia authored practice. A group codes twelve sessions across three coaches over five weeks. Week one shows the highest praise and lowest scold rates of the study, and the coaches had asked who would see the footage. Coder agreement is 0.88 in week one and not measured again until week five, when it is 0.61. One coder, told coach C is the most experienced, codes his ambiguous utterances as instruction and other coaches identical utterances as management. A fourth coder missed a session and completed the sheet that evening from memory. Diagnose each and say which are recoverable. The marks shown are an AskSia study allocation and are not the University marking scheme.
  • 4Name each of the four failures.
  • 4Say what is recoverable and what is not, with the reason.
  • 4Name the two habits that would have prevented three of the four.
Week one is observer reactivity. It is recoverable and the fix is cheap, since reactivity fades with habituation, and the right decision is to discard week one rather than keep it and mention it, because a habituation effect is a systematic shift in one direction rather than noise that averages out. The fall from 0.88 to 0.61 is observer drift, invisible without periodic re checks, which is why it surfaced only at the end. It is not fully recoverable: four weeks sit between two measurements and nothing identifies when the drift began, so the options are re coding the middle weeks against the original criterion or reporting the range as a limitation. The third is observer bias, worse than random error because it aligns with the comparison being made, and recoverable only by re coding blind with coach identity removed from the footage labels. The fourth is cheating in the technical sense, which needs no dishonesty: a sheet completed from memory is complete, plausible and worthless, and that session should be dropped. Three of the four would have been prevented by two habits, habituating before recording and scheduling the reliability check into the collection period rather than onto the end of it.
Sia tip — Write the reliability schedule into the project plan on the day you choose the instrument. Every guard in this chapter is something done before or during collection, and none of them can be applied afterwards.
Glossary

Key terms

Observer Bias
The steady tilting of borderline decisions by what a coder already expects of the coach or of the study, reduced by hiding whose session it is.
Intra Observer Reliability
The agreement of one coder with themselves on the same footage at two times, which is the direct test for drift.
Habituation
Filming a coach and squad repeatedly before any analysed session, so that the change caused by being watched has faded before data collection begins.
FAQ

Observer Error, Reliability and Coach Development FAQ

Why is drift the threat that costs most?

Because of what it does to a before and after design, which is the commonest shape in this field. If a coder definition of post instruction loosens across eight weeks, the second observation shows a higher rate whatever the coach did, and the study reports exactly the change the intervention was meant to produce while the coder has no sense of having changed anything.

That is why the reliability check belongs in the middle of the collection period rather than at the end, and why the check costs one session of double coding and saves the whole study.

Is it better to use a simpler instrument if agreement is poor?

Only if the question survives the simplification. More categories means more decisions per second and accuracy falls as load rises, so a smaller set will raise agreement, but simplifying may cost validity and produce a record two coders agree about that no longer captures the thing being studied. The test is the research question.

If the question is how much feedback a coach gives, merging feedback categories is fine; if it is what kind, merging them deletes the answer and the correct response is more coder training.

Study strategy

Assessment move

Code the same five minutes of footage twice, a week apart, without looking at your first sheet. The disagreement between you and yourself is your own drift rate, it takes twenty minutes to measure, and it is the fastest way to stop treating reliability as an administrative requirement rather than as a property of your own attention.

Working through Observer Error, Reliability and Coach Development in PES6252? Sia is AskSia’s AI Education tutor — ask any PES6252 Observer Error, Reliability and Coach Development question and get a clear, step-by-step explanation grounded in how PES6252 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 3 of your The Education University of Hong Kong subjects - and 1,000+ Bibles across every Australian university.
Sia - your PES6252 tutor, unlimited, worked the way the exam marks it
The full 3-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works