PHC5115 Chap.3 Retrieval Practice and the Testing Effect
Retrieval Practice and the Testing Effect
The claim the course is organised around
The second session puts its central result plainly. Recall is not simply an instrument for reading learning off; it is one of the things that manufactures learning in the first place.
The testing effect names a result: pulling material back out of memory leaves more of it behind, over the long run, than spending the identical stretch of time going over the same material again. The word equivalent carries the argument.
This is not a claim that more work produces more learning; it is a claim that the same time spent differently produces a different outcome.
The reference evidence is the crossover reported by Roediger and Karpicke in 2006. Participants worked through a passage of science writing.
One condition then went over it again and again; a second spent the identical stretch of time writing out whatever they could still produce, uncorrected. A final recall test followed either five minutes later or one week later.
At the short interval the restudy condition led, with about 83 per cent of the passage produced against about 71. A week out the ordering had flipped, at roughly 61 per cent for the retrieval group against 40 for the restudy group.
Same material, same time on task, opposite long-term outcome.
Three properties that follow from that shape
The effect is effortful rather than passive, so what pays is building the answer rather than laying eyes on it a second time. Open recall and short written answers beat a choice of options, because producing beats selecting.
And the gain is slight or negative when the check follows at once, then opens up as the interval grows; that ordering is why the better option feels like the worse one at the point of choosing.
Why effort is the mechanism rather than a side effect
The course explains the crossover through desirable difficulties, and makes it predictive by splitting memory into two quantities that move independently.
Storage strength describes how deeply a memory sits, and it moves in one direction only, upward. Retrieval strength is how accessible it is right now, and it fluctuates and fades with disuse. What follows is a rule about effort: among recalls that come off, the ones that cost most add most to how deeply the material sits.
Reaching for something already at your fingertips changes very little; reaching for something that nearly would not come is what makes it last.
At the level of tissue, three things happen. Bringing a memory back puts the same group of cells into play again, and driving that group together tightens the links that constitute it.
Effortful retrieval engages prefrontal cortex and sharpens the representation more than passive review does. And a reactivated memory becomes briefly labile before being re-stabilised, so each successful recall can update as well as strengthen the trace.
That third property is why feedback delivered after an attempt is worth far more than the same feedback delivered before one.
Recognition is not retrieval
This is the distinction the assessed question writing rests on.
Under recognition the candidate answer is already on the page and the learner only decides whether it looks acceptable; under retrieval nothing is on the page, so the answer has to be built from what remains in memory. A multiple-choice item solvable by elimination is recognition in disguise.
The fastest test is to hide the choices and see whether your own stem is answerable: if you cannot, the item needs rewriting.
Designing retrieval that a group will actually do has six properties, and they function as a set. Keep the weight in the grade near zero, so nerves stay down and learners stop hiding gaps. Run it often and keep each one short. Pitch it so the answer has to be built and usually arrives.
Let feedback follow the effort rather than precede it. Vary the format so recall generalises beyond the cue it was practised on. And make it spaced and cumulative rather than confined to the current week.
What this chapter covers
- 01
The testing effect stated as a mechanism rather than a measurement
- 02
The 2006 crossover result and why the delay decides the ranking
- 03
Storage strength and retrieval strength moving independently
- 04
Reinstatement, prefrontal engagement and updating during recall
- 05
Recognition against retrieval as the test for any practice item
- 06
Six design properties for low-stakes retrieval that a group will sustain
Converting a recognition item into a retrieval item
- 2State why the item can be answered without retrieval.
- 1Identify the second cue that makes it easier still.
- 2Rewrite the stem so the answer must be produced.
- 2Rebuild the distractors so elimination fails.
Key terms
- Testing Effect
- The result that pulling material back out of memory leaves more behind over the long run than spending the identical time going over it again.
- Storage Strength
- How deeply a memory sits. It moves upward only, and a costlier recall that comes off adds more of it.
- Retrieval Strength
- How reachable a memory is at this moment. It climbs after recent contact and slides away without it.
- Desirable Difficulty
- A condition that lowers immediate performance while improving long-term retention, provided the retrieval succeeds and is corrected.
- Recognition
- Deciding whether an option already on the page looks acceptable, on a sense of familiarity, without building anything.
- Reconsolidation
- The brief window after recall in which a reactivated memory can be updated as well as strengthened.
- Low-Stakes Retrieval
- Short, repeated recall carrying almost no weight in the grade, built so that building rather than judging is the point.
Retrieval Practice and the Testing Effect FAQ
What makes the 2006 crossover result so persuasive?
Both groups spent the same time on the same passage, so the comparison isolates what was done with that time. At five minutes the restudy group led and at one week the ranking reversed. Because the only variable is the activity, the result cannot be explained by effort, motivation or exposure, and it also explains why learners reliably choose the weaker option.
Why does the benefit grow with the gap rather than shrinking?
Recent exposure inflates immediate scores for both groups and then decays quickly, so it masks the difference at short intervals. What recall builds is the deeper property that does not decay in the same way, and it only becomes visible once the temporary advantage has gone. The longer the gap, the more of the score reflects that deeper property.
Is a multiple-choice question ever useful for practice?
Yes, when the stem carries the work. An item whose answer can be produced with the options covered demands genuine recall, and plausible distractors drawn from real misconceptions close the route through elimination. The format is weaker than free recall, but a well-built item is far better than a badly built short-answer question.
Should learners receive feedback before or after they attempt recall?
After. A reactivated memory becomes briefly open to updating, so correction arriving at that point is absorbed into the trace. Supplying the answer first removes the attempt altogether and converts the task back into reading, which is the activity the whole result argues against.
Why insist that practice be low stakes?
Because the purpose is learning rather than judgement, and stakes change behaviour. Low weighting keeps anxiety down and honesty up, so gaps are revealed rather than concealed. High-stakes testing measures reliably and teaches poorly, which is the opposite of what a weekly opener is for.
What happens when weekly recall questions start to feel easy?
Retrieval strength has risen while the questions stayed fixed, so the recall is no longer effortful and has stopped building. The response is not harder wording but a change of span and cue: make the set cumulative across earlier weeks, vary the format, and lengthen the interval before any item returns.
Assessment move
Write one practice item a week and run the covered-options test on it before anything else. If the stem cannot be answered alone, rewrite the stem rather than the options, because the stem is where retrieval demand lives and it is the part the marking criteria reward.
Working through Retrieval Practice and the Testing Effect in PHC5115? Sia is AskSia’s AI Education tutor — ask any PHC5115 Retrieval Practice and the Testing Effect question and get a clear, step-by-step explanation grounded in how PHC5115 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.