The University of Melbourne · FACULTY OF PSYCHOLOGY

PSYC10006 Chap.2 Operant Conditioning and Social Cognitive Learning

- one subject, every graph, every model, every mark
12 Chapters8-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 2 of 14 · PSYC10006

Operant Conditioning and Social Cognitive Learning

Classical conditioning explains reflexes and has nothing to say about behaviour that is emitted rather than elicited. Operant conditioning fills the gap: voluntary behaviour is shaped by what follows it. Every consequence takes two labels, and they have to be assigned in a fixed order because answering the second question first is the most reliable way to get the label wrong.

Did the behaviour become more likely or less likely? Was something added or removed? The module then moves through the four schedules and the response pattern each produces, shaping by successive approximation, and three conditions that govern whether punishment works at all.

It closes with two lines of evidence that end the behaviourist restriction to observable events: a stored spatial map, learning that stays invisible until there is a reason to use it, and consequences that reach a learner by observation rather than by experience.

In this chapter

What this chapter covers

  • 01

    Why voluntary behaviour needs a second account, and what an operant is

  • 02

    The two questions, in order: did behaviour rise or fall, and was something added or removed

  • 03

    Why negative means subtraction, and why negative reinforcement raises behaviour

  • 04

    Ratio against interval, fixed against variable, and the four schedules they generate

  • 05

    The pattern each schedule produces, and why ratio schedules run steeper

  • 06

    Shaping: reinforce the closest approximation, withhold, then move the criterion

  • 07

    The partial reinforcement effect, and why persistence is engineered separately

  • 08

    Contingency, contiguity and consistency as conditions on effective punishment

  • 09

    Why inconsistent punishment can leave a behaviour stronger than none would

  • 10

    Four costs of punishment, and three constructive alternatives to it

  • 11

    The blocked-route study, and why its design lets two accounts predict differently

  • 12

    Latent learning, and learning separated from performance by an incentive condition

Worked example · free

Design a shaping programme and state the schedule at each stage

Q [5 marks]. A support worker wants a client who currently stays in their room to join a shared meal at the far end of a common area. Design the programme, name the schedule at each stage, and say what you would do once the target behaviour exists. This mark allocation is our own and is not an official assessment scheme.
  • 1Define the target behaviour concretely: sits at the shared table for the length of one meal. An abstract target cannot be approximated, because you cannot tell whether a given behaviour is closer to it.
  • 1Find the closest current approximation, whatever the client already does that points at the target, and reinforce it on a continuous schedule, because continuous reinforcement is what establishes a behaviour fastest.
  • 1Withhold once the approximation is reliable. Behaviour becomes more variable when reinforcement stops, and that variability produces the slightly closer behaviour you will reinforce next.
  • 1Move the criterion and reinforce the new approximation continuously, repeating the cycle. Moving the criterion before the current step is reliable extinguishes the approximation and loses ground.
  • 1Switch schedules once the target behaviour exists. Move from continuous to a partial schedule so the behaviour resists extinction, and say explicitly that this is a different job from shaping.
Continuous reinforcement of successive approximations while the behaviour is being built, then a partial schedule to make it persist. The examinable point is that acquisition and persistence are separate problems with different solutions, so a scenario that mentions a behaviour being taught and then maintained is asking for both, in that order. This item is AskSia practice and carries no University marking scheme.
Sia tip — On scrap paper put up or down on one line and added or removed on the next, then read the label off the pair. Answer the direction question first: negative reinforcement and positive punishment both involve something unpleasant, and only the direction of the behaviour separates them.
Glossary

Key terms

Operant
A voluntary behaviour that acts on the environment to generate consequences, as distinct from a reflex elicited by a stimulus. The term marks the difference between behaviour that is emitted and behaviour that is drawn out.
Variable ratio schedule
Reinforcement delivered after an average number of responses rather than a set number. It produces the fastest and most persistent responding of the four schedules, because no moment is ever provably a bad time to respond.
Shaping
Building a behaviour that does not yet occur by reinforcing successive approximations to it, alternating reinforcement with withholding so that variability produces something closer to reinforce next.
Vicarious reinforcement
Reinforcement that reaches a learner by observing a consequence delivered to someone else. It changes how much of a learned behaviour an observer displays without changing how much they learned.
Latent learning
Learning that has occurred but is not expressed in behaviour until there is a reason to use it. It is the finding that forces a distinction between what has been acquired and what is being performed.
Cognitive map
A stored internal representation of the layout of an environment, rather than a sequence of reinforced turns. It is inferred from behaviour when a trained route is blocked and the learner heads towards the goal's location instead.
FAQ

Operant Conditioning and Social Cognitive Learning FAQ

Why is negative reinforcement not a kind of punishment?

Because reinforcement is defined by the behaviour becoming more likely, and punishment by it becoming less likely. Negative refers only to something being taken away. A behaviour that switches off something unpleasant becomes more frequent, so it is being reinforced even though the stimulus involved is aversive.

Which schedule of reinforcement is best?

The question has no answer without a goal, which is what a well-set item is testing. Continuous reinforcement establishes a behaviour fastest, so it is best while something is being acquired. Partial schedules produce behaviour that survives long stretches without reward, so they are best once the behaviour exists and has to persist.

Why does inconsistent punishment make a behaviour worse?

Because the occasions that go unpunished still deliver whatever the behaviour was for. That puts the behaviour's own reinforcer on a partial schedule, and partial schedules produce the behaviour most resistant to extinction. Inconsistency does not weaken the punishment so much as strengthen the habit.

What did the incentive condition in the observational learning study show?

That the difference between the groups was in performance rather than in learning. Offered a reward for demonstrating what they had seen, children who had watched a model being punished matched the other groups, which means they had acquired the behaviour and were choosing not to display it.

Do I need to know the four processes of observational learning?

Not from this module. It teaches observational learning through the incentive condition and the two vicarious terms, and does not cover an attention, retention, reproduction and motivation account. Presenting that framework as subject content would go beyond what is taught here.

Study strategy

Assessment move

The grid is the whole chapter, so build it first and drill it in the right order: write up or down on one line and added or removed on the next, then read the label off. Under time pressure the two questions bleed into each other, and each of the four labels has a plausible wrong twin produced by answering them together.

For the schedules, learn the two binary choices rather than four separate definitions, and attach one everyday case to each so an unfamiliar scenario has somewhere to land. Punishment is best revised as a diagnostic: given a case where it failed, name which of the three conditions was violated and what the learner was still getting.

Finally, practise stating why the blocked-route study is evidence at all, since the mark there is for the design rather than for the result: the route was blocked so that two accounts would predict different behaviour.

Working through Operant Conditioning and Social Cognitive Learning in PSYC10006? Sia is AskSia’s AI Psychology tutor — ask any PSYC10006 Operant Conditioning and Social Cognitive Learning question and get a clear, step-by-step explanation grounded in how PSYC10006 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 74 of your The University of Melbourne subjects - and 1,000+ Bibles across every Australian university.
Sia - your PSYC10006 tutor, unlimited, worked the way the exam marks it
The full 8-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works
Unlock the full PSYC10006 Bible + 74 The University of Melbourne subjects
$0.99 Trial