PSYC10006 Chap.2 Operant Conditioning and Social Cognitive Learning
Operant Conditioning and Social Cognitive Learning
Classical conditioning explains reflexes and has nothing to say about behaviour that is emitted rather than elicited. Operant conditioning fills the gap: voluntary behaviour is shaped by what follows it. Every consequence takes two labels, and they have to be assigned in a fixed order because answering the second question first is the most reliable way to get the label wrong.
Did the behaviour become more likely or less likely? Was something added or removed? The module then moves through the four schedules and the response pattern each produces, shaping by successive approximation, and three conditions that govern whether punishment works at all.
It closes with two lines of evidence that end the behaviourist restriction to observable events: a stored spatial map, learning that stays invisible until there is a reason to use it, and consequences that reach a learner by observation rather than by experience.
What this chapter covers
- 01
Why voluntary behaviour needs a second account, and what an operant is
- 02
The two questions, in order: did behaviour rise or fall, and was something added or removed
- 03
Why negative means subtraction, and why negative reinforcement raises behaviour
- 04
Ratio against interval, fixed against variable, and the four schedules they generate
- 05
The pattern each schedule produces, and why ratio schedules run steeper
- 06
Shaping: reinforce the closest approximation, withhold, then move the criterion
- 07
The partial reinforcement effect, and why persistence is engineered separately
- 08
Contingency, contiguity and consistency as conditions on effective punishment
- 09
Why inconsistent punishment can leave a behaviour stronger than none would
- 10
Four costs of punishment, and three constructive alternatives to it
- 11
The blocked-route study, and why its design lets two accounts predict differently
- 12
Latent learning, and learning separated from performance by an incentive condition
Design a shaping programme and state the schedule at each stage
- 1Define the target behaviour concretely: sits at the shared table for the length of one meal. An abstract target cannot be approximated, because you cannot tell whether a given behaviour is closer to it.
- 1Find the closest current approximation, whatever the client already does that points at the target, and reinforce it on a continuous schedule, because continuous reinforcement is what establishes a behaviour fastest.
- 1Withhold once the approximation is reliable. Behaviour becomes more variable when reinforcement stops, and that variability produces the slightly closer behaviour you will reinforce next.
- 1Move the criterion and reinforce the new approximation continuously, repeating the cycle. Moving the criterion before the current step is reliable extinguishes the approximation and loses ground.
- 1Switch schedules once the target behaviour exists. Move from continuous to a partial schedule so the behaviour resists extinction, and say explicitly that this is a different job from shaping.
Key terms
- Operant
- A voluntary behaviour that acts on the environment to generate consequences, as distinct from a reflex elicited by a stimulus. The term marks the difference between behaviour that is emitted and behaviour that is drawn out.
- Variable ratio schedule
- Reinforcement delivered after an average number of responses rather than a set number. It produces the fastest and most persistent responding of the four schedules, because no moment is ever provably a bad time to respond.
- Shaping
- Building a behaviour that does not yet occur by reinforcing successive approximations to it, alternating reinforcement with withholding so that variability produces something closer to reinforce next.
- Vicarious reinforcement
- Reinforcement that reaches a learner by observing a consequence delivered to someone else. It changes how much of a learned behaviour an observer displays without changing how much they learned.
- Latent learning
- Learning that has occurred but is not expressed in behaviour until there is a reason to use it. It is the finding that forces a distinction between what has been acquired and what is being performed.
- Cognitive map
- A stored internal representation of the layout of an environment, rather than a sequence of reinforced turns. It is inferred from behaviour when a trained route is blocked and the learner heads towards the goal's location instead.
Operant Conditioning and Social Cognitive Learning FAQ
Why is negative reinforcement not a kind of punishment?
Because reinforcement is defined by the behaviour becoming more likely, and punishment by it becoming less likely. Negative refers only to something being taken away. A behaviour that switches off something unpleasant becomes more frequent, so it is being reinforced even though the stimulus involved is aversive.
Which schedule of reinforcement is best?
The question has no answer without a goal, which is what a well-set item is testing. Continuous reinforcement establishes a behaviour fastest, so it is best while something is being acquired. Partial schedules produce behaviour that survives long stretches without reward, so they are best once the behaviour exists and has to persist.
Why does inconsistent punishment make a behaviour worse?
Because the occasions that go unpunished still deliver whatever the behaviour was for. That puts the behaviour's own reinforcer on a partial schedule, and partial schedules produce the behaviour most resistant to extinction. Inconsistency does not weaken the punishment so much as strengthen the habit.
What did the incentive condition in the observational learning study show?
That the difference between the groups was in performance rather than in learning. Offered a reward for demonstrating what they had seen, children who had watched a model being punished matched the other groups, which means they had acquired the behaviour and were choosing not to display it.
Do I need to know the four processes of observational learning?
Not from this module. It teaches observational learning through the incentive condition and the two vicarious terms, and does not cover an attention, retention, reproduction and motivation account. Presenting that framework as subject content would go beyond what is taught here.
Assessment move
The grid is the whole chapter, so build it first and drill it in the right order: write up or down on one line and added or removed on the next, then read the label off. Under time pressure the two questions bleed into each other, and each of the four labels has a plausible wrong twin produced by answering them together.
For the schedules, learn the two binary choices rather than four separate definitions, and attach one everyday case to each so an unfamiliar scenario has somewhere to land. Punishment is best revised as a diagnostic: given a case where it failed, name which of the three conditions was violated and what the learner was still getting.
Finally, practise stating why the blocked-route study is evidence at all, since the mark there is for the design rather than for the result: the route was blocked so that two accounts would predict different behaviour.
Working through Operant Conditioning and Social Cognitive Learning in PSYC10006? Sia is AskSia’s AI Psychology tutor — ask any PSYC10006 Operant Conditioning and Social Cognitive Learning question and get a clear, step-by-step explanation grounded in how PSYC10006 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.