The University of Hong Kong · FACULTY OF EDUCATION

MEDD6128 Chap.6 Assessment and Technology in Curriculum Design

- one subject, every graph, every model, every mark
10 Chapters5-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 6 of 7 · MEDD6128

Assessment and Technology in Curriculum Design

Two purposes pulling in different directions

Reformers almost always want assessment in their plans and frequently treat it as the main instrument of reform, which is why this session places it inside design rather than after it. The difficulty named at the outset is that any single act of assessment serves more than one master at the same time, serving the grade and serving the learning together.

That tension is how teachers come to value innovative assessment ideas while doing something far more limited in practice, and it is why an assessment decision cannot be delegated to whoever writes the reporting policy.

National arrangements encode a prior judgement about trust

The differences between assessment systems are significant and deep rooted.

One system pairs a settled suspicion of teacher judgement with a proliferation of formal testing alongside some school based work. Another gives formative work to teachers and handles all summative assessment externally. A third relies on national tests while trusting teachers to make summative judgements. A fourth has increased testing for accountability with rigid pacing toward standards.

Read any proposal to import a practice with this in mind, because the practice travels and the trust arrangement behind it does not.

Why the formative and summative split will not hold cleanly

On one influential account all assessment begins with a summative act, because a judgement is made, and formative assessment is that judgement plus feedback which the learner then uses.

If that is right, the label attaches to what happens after the judgement rather than to the instrument, and a department that renames its tests has changed nothing.

Two findings explain why reform in this area stalls: formative practice as written down and formative practice as meant come apart, with many teachers adopting it superficially rather than integrating it into planning; and where a summative instrument carries heavy consequences, its backward pressure squeezes out the room in which formative work could happen.

What the evidence base for these technologies actually contains

A systematic review of artificial intelligence applied to student assessment at primary and secondary level found nine original studies meeting its criteria across more than a decade, covering a few hundred participants in total.

Its reported contributions are specific: predicting performance, automating evaluation through neural networks or natural language processing, using educational robots to analyse the learning process, and detecting factors that make classes more engaging.

That is a real evidence base and a thin one, so a framework resting a system wide claim on it has overloaded the citation.

Four lenses on whether these systems reduce inequity or widen it

The central reading declines to answer in the abstract and supplies four lenses pointed at one socio technical system: the surrounding system, the data a model learned from, the algorithms themselves, and what happens when automated and human decisions interact.

Each makes a different repair available, so naming the wrong lens produces a remedy that cannot work. Disparities of access do not end when every student has a device, and the summarising finding is worth keeping beside any proposal: simply improving access is not enough.

In this chapter

What this chapter covers

  • 01

    Grading and learning as competing purposes of one act

  • 02

    How national assessment arrangements encode trust in teachers

  • 03

    The claim that formative assessment is a judgement plus usable feedback

  • 04

    Letter and spirit, and why formative reform stalls

  • 05

    Wash back from high stakes instruments

  • 06

    Benchmarking and value added assessment as accountability instruments

  • 07

    What a systematic review of assessment and artificial intelligence found

  • 08

    What generative tools changed, stated precisely

  • 09

    Four equity lenses, and the repair each one makes available

  • 10

    Social distance between developers and the people served

Worked example · free

Two objections to an automatically scored extended response

Q [9 marks]. AskSia authored practice. A department replaces its end of unit written report with an automatically scored extended response task, arguing that the tool marks faster, more consistently and without the variation between teachers that students complain about. What are the two strongest objections available in this course, and which is harder to answer? The marks shown are an AskSia study allocation and are not the University's marking scheme.
  • 3State the objection that concerns what the model learned from.
  • 3State the objection that concerns what the instrument can reach.
  • 3Say which is harder to answer and why.
The first objection is about the training data. Where a scoring model learns from human judgements it scales up whatever those judgements contained, including assessment biases against the written work of students from marginalised groups, so consistency here is not the absence of bias but its stabilisation, and the department has removed the variation that made it visible. The second objection is about the construct: automatic scoring struggles with the macrostructure and causal structure of an explanation, which in an extended response is usually the thing being assessed. The second is harder, because the first can in principle be addressed by auditing and reweighting training data while the second says the instrument does not reach the construct at all.
Sia tip — Decide what an instrument can reach before deciding whether it is fair. A valid measure of the wrong construct cannot be repaired with better data.
Glossary

Key terms

Wash Back
The effect a high stakes summative instrument has on teaching and learning before it is taken, whose backward pressure squeezes out much of the room formative work needs.
Benchmarking
Setting one school's measured performance beside the results of schools that resemble it closely enough to make the comparison mean something, usually attached to published comparison.
Value Added Assessment
Summative assessment in which unadjusted results are corrected for who the school actually enrolled, so that comparisons account for who was enrolled.
System Lens
The view of equity that examines the surrounding socio technical arrangement, covering access, adoption and who was represented in the design process.
Data Lens
The view of equity that examines the records a model learned from, asking whose learners and contexts were overrepresented and what historical inequity the data encodes.
Model Misspecification
The situation in which a statistical model of learning has the wrong form for the process it represents, which can leave slower learners with less practice than they need.
Product Problem
The long standing teaching question of whether a student produced the right answer, distinguished from the process problem of knowing whether learning occurred.
Social Distance
The gap between the people who design an educational technology and the people it is meant to serve, identified as a contributor to inequity even when intentions are good.
FAQ

Assessment and Technology in Curriculum Design FAQ

Is a low stakes quiz automatically formative?

No, and assuming so is the commonest design error in this area. The label attaches to whether a judgement is followed by feedback the learner actually uses, not to the frequency or the stakes of the task.

A fortnightly quiz whose results are averaged into a final grade is summative assessment at fortnightly intervals, and attaching a consequence to it also introduces the wash back that limits formative opportunity elsewhere in the course.

Can a school make its assessment formative by renaming its tasks?

Not on the account this session teaches. If formative assessment is a judgement plus feedback the learner uses, then building it requires producing usable feedback and providing the time in which a student can act on it, which makes it a timetable decision as much as an assessment one.

A department that changes only its vocabulary has changed its documents, and the gap between formative work as written down and formative work as meant is exactly what the research on stalled reform describes.

What did generative tools actually change about assessment?

The reading is careful rather than alarmed. It notes that claims about a new technology solving education are not new and have repeatedly failed, and locates what is genuinely different in the capacity to turn out original prose that no reader can reliably tell from a person's.

Its conclusion is limited: these tools are unlikely to destroy education but may destroy the legitimacy of some long held practices, with exposure concentrated where a course relies on a narrow set of task types completed away from an instructor.

How strong is the evidence for artificial intelligence improving assessment?

Thinner than the volume of discussion suggests, which is the reason to state the numbers. A systematic review covering more than a decade found nine original studies meeting its criteria, with a few hundred participants between them, and reported contributions in performance prediction, automated evaluation, robot supported analysis of learning and the detection of factors that engage students.

Citing it for a system wide claim overloads it, whereas citing it for a specific and bounded one is entirely defensible.

Does giving every student a device solve the access problem?

It solves part of one lens and leaves the rest untouched. Households with nominal access can still be limited by slow connections, shared machines, mobile only access and capped data; guidance in a single language limits who can use a system; and better resourced schools tend to use the same technologies in more inventive ways so that equal provision can still widen a gap.

Where a disparity persists among students with equivalent access, the finding implicates the algorithmic lens and no amount of hardware repairs it.

Study strategy

Assessment move

Take one assessment you set or sat and trace what it actually rewards, then compare that with what the surrounding curriculum says it values. Where they diverge, the assessment wins, and that is the claim this chapter exists to establish. Then ask the harder question about any tool you would propose in its place: not whether the tool is fair, but whether the thing it measures is the thing you meant to measure.

Working through Assessment and Technology in Curriculum Design in MEDD6128? Sia is AskSia’s AI Education tutor — ask any MEDD6128 Assessment and Technology in Curriculum Design question and get a clear, step-by-step explanation grounded in how MEDD6128 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 8 of your The University of Hong Kong subjects - and 1,000+ Bibles across every Australian university.
Sia - your MEDD6128 tutor, unlimited, worked the way the exam marks it
The full 5-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works