The University of Hong Kong · FACULTY OF EDUCATION

MEDD8001 Chap.10 Mixed Methods, Validity and Substantive Significance

- one subject, every graph, every model, every mark
10 Chapters5-page Bible
Our own words - no uploaded lecturer files
Updated for this semester
Chapter 10 of 11 · MEDD8001

Mixed Methods, Validity and Substantive Significance

A chapter about judgement rather than new content

The final lecture returns to designs already taught, adds two that combine or extend them, and then spends its weight on two questions a reader of research has to be able to ask: is this finding an accident, and if it is not, is it big enough to matter.

Those are different questions with different answers, and conflating them is the most consequential error the course warns against.

Everything in the chapter serves one purpose, which is producing the critical mindset the lecture is named after.

Five reasons to mix, and they are not interchangeable

Mixing means running numeric and non-numeric work together so that each strand contributes something the other cannot. Five purposes are named.

Triangulation checks one finding by reaching it along more than one route, whether through separate procedures, separate sources or separate observers. Complementarity lets the second strand spell out, illustrate, sharpen or qualify what the first strand found.

Development takes what one approach produced and uses it to build the other.

Initiation treats a contradiction between the two approaches as grounds for reframing the theory. Expansion widens what a study reaches by putting a different method on each of its components.

Triangulation is the one students name most and the hardest to deliver, since it requires the routes to be genuinely independent; initiation is the most interesting and the least chosen, because it treats disagreement between strands as the finding.

Three timings of the same two strands

A concurrent design collects the two kinds of data separately but at about the same time and integrates the inferences.

A parallel design collects and analyses separately and makes inferences separately. A sequential design lets the analysis of the earlier phase influence the method choices of the next.

Each buys something and charges for it: concurrent is fastest and gives least opportunity for one strand to inform the other; parallel protects each strand and gives up the integrated account; sequential is the only one that lets the first phase change the second and is therefore the only one that cannot be compressed if the first phase runs late.

Longitudinal designs sit alongside, and only they can distinguish change over time from a difference between cohorts.

Five families of validity threat

The lecture consolidates the validity vocabulary into five questions. Is the hypothesis itself sound? Is the construct being measured properly? Is causality actually being shown? Do the findings reach beyond this sample? Is the statistical work sound?

Those five questions are, in order, the conceptual one, the one about constructs, the internal one, the external one, and the one about statistical conclusions.

Four of the five appeared in the experimental chapter; the first is the addition and it is the one no analysis can repair, because a study testing a poorly formed hypothesis can be flawless in every other respect and still produce nothing.

What a p-value carries, and what an effect size adds

Statistical significance is described as the likelihood that a gap between two groups is an accident of who happened to be sampled.

Draw two samples out of one population and they will differ a little regardless, so the question is whether the gap in front of you exceeds that ordinary wobble; the p-value answers it by giving how often a gap at least this wide would appear if the populations behind the samples were identical, and anything under five per cent is conventionally called significant.

The trouble is that how wide the gap is and how many people were measured both feed that one figure, so thousands of cases can make a negligible gap significant while a handful can leave a wide one undetected.

Substantive importance is a question about magnitude, answered by an effect size carrying its likely margin for error, and the payoff is that you stop asking whether a thing works and start asking how well, and where.

In this chapter

What this chapter covers

  • 01

    Five purposes: confirm, elaborate, build, reframe, widen

  • 02

    Why triangulation is the hardest of the five to deliver

  • 03

    Concurrent, parallel and sequential timings

  • 04

    What each timing costs when a deadline is fixed

  • 05

    Longitudinal designs and the cohort problem

  • 06

    Five questions that consolidate validity

  • 07

    Conceptual validity: the one no analysis repairs

  • 08

    What a p-value is, and the convention around it

  • 09

    Why the same number carries effect and sample size together

  • 10

    Effect size, confidence interval and proportion of variance explained

Worked example · free

Two studies, opposite conclusions, neither of them supported

Q [9 marks]. AskSia-authored practice. Two studies report on the same reading programme. Study A has 2,400 students, reports a statistically significant gain and calls the programme effective. Study B has 90 students, reports a gain that is not statistically significant, and says the programme shows no effect. Study A's reported gain is about a fifth of the size of Study B's. Which study supports adopting the programme? The marks shown are an AskSia study allocation and are not the University's marking scheme.
  • 3Say what each study has actually established.
  • 3Explain how sample size produced the opposite verdicts.
  • 3Name the reporting that would settle the question.
Neither, on what has been reported, because both have answered a question about sampling accident and neither has answered a question about magnitude. A p-value is fed both by how wide the gap is and by how many people were measured, so a very large sample can make a tiny effect significant and a small sample can leave a large effect indistinguishable from chance. Study A has probably demonstrated a real but small difference; Study B may have a substantial one its sample cannot resolve. What would settle it is the effect size in each study with its confidence interval, which converts does it work into how well does it work and under what conditions. Until that appears, effective in Study A and no effect in Study B are both overclaims.
Sia tip — Whenever you read a significance verdict, ask what the sample size was before deciding what the verdict means. Very large samples make significance cheap and very small ones make it expensive, and the word significant hides both.
Glossary

Key terms

Triangulation
Checking a finding by reaching it along more than one route, whether separate procedures, separate sources or separate observers, which only works if the routes are genuinely independent.
Sequential Design
A mixed-methods design in which the analysis of the earlier phase influences the method choices of the later one, and therefore the only timing that cannot be compressed.
Conceptual Validity
Whether the hypothesis under test is itself sound, which is the threat family no amount of careful analysis can repair after the fact.
Statistical Significance
The likelihood that a gap seen between groups is an accident of who was sampled, given as how often a gap at least that wide would appear if the underlying populations were identical.
Effect Size
A measure of the magnitude of a difference between groups, reported so that size is not confounded with sample size.
Confidence Interval
The stated range of uncertainty around an estimate, which is what turns a point value into a claim a reader can evaluate.
FAQ

Mixed Methods, Validity and Substantive Significance FAQ

What does a p-value below five per cent actually tell me?

That a gap at least this wide would rarely appear from sampling accident alone if the two populations were in fact identical. That is a statement about accident, not about importance. Because how wide the gap is and how many people were measured both feed one figure, the same value can come from a wide gap in a handful of cases or a negligible one in thousands, and the figure will not tell you which.

Reporting an effect size with its margin for error is the fix the course recommends.

When should I use mixed methods rather than one design?

When you can state in one sentence what the second strand does for the first, using one of the five named purposes. A study that adds interviews because the numbers looked thin, or a survey because interviews felt anecdotal, has hedged rather than mixed.

Given a single-semester timetable, a sequential design is the riskiest choice because it cannot be compressed if the first phase runs late, which is why most student mixed-methods studies should be concurrent.

Why is a longitudinal design so often the honest answer?

Because a great many education questions are secretly about change: whether a gap widens, whether an effect persists, whether a benefit decays. Only observing the same cases more than once makes change visible as change.

A cross-sectional study of two year groups can describe a difference between them and cannot distinguish a change over time from a difference between cohorts, so if your question contains a change verb and your design is cross-sectional, the write-up has to say so.

Study strategy

Assessment move

For the next five results you read, write two numbers in the margin: the sample size and the effect size. If the second is missing, write missing rather than moving on. Doing that across a handful of papers rebuilds your instinct faster than any amount of reading about significance, and it gives you a concrete sentence for your own analysis plan about which effect size you will report and why.

Working through Mixed Methods, Validity and Substantive Significance in MEDD8001? Sia is AskSia’s AI Education tutor — ask any MEDD8001 Mixed Methods, Validity and Substantive Significance question and get a clear, step-by-step explanation grounded in how MEDD8001 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.

A+Everything unlocked
Unlocks this Bible + all 8 of your The University of Hong Kong subjects - and 1,000+ Bibles across every Australian university.
Sia - your MEDD8001 tutor, unlimited, worked the way the exam marks it
The full 5-page Bible + practice bank with worked solutions
Chrome extension - sync your LMS so Sia knows your deadlines
Bilingual EN / Chinese on every Bible and every Sia answer
$0.99 Trial
30-day money-back · cancel in one tap · how it works