PSY2042 Chap.3 Measurement, Evidence and Score Interpretation
Measurement, Evidence and Score Interpretation
Personality measurement begins with a construct definition and an operational link to observable indicators. Standard administration, scoring and interpretation distinguish formal assessment from informal impression. Reliability concerns precision across relevant items, occasions, forms or raters. Validity concerns whether evidence supports a particular interpretation and use.
A consistent measure can still assess the wrong construct, so the two ideas must not be collapsed.
Evidence can concern content coverage, concurrent or predictive criteria, convergence, discrimination, factor structure and incremental prediction. Self reports, informants, observation, experience sampling and digital traces sample different information and contexts.
Disagreement can reveal error or perspective-specific expression. Scores also depend on norms, response processes and fairness. There is no perfect method; the defensible choice fits the construct, population, purpose and consequence of error.
The audit runs from construct definition through indicators, scoring, reliability, validity, norms and justified use.
Each reliability form targets a different source of inconsistency, and each validity source supports a bounded inference. Multimethod disagreement is analysed through visibility, opportunity, reference points and bias before it is labelled error. Norm-referenced scores are interpreted relative to a named comparison group and measurement uncertainty, never as percentages of a trait.
Newer digital and experience-sampling methods add temporal detail but also require transparent operationalisation, external validation, fairness checks and governance proportionate to the consequences of a decision.
What this chapter covers
- 01
Construct definition and operationalisation
- 02
Internal consistency, retest, alternate form and inter-rater reliability
- 03
Content, criterion, construct and incremental validity evidence
- 04
Self report design and response styles
- 05
Informant reports and multimethod convergence
- 06
Norm referenced score interpretation
- 07
Experience sampling and behavioural observation
- 08
Digital traces, prediction, fairness and method selection
Select a measure for a high stakes decision
- +1Define the job relevant construct and check whether item content represents it rather than a convenient proxy.
- +1Verify relevant reliability in the target applicant population, not only the development sample.
- +1Examine whether supervisor ratings are an independent, reliable and appropriate criterion.
- +1Test predictive performance in held out applicants and compare it with simpler existing information.
- +1Check norms, subgroup functioning, accessibility and fairness because the use is consequential.
- +1Inspect response style and incentives created by a high stakes self report setting.
- +1Conclude that the current evidence is promising but insufficient for operational selection without external and fairness validation.
Key terms
- Formal assessment
- A systematic process with standard administration, scoring and interpretation used to describe or predict functioning.
- Internal consistency
- Reliability evidence concerning the coherence of item responses intended to contribute to a score.
- Test retest reliability
- Reliability evidence concerning score stability across an interval when the construct should remain stable.
- Content evidence
- Evidence that indicators adequately represent the defined construct domain.
- Convergent evidence
- Evidence that scores relate as expected to other measures of similar constructs.
- Incremental validity
- Evidence that a score improves prediction beyond information already available.
- Reference group
- The population distribution used to interpret a person’s relative score standing.
- Experience sampling
- Repeated assessment of states, behaviour or context close to daily-life occurrence.
Measurement, Evidence and Score Interpretation FAQ
Can a reliable measure still be invalid?
Yes. Scores can be consistent while the indicators omit the target construct or represent something else. Reliability limits random imprecision, while validity evaluates the interpretation and purpose supported by the full evidence network.
What is the difference between convergent and criterion evidence?
Convergent evidence compares the score with another measure of a related construct. Criterion evidence relates it to a relevant outcome or benchmark. The comparison target, not the size of the association alone, determines the label.
Why might self and informant reports disagree?
They may observe different contexts and information, use different reference points or contain different biases. Map opportunity and visibility, check score precision and seek behavioural evidence before deciding whether disagreement is error or meaningful specificity.
Are digital traces objective personality measures?
They can be recorded automatically, but their psychological meaning still requires operational and validity evidence. A trace can predict in one platform while reflecting device, demographic or contextual artifacts instead of the intended construct.
Exam move
Create an evidence matrix with measurement claims down the side and methods across the top. For each example, state the relation before naming the technical concept: same items, later occasion, different raters, related construct, external criterion or added prediction. Practise auditing from construct to indicators, scoring, reliability, validity, norms and use.
Keep percentile rank separate from percentage and interpret any standard score relative to its reference group and uncertainty. End high stakes cases with a decision and the evidence gap most likely to change it, rather than saying only that more research is needed.
Working through Measurement, Evidence and Score Interpretation in PSY2042? Sia is AskSia’s AI Psychology tutor — ask any PSY2042 Measurement, Evidence and Score Interpretation question and get a clear, step-by-step explanation grounded in how PSY2042 is taught and assessed. Read this chapter free, then take your hardest questions to Sia.