Measurement and the Basis of Validity
One slide carries the whole of this week's treatment of measurement: a definition, a dependency, an error claim and three scales. This page takes each of the four as far as it will go and marks the point where you will need a second source.
The entire treatment, on one slide
Measurement gets a single slide in the week that puts measurement at the centre of the approach. It is short enough to reproduce in full, and reproducing it in full is the fairest way to show what you are working with.
Four things are done here: measurement is defined, a dependency between reliability and validity is asserted, a relationship between error and validity is asserted, and three scales are listed. Each is taken in turn below. The counts — three scales, three nominal codes, two ordinal examples, three ratio examples — are derived from the lists; the slide states none of them, and its own hedge is the word "include", which declares the list partial without saying what is missing.
Reliability before validity: two halves of one relationship
The key-terms slide defines reliability as "consistency or stability of test scores" and validity as "the instrument measures what it claims to measure", and closes with a parenthesis: "(Reliability in itself does not ensure validity)". The measurement slide then adds a claim the key-terms slide does not make: "Must be reliable in order to claim validity of findings."
THE TWO STATEMENTS, AND WHAT EACH ONE ALONE LEAVES OUT
| Statement | Where | What it establishes | What a reader of that slide alone misses |
|---|---|---|---|
| "Reliability in itself does not ensure validity" | Key-terms slide | Reliability is not sufficient for validity — a consistently wrong instrument is still wrong | That reliability is nonetheless required |
| "Must be reliable in order to claim validity of findings" | Measurement slide | Reliability is necessary for validity — an unstable instrument cannot support a validity claim | That reliability alone does not get you there |
The two are compatible: reliability is necessary and not sufficient. The source never says so in one place, and neither slide refers to the other.
Put together, they give you the one genuinely operational idea in the week's measurement material, and it has a direct consequence for how you sequence your instrument work. A reliability problem is prior: until your instrument gives the same answer on the same case twice, no argument about whether it measures the right thing can even be started. A validity problem is subsequent and is not fixed by tightening consistency.
The error claim, and what the week's own reading says instead
The third line on the slide reads: "The more errors there are in the data the less valid are claims or generalisations." As written that is a monotonic law — validity falls as error count rises — stated without qualification, without a threshold and without any means of measuring either quantity.
The article the same week sets as required reading takes the opposite view of the same relationship, and the difference is not a nuance. It treats poor measurement not as something that destroys validity but as something that costs you subjects. That article is external material: it is from a sport and exercise science web journal, published in 2000, and its subjects are athletes and physiological measures.
Three scales, no definitions, and one wrong example
The slide's last item lists scales of measurement. It is the material another week of this subject points at when it defers its own measurement content elsewhere — the two decks share the nominal coding example word for word — so this list is doing more work in the subject than its four lines suggest.
THE SCALES AS THE SLIDE GIVES THEM
| Scale | What the slide supplies | What it does not |
|---|---|---|
| Nominal | A coding example only: 1 = English, 2 = French, 3 = Chinese | Any statement that the numbers do not order or measure anything — which is the whole point of the category |
| Ordinal | A three-word gloss, "ranking of order", and two examples: frequency, grades achieved | A definition, and any warning about what may not be done with ranked codes |
| Ratio | Three instances: weight, height, speed | The property that defines it — a true zero — which is never mentioned anywhere in the subject |
| Interval | Absent from this list entirely | It appears in another week's list, which in turn omits ratio |
The scales are source examples, not a classification you can apply: no criterion separates any two of the categories. Between the two lists the subject holds, all four conventional levels are named, and no single place in the supplied material lists all four.
The practical effect is that a reader finishes this week able to recognise the words nominal, ordinal and ratio, and unable to assign a new variable to one of them with any confidence. If your instrument needs a measurement level stated for each item — and any instrument that will be analysed statistically does — take the scheme from Levels of Measurement in Structured Questions or from a methods text, and cite it where you took it from.
Writing a measurement section you can defend
The sequence below is built on what the source supports and is not a procedure the source states. Each step names what the material gives you and what you must supply yourself.
Six steps from a concept to a defensible measure
Write down the thing you are trying to capture
The week uses "concept" four times and defines it none, so the step from concept to measure is one the material never performs. Perform it explicitly and in writing.
State the assignment rule
What symbol or number gets attached to what observation, and by whom. This is the source's own definition of measurement; supplying the rule is your work, not its.
Give every item a measurement level, and name the scheme
Say which of the two classifications in this week you used, or cite an external one. Do not mix them silently.
Establish stability before you argue relevance
The dependency runs one way. Until the instrument returns the same answer on the same case, the validity argument has nothing to attach to.
State what you are claiming validity of
The source's definition is about the instrument measuring what it claims to measure. Write down the claim, so a reader can see what would falsify it.
Record what your measurement error does to your conclusions
State the position you are taking on error and attribute it. Do not attribute the article's sample-size compensation to the teaching material, or the teaching material's monotonic claim to the article.
Before your measurement section goes to a supervisor
- Every measure has a stated assignment rule, not just a name
- Every item has a measurement level, and the classification scheme is cited
- The reliability claim comes before the validity claim and is separately evidenced
- No figure from the week's set reading appears without its discipline attached
- Where the material contradicts itself, you have quoted both statements rather than choosing one
The objective this material sits under
The week's second stated objective, quoted exactly and without punctuation as the source prints it, is "Perceive reliability and validity generalization and replication". It admits at least three readings — four items, three items, or two paired items — and the source does not resolve it. The audit records the objective as not delivered, and the coverage table below is why.
WHAT THE WEEK SUPPLIES AGAINST ITS OWN SECOND OBJECTIVE
| Term | Coverage across the whole week |
|---|---|
| Reliability | One definition of eight words, plus two uses in the dependency claim. No procedure, no measure, no test, no threshold. |
| Validity | One definition of eight words, plus three further uses. No procedure, no type of validity named, no threat named. |
| Generalisation | Two occurrences: the objective, and "generalisations" on the measurement slide. Never defined. The verb form appears only in borrowed text. |
| Replication | One occurrence in the entire week — inside the objective itself. Zero across the slides beyond that line and zero in the notes. |
What to carry forward
- Measurement is a rule-governed assignment of symbols and numbers. The source says so and supplies no rule.
- Reliability is necessary for validity and not sufficient for it. The source states each half on a different slide.
- The slide's error claim and the set reading's compensation argument are incompatible as written. Attribute whichever you use.
- Three scales are listed, none defined, one ordinal example misfiled, and interval missing. No single place in the material lists all four levels.
- Numbers connecting measurement quality to sample size come from a sport and exercise science article and were deleted from the teaching notes. They are not standards.
Frequently asked questions
Does the source say reliability is required for validity, or only that it is not enough?
Both, on two different slides. The key-terms slide closes with "Reliability in itself does not ensure validity"; the measurement slide states "Must be reliable in order to claim validity of findings". The two are compatible — reliability is necessary and not sufficient — but the source never says that in one place, so quote both if you are asked to state its position.
How many levels of measurement does this material teach?
Three, and a different three depending which week you read. This week lists nominal, ordinal and ratio; another week lists nominal, ordinal and interval. All four conventional levels are named across the subject and no single place lists them all. The property that separates ratio from interval, a true zero, is never mentioned anywhere.
Is "frequency" really an ordinal variable?
Not on the slide's own terms. A frequency is a count with a true zero and equal intervals, which places it in the ratio category the same slide lists two lines below. This library reports the example as printed rather than correcting it, and notes that the slide gives no definition of either category that would let a reader catch the error.
Can I use the sample sizes the set reading gives for poor validity?
Not as a standard. Those figures — thousands rather than hundreds for a descriptive study with poorly valid measures, twenty per group for a controlled trial with a highly reliable one — are one author's statements about sport and exercise science in 2000. The teaching material deleted that passage entirely and states no equivalent figure in its own voice.
What does the source give me for testing an instrument before I use it?
Nothing procedural. No pilot, retest, split-half or inter-rater check appears in the week, and the passage on pilot studies in its set reading is one of the five the study notes deleted. Plan the pilot anyway, and cite the source you took its design from.
References and source attribution
- Hopkins 2000, 'Quantitative Research Design', a sport-science web journal, vol. 4, issue 1 — journal title, address and author given names scrubbed. The week's set reading, and the source of the deleted passage connecting validity and reliability to sample size.
- Bryman, A. 2016, Social Research Methods, 5th ed., Oxford University Press, Oxford.
- O'Leary, Z. 2017, The Essential Guide to Doing Your Research Project, 3rd ed., Sage Publications, London.
- Veal, A. J. 2005, Business Research Methods: A Managerial Approach, Longman.
- The supplied teaching source: the slide deck on the nature of quantitative research — in particular its key-terms and measurement slides — and both copies of the accompanying study notes.
Suggested questions for Ask KEVOS
- Write the measurement subsection of my methodology chapter using only what this material supports.
- Give every item in my instrument a measurement level and tell me which scheme you used.
- How do I evidence reliability before I make a validity claim about my instrument?
- What exactly does this week say about validity, and what do I have to source elsewhere?
- Explain the difference between measurement error and sampling error, since the source never separates them.
