KEVOS
ArticlesServicesCase studiesAboutContact
ArticlesServicesCase studiesAboutContact
← ArticlesMeasurement and the Basis of ValidityProject Delivery · Research ProjectsLesson 187/216← PrevNext →
GuidePublished 16 Aug 202614 min readBy KEVOS Editorialmeasurement in researchreliability and validityscales of measurementnominal ordinal ratio
On this page

Ask about this page

KEVOS AIMeasurement and the Basis of Validity

KEVOS knowledge first · trusted web sources when needed

KEVOS/Project Delivery/Research Projects/The Nature of Quantitative Research
Project DeliveryResearch ProjectsCoreQuantitative Research

Measurement and the Basis of Validity

One slide carries the whole of this week's treatment of measurement: a definition, a dependency, an error claim and three scales. This page takes each of the four as far as it will go and marks the point where you will need a second source.

Reading time15 minutes
LevelCore
Topic streamQuantitative Research
Source materialThe Nature of Quantitative Research
Updated2026-08-16

In brief

  • Measurement is defined as an act — assigning symbols and numbers by rule — and not one rule of assignment appears anywhere in the week.
  • Reliability is made a precondition of validity on one slide and explicitly insufficient for it on another. Both statements are true together; neither appears on the other's slide.
  • The error claim is stated as a monotonic law: more error, less validity. The article this week sets as its reading takes the opposite position, and the passage saying so was deleted from the notes.
  • Three scales of measurement are listed, none is defined, and one of the two ordinal examples belongs in the ratio category by the slide's own reckoning.
  • The objective covering reliability, validity, generalisation and replication is recorded as not delivered.

The entire treatment, on one slide

Measurement gets a single slide in the week that puts measurement at the centre of the approach. It is short enough to reproduce in full, and reproducing it in full is the fairest way to show what you are working with.

From the source

The measurement slide, quoted in full and in order

"The act of measuring by assigning symbols, and numbers based on a set of rules"

"Must be reliable in order to claim validity of findings"

"The more errors there are in the data the less valid are claims or generalisations"

"Scales of measurement include:" — "Nominal e.g. 1= English, 2 = French, 3 = Chinese" · "Ordinal – ranking of order e.g. frequency, grades achieved" · "Ratio e.g. weight, height, speed"

Four things are done here: measurement is defined, a dependency between reliability and validity is asserted, a relationship between error and validity is asserted, and three scales are listed. Each is taken in turn below. The counts — three scales, three nominal codes, two ordinal examples, three ratio examples — are derived from the lists; the slide states none of them, and its own hedge is the word "include", which declares the list partial without saying what is missing.

Source gap

A definition of an act, with the rules withheld

"The act of measuring by assigning symbols, and numbers based on a set of rules" names the rules and supplies none. No rule of assignment appears anywhere in the week: not for deciding what number goes with what category, not for deciding how many categories there should be, not for deciding when a thing is measurable at all.

That is worth stating rather than glossing over, because the definition is otherwise a good one. It gets the crucial point across — measurement is a rule-governed assignment, not an observation — and then leaves the reader with no rule.

Reliability before validity: two halves of one relationship

The key-terms slide defines reliability as "consistency or stability of test scores" and validity as "the instrument measures what it claims to measure", and closes with a parenthesis: "(Reliability in itself does not ensure validity)". The measurement slide then adds a claim the key-terms slide does not make: "Must be reliable in order to claim validity of findings."

THE TWO STATEMENTS, AND WHAT EACH ONE ALONE LEAVES OUT

StatementWhereWhat it establishesWhat a reader of that slide alone misses
"Reliability in itself does not ensure validity"Key-terms slideReliability is not sufficient for validity — a consistently wrong instrument is still wrongThat reliability is nonetheless required
"Must be reliable in order to claim validity of findings"Measurement slideReliability is necessary for validity — an unstable instrument cannot support a validity claimThat reliability alone does not get you there

The two are compatible: reliability is necessary and not sufficient. The source never says so in one place, and neither slide refers to the other.

Put together, they give you the one genuinely operational idea in the week's measurement material, and it has a direct consequence for how you sequence your instrument work. A reliability problem is prior: until your instrument gives the same answer on the same case twice, no argument about whether it measures the right thing can even be started. A validity problem is subsequent and is not fixed by tightening consistency.

Practice note

What the dependency does to your instrument decisions

The sequencing below is this page's, drawn from the two statements; the source states the dependency and prescribes nothing procedural.

Deal with stability first — wording that means one thing, categories that do not overlap, a coding rule someone else could apply to your data and reach your answer. Only then argue that the stable thing you are capturing is the thing your research question is about. If a pilot shows the instrument is unstable, do not proceed to the validity argument; there is nothing yet to make it about.

Source gap

Neither term gets a procedure

Reliability is one line and validity is one line. Nothing in the week says how either is established, tested, reported or threatened; no type of validity is named; no threat to validity is named; no coefficient, retest, split-half or inter-rater check appears. The subject defines the pair in three separate places with three different scopes and never assembles them.

The phrase "test scores" in the reliability definition is also the only appearance of that phrase in the week — no test is described, administered or scored anywhere in it. For the fuller treatment held in this library, see Validity and Reliability.

The error claim, and what the week's own reading says instead

The third line on the slide reads: "The more errors there are in the data the less valid are claims or generalisations." As written that is a monotonic law — validity falls as error count rises — stated without qualification, without a threshold and without any means of measuring either quantity.

The article the same week sets as required reading takes the opposite view of the same relationship, and the difference is not a nuance. It treats poor measurement not as something that destroys validity but as something that costs you subjects. That article is external material: it is from a sport and exercise science web journal, published in 2000, and its subjects are athletes and physiological measures.

Note

The passage that connects measurement quality to sample size — deleted from the notes

The article's own words: "the worse your measurements, the more subjects you need to lift the signal (the effect) out of the noise (the errors in measurement)"; "if the validity of the main variables is poor, you may need thousands rather than hundreds of subjects"; "the more reliable a measure, the less subjects you need to see a small change".

That sub-section was cut from the study notes along with four other passages. It is the only place in the supplied material that attaches a consequence to validity or reliability, and the teaching material removed it. Every figure in it belongs to that article and that discipline — "thousands rather than hundreds", and the twenty and ten subjects it offers elsewhere for a controlled trial and a crossover, are its statements about physiological measurement, not thresholds for a project-management study. The teaching material states no such threshold anywhere.

Caution

The two positions cannot be reconciled from the source

The slide says error destroys validity. Its own set reading says error is compensable by sample size. Neither document refers to the other, and nothing in the supplied material distinguishes measurement error from sampling error — which is the distinction that would let a reader hold both.

If you write about measurement quality, state which position you are taking and whose it is. Do not present the slide's claim and the article's claim as one account, and do not use the article's numbers to give the slide's claim a precision it never had. The sizing question is handled separately at Sample Size in Quantitative Studies.

Three scales, no definitions, and one wrong example

The slide's last item lists scales of measurement. It is the material another week of this subject points at when it defers its own measurement content elsewhere — the two decks share the nominal coding example word for word — so this list is doing more work in the subject than its four lines suggest.

THE SCALES AS THE SLIDE GIVES THEM

ScaleWhat the slide suppliesWhat it does not
NominalA coding example only: 1 = English, 2 = French, 3 = ChineseAny statement that the numbers do not order or measure anything — which is the whole point of the category
OrdinalA three-word gloss, "ranking of order", and two examples: frequency, grades achievedA definition, and any warning about what may not be done with ranked codes
RatioThree instances: weight, height, speedThe property that defines it — a true zero — which is never mentioned anywhere in the subject
IntervalAbsent from this list entirelyIt appears in another week's list, which in turn omits ratio

The scales are source examples, not a classification you can apply: no criterion separates any two of the categories. Between the two lists the subject holds, all four conventional levels are named, and no single place in the supplied material lists all four.

Caution

One of the two ordinal examples is in the wrong category

"Grades achieved" is a sound ordinal example. "Frequency" is not. A frequency is a count: it has a true zero and equal intervals, which puts it in the slide's own ratio category alongside weight and speed.

The error is reported and not corrected here, and it matters more than a stray example usually would, because the slide gives no definition of either category that would let a reader catch it. The same week's study notes compound it: they define ordinal as a variable that "can't decide whether it's numeric or nominal", and a frequency is unambiguously numeric. See Population, Sample, Variables and Data for the second classification system and where it collides with this one.

The practical effect is that a reader finishes this week able to recognise the words nominal, ordinal and ratio, and unable to assign a new variable to one of them with any confidence. If your instrument needs a measurement level stated for each item — and any instrument that will be analysed statistically does — take the scheme from Levels of Measurement in Structured Questions or from a methods text, and cite it where you took it from.

Writing a measurement section you can defend

The sequence below is built on what the source supports and is not a procedure the source states. Each step names what the material gives you and what you must supply yourself.

Six steps from a concept to a defensible measure

  1. Write down the thing you are trying to capture

    The week uses "concept" four times and defines it none, so the step from concept to measure is one the material never performs. Perform it explicitly and in writing.

  2. State the assignment rule

    What symbol or number gets attached to what observation, and by whom. This is the source's own definition of measurement; supplying the rule is your work, not its.

  3. Give every item a measurement level, and name the scheme

    Say which of the two classifications in this week you used, or cite an external one. Do not mix them silently.

  4. Establish stability before you argue relevance

    The dependency runs one way. Until the instrument returns the same answer on the same case, the validity argument has nothing to attach to.

  5. State what you are claiming validity of

    The source's definition is about the instrument measuring what it claims to measure. Write down the claim, so a reader can see what would falsify it.

  6. Record what your measurement error does to your conclusions

    State the position you are taking on error and attribute it. Do not attribute the article's sample-size compensation to the teaching material, or the teaching material's monotonic claim to the article.

Before your measurement section goes to a supervisor

  • Every measure has a stated assignment rule, not just a name
  • Every item has a measurement level, and the classification scheme is cited
  • The reliability claim comes before the validity claim and is separately evidenced
  • No figure from the week's set reading appears without its discipline attached
  • Where the material contradicts itself, you have quoted both statements rather than choosing one

The objective this material sits under

The week's second stated objective, quoted exactly and without punctuation as the source prints it, is "Perceive reliability and validity generalization and replication". It admits at least three readings — four items, three items, or two paired items — and the source does not resolve it. The audit records the objective as not delivered, and the coverage table below is why.

WHAT THE WEEK SUPPLIES AGAINST ITS OWN SECOND OBJECTIVE

TermCoverage across the whole week
ReliabilityOne definition of eight words, plus two uses in the dependency claim. No procedure, no measure, no test, no threshold.
ValidityOne definition of eight words, plus three further uses. No procedure, no type of validity named, no threat named.
GeneralisationTwo occurrences: the objective, and "generalisations" on the measurement slide. Never defined. The verb form appears only in borrowed text.
ReplicationOne occurrence in the entire week — inside the objective itself. Zero across the slides beyond that line and zero in the notes.
From the source

Objective delivery, stated once

The consolidated audit across the six weeks now on record shows fourteen objectives stated, none fully delivered, four partially delivered and ten not delivered at all. This week's second objective is one of the ten. Those are the reconciled figures and the only ones this library cites; the full reckoning is at The Six-Week Objective Audit.

The practical consequence for you is narrow and worth being precise about: you can take the definitions and the dependency from this material, and you will have to take everything you do with them from somewhere else.

What to carry forward

  1. Measurement is a rule-governed assignment of symbols and numbers. The source says so and supplies no rule.
  2. Reliability is necessary for validity and not sufficient for it. The source states each half on a different slide.
  3. The slide's error claim and the set reading's compensation argument are incompatible as written. Attribute whichever you use.
  4. Three scales are listed, none defined, one ordinal example misfiled, and interval missing. No single place in the material lists all four levels.
  5. Numbers connecting measurement quality to sample size come from a sport and exercise science article and were deleted from the teaching notes. They are not standards.

Frequently asked questions

Does the source say reliability is required for validity, or only that it is not enough?

Both, on two different slides. The key-terms slide closes with "Reliability in itself does not ensure validity"; the measurement slide states "Must be reliable in order to claim validity of findings". The two are compatible — reliability is necessary and not sufficient — but the source never says that in one place, so quote both if you are asked to state its position.

How many levels of measurement does this material teach?

Three, and a different three depending which week you read. This week lists nominal, ordinal and ratio; another week lists nominal, ordinal and interval. All four conventional levels are named across the subject and no single place lists them all. The property that separates ratio from interval, a true zero, is never mentioned anywhere.

Is "frequency" really an ordinal variable?

Not on the slide's own terms. A frequency is a count with a true zero and equal intervals, which places it in the ratio category the same slide lists two lines below. This library reports the example as printed rather than correcting it, and notes that the slide gives no definition of either category that would let a reader catch the error.

Can I use the sample sizes the set reading gives for poor validity?

Not as a standard. Those figures — thousands rather than hundreds for a descriptive study with poorly valid measures, twenty per group for a controlled trial with a highly reliable one — are one author's statements about sport and exercise science in 2000. The teaching material deleted that passage entirely and states no equivalent figure in its own voice.

What does the source give me for testing an instrument before I use it?

Nothing procedural. No pilot, retest, split-half or inter-rater check appears in the week, and the passage on pilot studies in its set reading is one of the five the study notes deleted. Plan the pilot anyway, and cite the source you took its design from.

References and source attribution

  1. Hopkins 2000, 'Quantitative Research Design', a sport-science web journal, vol. 4, issue 1 — journal title, address and author given names scrubbed. The week's set reading, and the source of the deleted passage connecting validity and reliability to sample size.
  2. Bryman, A. 2016, Social Research Methods, 5th ed., Oxford University Press, Oxford.
  3. O'Leary, Z. 2017, The Essential Guide to Doing Your Research Project, 3rd ed., Sage Publications, London.
  4. Veal, A. J. 2005, Business Research Methods: A Managerial Approach, Longman.
  5. The supplied teaching source: the slide deck on the nature of quantitative research — in particular its key-terms and measurement slides — and both copies of the accompanying study notes.

Suggested questions for Ask KEVOS

  • Write the measurement subsection of my methodology chapter using only what this material supports.
  • Give every item in my instrument a measurement level and tell me which scheme you used.
  • How do I evidence reliability before I make a validity claim about my instrument?
  • What exactly does this week say about validity, and what do I have to source elsewhere?
  • Explain the difference between measurement error and sampling error, since the source never separates them.

Related KEVOS knowledge

Validity and ReliabilityCore · research designLevels of Measurement in Structured QuestionsCore · data collectionWhat Quantitative Research IsFoundation · quantitative researchPopulation, Sample, Variables and DataFoundation · quantitative researchChoosing What to MeasureAdvanced · quantitative researchSample Size in Quantitative StudiesAdvanced · quantitative research
KEVOS® · Project Delivery · Research Projects Page KVS-PM-RES-0187 · v1.0.0 · content 2026.08 Last reviewed 2026-08-16

Continue learning

Population, Sample, Variables and DataGuide · Research ProjectsNEXT LESSON →Sampling Methods: The Six Named TypesGuide · Research ProjectsWhat Quantitative Research IsGuide · Research ProjectsHow Large Should a Sample Be?Guide · Research Projects
KEVOS · Engineering, manufacturing and project improvement
ArticlesServicesCase studiesAboutContact
© 2026 KEVOS®