Coding Quantitative Data
Two slides tell you what coding is, what it achieves, what dictates the method, and whether the codes come before or after collection. Between them they never show a code, a category or a variable, and this page is careful not to supply one on their behalf.
What the deck means by coding, and the two things it means at once
Coding is where a pile of completed instruments becomes a dataset. Get it wrong and every statistic downstream is a statistic about your clerical decisions rather than about your respondents, which is why the deck's brevity here is worth mapping precisely.
Read those four statements together and they describe two different jobs. Transforming data into a form software can read is mechanical: a value goes in a column in an agreed format. Refining data into categories and sub-categories is interpretive: somebody decides what the categories are, and two competent people can decide differently.
Coding as data entry
- "The transformation of data into a form understandable by computer software".
- Preparation "for computer processing with statistical software".
- Entry "in a predetermined format".
- Judged by whether the software can read it.
Coding as classification
- Data "refined into smaller units".
- Through "creating categories and sub- categories that can then be analysed".
- Categories developed before or after collection depending on research logic.
- Judged by whether the categories carry the meaning - and the deck states no test for that.
When the codes get made: the one classification the deck offers
This is the deck's single substantive coding distinction, and it is a good one: it ties the timing of your coding scheme to the logic of your study rather than to your convenience. It is set out in one line each way, with no elaboration.
THE TIMING RULE, AS THE SLIDE STATES IT
| Research logic | When categories and codes are developed | Trigger stated by the source |
|---|---|---|
| Deductive | Before data is collected | "When testing a hypothesis" |
| Inductive | After examining the collected data | "When generating a theory" |
The source states no count and names no other timing. Note the hedging: deductive codes "can be developed" before collection; inductive codes "are generated" after.
The rule as a decision, in the source's own terms
The underlying distinction is treated properly elsewhere in the library - see Inductive and Deductive Reasoning in Research - and this slide is its only appearance in the quantitative week. If your design is mixed, you are choosing your own coding timing, because the material does not cover the case.
What coding is for, and the four properties of good coding
The four properties named for good coding
The study can be repeated
Somebody else, holding your codes, could code the same raw material the same way. The deck states this as an outcome of good coding and gives no test for it.
The study can be validated
The coding can be checked. What checking would consist of, who would do it and against what standard is not stated anywhere in the week.
Comparison with other studies is possible
Your categories can be set beside somebody else's. The deck does not mention published coding frames, standard variables or any mechanism by which comparability would be achieved.
The method is transparent
"By recording analytical thinking used to devise codes". This is the one property the slide attaches a mechanism to, and the mechanism is a record - which is what the code book, defined and abandoned, would have been.
The determinant the deck cannot support
"The way a variable has been measured" is stated as the first thing that dictates your coding method. It is the right dependency: how you coded an item constrains what you can legitimately do with it later. The problem is that the supplied material never puts the measurement levels in your hands.
The second determinant - "The way you want to communicate your findings" - points forward to the presentation half of the week and is not connected to it. Nothing in the deck says which coding decisions constrain which display, and nothing in the presentation slides refers back to coding.
The code book, defined once and abandoned
That is the whole of it. The term is defined on the key-terms slide and appears nowhere else in the week - including on the two slides devoted to coding. No code book is shown, no format is described, no fields are listed and no entry is exemplified.
The worked example the deck never gives
For a topic whose whole content is the transformation of answers into values, the week contains no instance of that transformation. It is worth being precise about what is missing, because the list is what your second source has to supply.
WHAT A READER WOULD NEED, AND WHAT THE WEEK SUPPLIES
| Element | In the supplied week? | Note |
|---|---|---|
| A single code | Absent | No code of any kind is printed in the deck |
| A category and a sub-category | Absent | Named as the unit of refinement; never exemplified |
| A variable name | Absent | No naming rule and no instance |
| A value label | Absent | No convention stated |
| The "predetermined format" | Absent | Named on the coding slide and never specified |
| A code book or extract of one | Absent | Defined on the key-terms slide, never used again |
| A named coding option | Absent | See the gap note below |
| Software steps for entering codes | Absent | The week's software content is one sentence pointing at an external chapter |
The only coded example anywhere in the supplied source is a three-value coding of a language question on one slide of the data collection week, reproduced at Levels of Measurement. It is not referred to by the quantitative week.
Repeatability and the note-taking instruction that undercuts it
The claim that good coding "allows a study to be repeated" is the week's single quality criterion, and the week before it instructs the researcher in the opposite direction.
Using the material honestly in your own project
What the deck gives you is a purpose, a timing rule, two determinants and four properties. That is a frame for a coding section, not a coding scheme, and it can carry a methods chapter as long as you do not pretend it carried more.
Practical steps this material supports - with the source's limits marked
- State your research logic and derive your coding timing from it. This is the one rule the source actually supplies.
- Say how each variable was measured before you say how it was coded, and cite the text you took the measurement levels from - the supplied material does not supply them usably.
- Write the code book anyway, in whatever format your cited text specifies. The source names the artefact and defines nothing about it.
- Record the reasoning behind each category, not only the category. The source's own transparency property is about "recording analytical thinking used to devise codes".
- Do not claim your coding scheme is standard, comparable or validated unless you can say against what. The source names those properties and supplies no test for any of them.
- This checklist goes beyond the source in its detail. Only the first item and the last clause of the fourth are stated in the supplied material.
What to carry forward
- One definition of coding is made to cover two different jobs - entry and classification - and the deck never separates them.
- The timing rule is the week's one solid contribution: hypothesis testing means codes before collection, theory generation means codes after.
- Four properties of good coding are named and no criterion is given for any of them, in a week whose objective is to evaluate coding options.
- The code book is defined once and abandoned; "a predetermined format" is named and never specified. Get both from a text you cite.
- No code, category, sub-category or worked coding example appears anywhere in the week. Treat the absence as your reading list, not as a gap to fill from memory.
Frequently asked questions
What coding options does this material actually teach?
None by name. It teaches when codes are developed - before collection for a deductive study, after examining the data for an inductive one - and it names two things that dictate the method: how the variable was measured and how you want to communicate your findings. No coding scheme, technique or family is named anywhere in the week, which is why its own objective of evaluating coding options is recorded as undelivered.
What should a code book contain?
The supplied material does not say. It defines a code book, once, as the guide to all of the codes used in coding data for input into a software program, and never mentions it again - no fields, no format, no example. Take a template from a methods text you have read and cite it, rather than attributing one to this material.
Is coding the same in qualitative and quantitative work?
The supplied deck applies one definition to both without saying so. Its definition is machine-facing - transforming data into a form software can understand - while its description of creating categories and sub-categories is the interpretive activity the data collection week describes as coding emergent themes. The deck does not distinguish them, so treat any claim about the difference as coming from your own reading.
Why does the coding method depend on how a variable was measured?
The slide states the dependency and does not explain it. That explanation would require the levels of measurement, and the only slide in the supplied source that covers them gives three levels rather than four, defines one, and files its worked example under a level that contradicts its own definitions. So the dependency is asserted here and its foundation is not available in the supplied material.
Can I code before I collect the data?
Yes, according to the source, if you are testing a hypothesis - it states that categories and codes can be developed before data is collected in a deductive study. If you are generating a theory, it states that categories and codes are generated after examining the collected data. It gives no guidance for a design that does both.
Does good coding mean somebody else could repeat my study?
That is what the slide claims: good coding allows a study to be repeated and validated, allows comparison with other studies, and makes methods transparent by recording the analytical thinking used to devise the codes. It supplies no test for any of the four, and the previous week's instruction to develop your own shorthand for recording responses points the opposite way.
References and source attribution
- Bryman, A. 2016, Social Research Methods, 5th ed., Oxford University Press, Oxford - cited elsewhere in the supplied source; a place to obtain code book formats and named coding schemes, which this week does not supply.
- Veal, A. J. 2005, Business Research Methods: A Managerial Approach, Longman - the one work fully cited in the supplied quantitative week.
- O'Leary, Z. 2017, The Essential Guide to Doing Your Research Project, 3rd ed., Sage Publications, London.
- Naoum, S. G. 2013, Dissertation Research & Writing for Construction Students, 3rd ed., Routledge.
- The supplied teaching source: the quantitative analysis and presentation slide deck, slides 4, 6 and 7, read against the data collection week's slides and study notes. The quantitative week's study notes are listed in the upload manifest and are not present in the supplied files.
Suggested questions for Ask KEVOS
- Draft a code book layout for a forty-item project delivery survey and tell me which text to cite for it.
- Which coding schemes exist for closed survey questions, and which suits an ordinal scale?
- My design is mixed methods - should my codes come before or after collection?
- Write the coding paragraphs of a methods chapter that is honest about where the scheme came from.
- What would a second coder need from me to reproduce my coding exactly?
