Coding Qualitative Data in Practice
The week defines coding three times, in three documents, and the three do not agree on what gets marked, what it gets marked with, or whether the act is objective. All three are here, with the two coding vocabularies that share no term and the worked tables that do not code the same data.
Coding, defined three times, differently
The definition to work from is the fullest one, and it is worth quoting before anything else because it is the only one that says what kind of act coding is.
Two words in that sentence do real work. Interpretive means a code is a reading, not a measurement. Subjective means a second analyst could code the same segment differently without either of you being wrong. Together they set the standard of proof you owe a reader: not that your codes are correct, but that they are traceable. The place this definition sits in the week's own sequence is set out on The Steps in Qualitative Data Analysis.
THE THREE DEFINITIONS SIDE BY SIDE
| Where | Definition, verbatim | Marks what | Marks with what | Interpretive? |
|---|---|---|---|---|
| Study notes, section 4 | "Coding usually starts with a summary of the text you are examining" | The whole text | A summary | Not stated |
| Slide glossary | "Marking the segments of data with symbols or names" | Segments of data | Symbols or names | Not stated |
| The coding slide | "the interpretive and subjective application of symbols and words to segments of transcribed text" | Segments of transcribed text | Symbols or words | Yes - stated explicitly |
Source: the week's study notes and slide deck. Three definitions in three places; none cites either of the others.
The three differences you have to carry
A whole text or a segment
The notes code a whole text with a summary; the slides code segments. The glossary defines segmenting as "Dividing the data into meaningful units" and the notes never mention it - so the notes' procedure has no segmentation step in it.
Names or words
One slide says "symbols or names". The next says "symbols and words". Names and words are not the same class of thing, and the change is silent.
Subjective here, repeatable there
Only the coding slide says the act is interpretive and subjective. The glossary line reads as a clerical operation, and the previous week claims good data coding "allows a study to be repeated and validated". The tension is never addressed.
Transcribed text that is never required
The coding slide requires transcribed text. Neither the four-step model, the stage diagram nor the week's worked example involves a transcript - the worked example codes written survey responses that were never transcribed from anything.
What walking in the participant's shoes actually requires
The metaphor is the deck's only statement about how to stand in relation to your data, and it does more work than it looks. Read with the interpretive and subjective clause beside it, it says the code you assign is your reading of what the participant meant, not a category the participant chose. Three consequences follow, none of them stated by the source.
Where coding starts, according to the study notes
The notes give a two-stage progression, and it is the only account in the week that describes a starting point and a direction of travel.
That progression is worth keeping: a defensible first pass, summarise what is there, and a defensible second, go beyond the summary to something that answers the question. It also, quietly, makes categories the input to coding rather than its output - the reverse of what the previous week says. Neither week acknowledges the other, and both readings are in the material.
Flat and tree coding: the two structures
The notes then give the two shapes a code list can take, with a worked table for each. Both tables come from a third-party study of a public library's computer facilities, which the notes were copied from along with everything else in them; product names and one place name inside the data have been removed here.
Flat, or non-hierarchical
- "Coding can be flat or non- hierarchical, like a list, there are no sub-code levels."
- Two columns, one per survey question, each an undelimited list of responses.
- Left: a computer-literacy certificate, library induction tutorial, genealogy research, holiday research, shopping, visits to a capital city, magazines, books.
- Right: college course, computer programming, word processing, school, introduction to computers, self taught, library courses, online course, the same certificate as an acronym, no training.
- Eight items and ten - counts derived here by enumeration, not stated by the source.
Tree, or hierarchical
- "It can also use tree or hierarchical coding which, like a tree, has a branching arrangement of sub-codes."
- "Ideally, codes in a tree relate to their parents by being 'examples of...', or 'contexts for...' or 'causes of...' or 'settings for...' and so on."
- Left, three parents: services and courses offered by the library; specific websites and internet based services; other services and resources used, marked non-ICT.
- Right, five parents: college course, school, self taught, library courses, no training - the same ten responses regrouped.
- Seventeen nodes and ten - counts derived here, not stated by the source.
The four relation types named above the tree table are the one instruction the source gives about building a hierarchy, and they are the most useful sentence in the section. A child code should be an example of its parent, a context for it, a cause of it or a setting for it. If you cannot say which holds, the parent is a heading rather than a code.
Factual, axial and selective coding
The notes then say "There are several types of coding that can be undertaken:" and give three. Descriptive and analytic coding, named eight lines earlier, are not among them - so one section names five kinds of coding under two framings and never says whether they are alternatives, a sequence, or different cuts of the same thing.
THE THREE TYPES, AS THE SOURCE LEAVES THEM
| Type | What the source says it does | Worked example given? | Result stated? |
|---|---|---|---|
| Factual | Data "are studied and coded according to what each piece of data is an example of" | Yes - six distinct labels, one of which is printed twice | No |
| Axial | "building connections within categories, and between categories and sub-categories", giving "a picture of the relationships" | Yes - two questions to compare | No |
| Selective | "The relationship between a core category and related categories; by focusing on one aspect of the core categories" | Yes - the subgroup who had received no training | No |
Three types, three examples, no example carried through to a finding. Each names the data to look at and stops.
The axial example is the costliest omission, because the data to do it with is already printed on the page. It proposes comparing training received against use made of the library computers, "to see if there are any links" - and then no link is proposed, no result is stated and no matrix is shown. The section stops one step short of the only analysis it could have carried out with the material in front of the reader.
Seven coding names in two families that share no term
THE TWO CODING VOCABULARIES, WITH ZERO OVERLAP
| Named in the study notes | Named on the slides |
|---|---|
| Descriptive coding | A priori codes - "codes developed before data is examined" |
| Analytic or Theoretical coding | Inductive codes - "codes developed whilst data is examined" |
| Factual Coding | Not defined in source |
| Axial coding | Not defined in source |
| Selective coding | Not defined in source |
Five names on one side, two on the other, no term in common. The notes divide coding by what the code does; the slides divide it by when the code is made. Nothing in the week says whether the two schemes are alternatives, orthogonal or compatible.
This is the practical problem the week leaves you with: seven coding names in two unreconciled families and no criterion for choosing, in a week whose first stated objective is to evaluate - a verb that needs criteria and a comparison. No two coding types are compared on any criterion anywhere in the material.
The procedure the source omits
What you will not find anywhere in this week
- A segmentation step in the notes' procedure. Segmenting is defined once in the slide glossary and never enters the notes' account of coding.
- A rule for deciding that two pieces of data belong under the same code, or what to do with a datum that fits two.
- A stopping rule. Nothing says when a code list is finished or how many codes is too many.
- A second coder, an agreement check or any audit of the coding - despite the same week's notes requiring that qualitative work be auditable.
- A worked example that uses any of the five coding types the notes name. The week's one complete analysis categorises rather than codes.
One last piece of context. This week states two learning objectives and the first, "Evaluate presentation styles for qualitative data", is not delivered anywhere in it. Across the six weeks now on record, fourteen objectives are stated, none is fully delivered, four are partially delivered and ten are not delivered at all - see The Six-Week Objective Audit. For the one complete analysis the subject contains, see Inductive Categorisation; for what software does and does not add, Computer-Assisted Qualitative Data Analysis.
What to carry forward
- Work to the fullest definition: coding is the interpretive and subjective application of symbols and words to segments of text. That sets your burden of proof at traceability, not correctness.
- The material gives three definitions, two vocabularies and seven names, with no criterion for choosing. Name the scheme you used and where it came from.
- The four tree relations - example of, context for, cause of, setting for - are the one structural rule the source states, and it never applies them.
- The two worked coding tables do not code the same data. Use them as illustrations of two shapes, not as a demonstration of one transformation.
- No segmentation step, no stopping rule, no second coder and no agreement check appear anywhere. Supply those from a cited text and say that you did.
Frequently asked questions
Do categories come before codes or after them?
Both, according to this material. The study notes make categories the input - data is coded according to categories identified by reading and re-reading - and the four-step slides agree, grouping cases into categories in step 1 and developing code labels in step 2. The previous week of the subject defines coding as creating categories and sub-categories, making them the output. Neither week acknowledges the other, so state the order you used.
If coding is subjective, how do I defend it to an examiner?
Not by claiming it is objective. The slide that calls coding interpretive and subjective is the strongest definition the week gives, and the defence it implies is traceability: keep the segment attached to the code, keep memos of why you read it that way, and describe the process in enough detail for someone else to follow it. The material's own trustworthiness criteria ask for exactly this and its one worked example satisfies none of them.
Should I use flat or tree coding?
The material describes both and gives no criterion for choosing. What it does give is a test for whether a tree is warranted: codes in a tree should relate to their parents as examples of, contexts for, causes of or settings for them. If your candidate parents do not hold that kind of relation to their children, a flat list is the honest structure.
What is the difference between factual, axial and selective coding?
As the source has it: factual codes each datum by what it is an example of; axial builds connections within categories and between categories and sub-categories; selective relates a core category to related categories by focusing on one aspect. Each gets an example that names the data to look at and stops, so none is carried through to a finding, and "core category" is never defined.
Do I have to transcribe before I code?
The coding slide says coding applies to segments of transcribed text, but the week never requires transcription as a step. It appears once in the notes' ordering exercise as the first action and once in the glossary as data entry, and neither the four-step model nor the stage diagram nor the worked example involves a transcript. If your data arrives as written responses, it is already text and there is nothing to transcribe.
How many codes should I end up with?
The material states no number, no range and no stopping rule at any point. It instructs you to combine similar codes and condense, and does not say when to stop. Any figure you use has to come from a text you cite, and the sensible working answer is the smallest set that still distinguishes the things your research question turns on.
References and source attribution
- O'Leary, Z 2010, Doing your research project, Sage Publications - the single reference listed on the week's own closing slide.
- O'Leary, Z. 2017, The Essential Guide to Doing Your Research Project, 3rd ed., Sage Publications, London - the later edition cited elsewhere in the supplied source.
- Bryman, A. 2016, Social Research Methods, 5th ed., Oxford University Press, Oxford - cited elsewhere in the supplied source; a place to obtain the segmentation, stopping and agreement procedures this week does not supply.
- Quinlan, C. 2011, Business Research Methods, 1st ed., Cengage Publishing.
- The supplied teaching source: the qualitative analysis week's study notes, section 4, and its slide deck, slides 9 and 10. The study notes are a capture of a third-party academic study-skills website and their worked coding tables come from a public-library computer-training study unconnected to project delivery. Two near-identical captures of the notes exist and differ in three words, two of them in the section quoted here.
Suggested questions for Ask KEVOS
- Build me a first-pass code list for twelve interview transcripts on project handover failures.
- Which of the seven coding types in this material fits an exploratory study, and what should I cite for it?
- How do I write the coding paragraphs of a methodology chapter when the teaching material gives no procedure?
- Give me a check for whether my code tree's parent-child relations actually hold.
- What would a second coder and an agreement check look like on a study of thirty open-ended responses?
