KEVOS
ArticlesServicesCase studiesAboutContact
ArticlesServicesCase studiesAboutContact
← ArticlesData Screening and CleaningProject Delivery · Research ProjectsLesson 170/216← PrevNext →
GuidePublished 16 Aug 202614 min readBy KEVOS Editorialdata screeningdata cleaningdata preparation researchmissing data
On this page

Ask about this page

KEVOS AIData Screening and Cleaning

KEVOS knowledge first · trusted web sources when needed

KEVOS/Project Delivery/Research Projects/Quantitative Data Analysis
Project DeliveryResearch ProjectsCoreQuantitative Analysis

Data Screening and Cleaning

Screening is the first thing that happens to a dataset after collection and the last cheap moment to discover it is broken. This page reproduces what the supplied deck asks you to look for, in what order it says to do it, and the several places where its two accounts of the same activity part company.

Reading time16 minutes
LevelCore
Topic streamQuantitative Analysis
Source materialQuantitative Data Analysis
Updated2026-08-16

In brief

  • The supplied deck defines screening and cleaning twice, on consecutive slides, with different defect lists and different scope. Both definitions are printed here and neither is preferred over the other.
  • Across the two slides you are given nine defect labels and no canonical set. Only one item recurs in recognisable form between the lists: misclassification on the first, miscoded on the second.
  • The two slides also give two different orderings of collection, entry, screening, cleaning and analysis, and no rule for which applies.
  • Two of the four things you are told to check for - "internal consistence" and "messy data" - are bare labels with no definition, no test and no remedy anywhere in the week.
  • No threshold appears anywhere: nothing states how much missing data is too much, or what makes an incomplete response unusable. No procedure for handling missing data is given at any point.

What screening and cleaning are for in this material

Screening and cleaning sit at the front of the week's data-preparation sequence, before coding and before anything is analysed. They are the operations that decide what your dataset actually contains, which is why a defect missed here becomes a finding later.

Note

Week 11 is a slide deck and nothing else

The supplied material for this week is a single deck of 25 slides. The week's study notes appear in the upload manifest and are not present in the supplied files, which makes this the only week in the subject with no study notes at all. In every other week it is the notes, not the slides, that carry the definitions.

Every gap recorded on this page is therefore a gap in the supplied material. It is not evidence that the subject as taught omitted these things - the missing notes may well have contained them. It does mean you cannot obtain them from what was supplied, and this library will not invent them on the source's behalf.

The deck files this material under a part it announces as "Analysing Data". What the part contains is data management, screening, coding and outliers - preparation rather than analysis - and the deck performs no analysis at any point. None of the week's three stated learning objectives names screening or cleaning. Of those three objectives, one is partially delivered and two are not delivered at all; the consolidated audit is on What This Material Does Not Teach.

From the source

The two definitions, quoted in full and in the order they appear

Slide 4, in the week's key-terms list: "Data Screening/cleaning involves going over the data and information you have collected and identifying any gaps, errors, incomplete questions, misclassification or illegible responses".

Slide 5, the following slide, under the heading Data screening: "Data screening involves going over the data collected and identifying any errors." It then instructs: "Check for internal consistence, miscoded, missing, or messy data e.g. do you have all the data requested, have you collected all the questionnaires?"

The first definition treats the two words as one activity - the slide's own heading is "Data Screening/cleaning" - and gives it five kinds of defect to find. The second narrows the activity to "any errors", one kind of defect, and then lists four categories to check that only partly overlap the first list.

Slide 4 - the key-term definition

  • Screening and cleaning named as one activity.
  • Applies to "the data and information you have collected".
  • Five defect types: gaps, errors, incomplete questions, misclassification, illegible responses.
  • Implies inspection of collected material, before entry.

Slide 5 - the screening slide

  • Screening named alone; cleaning appears only in the closing sentence.
  • Applies to "the data collected" and identifies "any errors".
  • Four check categories: internal consistence, miscoded, missing, messy data.
  • States the sequence as entry, then cleaning, then analysis.

THE NINE DEFECT LABELS ACROSS THE TWO SLIDES

Label as printedSlideDefined or explained anywhere?
Gaps4Label only
Errors4 and 5Label only; slide 5 makes it the whole of screening
Incomplete questions4Label only; no rule for when an incomplete response is unusable
Misclassification4Label only
Illegible responses4Label only
Internal consistence5Label only, and misspelled in the source
Miscoded5Label only; the nearest treatment is the coding slides, which give no example of a code
Missing5Label only; no procedure for handling missing data appears in the week
Messy data5Label only; no definition and no remedy

Nine labels, two lists, no canonical set. Only misclassification (slide 4) and miscoded (slide 5) are recognisably the same idea, and the deck does not say so.

Source gap

Two of the four checks are labels with nothing behind them

"Internal consistence" is the deck's only allusion to a reliability concept in the screening context. No definition, no procedure and no statistic is attached to it, and the word is misspelled in the source. Elsewhere in the same deck, "Reliability and Validity" appears as one of sixteen unexplained labels in a list of analysis issues. The one definition of those terms anywhere in the supplied material is in the research design week - see Validity and Reliability.

"Messy data" is a bare label with no definition and no remedy. You are told to check for it and given nothing that would let you recognise it or fix it.

Four checks, two prompts and no threshold

Slide 5 is the closest thing in the week to a screening procedure. It supplies four categories to check and two worked prompts, and stops there.

The screening checks, exactly as the slide states them

  • Check for internal consistence.
  • Check for miscoded data.
  • Check for missing data.
  • Check for messy data.
  • Do you have all the data requested?
  • Have you collected all the questionnaires?
Source example — illustrative only

The two prompts are examples, not a specification

"Do you have all the data requested, have you collected all the questionnaires?" is offered by the slide as an illustration of what checking looks like, introduced with "e.g.". Treat it as an example of the kind of question to ask, not as the definition of a complete check.

Source gap

No number, threshold or tolerance appears anywhere

The slide does not say how much missing data is too much, what proportion of an incomplete questionnaire makes it unusable, or what level of inconsistency should stop a case being analysed. No figure of any kind appears in the screening material.

No procedure for handling missing data is given anywhere in the deck either. Two slides later, in the outlier sequence, the deck asks "Is there a pattern to the missing data" - and no answer, no technique, no substitution method and no deletion rule follows. The question is posed and abandoned.

Which comes first: screening, entry or cleaning

The two slides put the same three steps in two different orders, and nothing in the week reconciles them. This matters more than it looks, because the order determines what you are inspecting: a paper form and a spreadsheet row fail in different ways.

THE TWO ORDERINGS, AS EACH SLIDE STATES OR IMPLIES THEM

SlideOrderingHow the source words it
4Collect, then screen, then enterScreening is inspection of "the data and information you have collected" - the material as gathered, which places it before entry.
5Collect, then enter, then clean, then analyse"Once the data has been entered into the computer system and cleaned you can start to analyse it".

The deck states no rule for which ordering applies, and does not acknowledge that it gives two.

Caution

Do not resolve the ordering by assuming one is a typo

Both statements are load-bearing in their own slides and neither is presented as an aside. If your method chapter needs to state when screening happened, state the order you actually used and why, rather than citing the material for an order it gives two ways.

The remedy the deck offers, and when you will not be able to use it

Having told you to find defects, the slide offers exactly one remedy, and it is a remedy that depends on being able to trace a response back to the person who gave it.

From the source

The remedy, quoted in full

"By reviewing the data you can seek to go back to the respondent or original source of the data e.g. review any video/tape recordings etc."

Caution

This instruction and the week before it cannot both be satisfied

Going back to the respondent requires knowing which respondent gave which response. The data collection week instructs the researcher to "Ensure anonymity", and in a genuinely anonymous survey there is no route from a defective record back to the person who produced it. Neither slide notes the conflict.

The two positions are set out at Confidentiality and Anonymity, and the recording practices the remedy assumes are covered at Conducting, Recording and Transcribing Interviews.

Practice note

Where this decision actually gets made

The source does not prescribe this; in practice, whether you can go back to a respondent is settled when you design the instrument and the consent wording, not when you screen. A code that links a response to a participant list held separately is one arrangement; full anonymity is another; the material describes neither.

Decide it before you collect, write it into your ethics application, and expect your screening options to be whatever that decision left you.

Data management: one term, two incompatible definitions in two weeks

Screening arrives as the second of five key terms, under a first term that is defined here in flat contradiction of the week before it. This is the most consequential single disagreement between the two weeks, and neither document acknowledges the other.

"DATA MANAGEMENT" AS THE TWO WEEKS DEFINE IT

WhereDefinition as givenWhat it includes and excludes
Data collection week, study notes"The process of data collection and analysis"Includes analysis; excludes software. The whole research pipeline.
This week, slide 4"becoming familiar with appropriate software, logging in, entering data and cleaning data"Includes software; excludes both collection and analysis. Four clerical operations.

The two definitions do not overlap. If you use the term in a methods chapter, say which sense you mean.

Source gap

Cleaning is inside data management and beside it at the same time

Slide 4 lists "cleaning data" as one of the four operations that make up data management, and then presents data cleaning as a separate key term of equal standing in the same list of five. Two of the five terms stand in a whole-part relation and the slide sets them out as parallel without saying so.

The five key terms this slide sets up

Screening does not arrive alone. The slide that defines it defines four other terms, and the rest of the week's preparation content runs off them. Three are treated in full on other pages of this library.

THE FIVE KEY TERMS, QUOTED AS THE SLIDE GIVES THEM

TermDefinition as givenWhere it is treated
Data management"involves becoming familiar with appropriate software, logging in, entering data and cleaning data"This page - and see the contradiction above
Data Screening/cleaning"involves going over the data and information you have collected and identifying any gaps, errors, incomplete questions, misclassification or illegible responses"This page
Coding Keys"also called a Code book. Guide to all of the codes used in coding data to input data into a computerised software program"Coding Quantitative Data
Data Coding"is the transformation of data into a form understandable by computer software"Coding Quantitative Data
Outliers"an outlying observation, or outlier, is one that appears to deviate markedly or is very distant from other members of the sample in which it occurs"Handling Outliers

All five definitions face a computer and no software is named on the slide. "Appropriate software", "a computerised software program" and "computer software" are as specific as it gets; a package is named only in the week's last two slides, and only as a pointer to an external chapter.

Working a screening pass from material this thin

What the supplied deck gives you is a list of things that go wrong and a pair of prompts. That is genuinely useful as a prompt list and genuinely insufficient as a procedure, and the honest way to use it is to treat it as the former.

What the source supports, and where you have to go elsewhere

  1. Take the nine labels as your defect vocabulary

    Gaps, errors, incomplete questions, misclassification, illegible responses, internal consistence, miscoded, missing, messy. The deck supplies the vocabulary and no test for any of them.

  2. Fix your own order and record it

    The material gives two orderings of collect, enter, screen, clean and analyse. Choose one, apply it consistently, and describe what you did rather than citing the week for a sequence it does not settle.

  3. Set your own thresholds before you look

    The source states none. Deciding after you have seen the data how much missing data you will tolerate is how a screening rule becomes a result. This is guidance beyond the source, not a source requirement.

  4. Take missing-data technique from a cited text

    No substitution, imputation or deletion method appears anywhere in the week. Use a methods text you have actually read and cite it, rather than attributing a method to this material.

  5. Log every change you make to the dataset

    The deck does not require this. The week's own criterion for good coding - that a study can be repeated and validated - cannot be met for cleaning decisions that were never written down.

Check before you proceed

Before you move on to coding

Can you state, in one sentence each: what you checked for, in what order you did it relative to data entry, what you did with incomplete cases, and how many cases and values you changed or dropped? If any of the four is unanswerable, the problem is in your record-keeping rather than in the data, and it is much cheaper to fix now than in a viva.

What to carry forward

  1. The supplied deck defines screening and cleaning twice, differently, on consecutive slides. Both definitions stand; neither is the canonical one.
  2. Nine defect labels, four checks, two prompts, two orderings, no thresholds and no missing-data procedure. That is the entirety of the material.
  3. "Internal consistence" and "messy data" are labels only. If your method chapter needs them, get the definitions from a text you cite.
  4. The one remedy offered - go back to the respondent - conflicts with the anonymity requirement stated in the previous week, and the source does not resolve it.
  5. "Data management" means the whole research pipeline in one week and four clerical operations in the next. Say which you mean whenever you use the term.

Frequently asked questions

Are screening and cleaning the same thing?

The supplied material says both. Its key-terms slide runs them together as "Data Screening/cleaning" and defines them as one activity of finding defects; the following slide defines screening alone as identifying errors, and mentions cleaning only as something that happens after data entry and before analysis. No statement anywhere in the week distinguishes the two operations, so if you need the distinction, take it from a cited text.

How much missing data is too much?

The supplied material states no threshold, proportion or tolerance of any kind, and gives no procedure for handling missing data. It asks whether there is a pattern to the missing data and never answers the question. Any figure you use will have to come from a source outside this material and should be cited as such.

Should I screen before or after entering the data?

The deck gives both orders on consecutive slides and no rule for choosing. In practice the two catch different things - a paper form can be illegible, a spreadsheet row cannot - and the defensible move is to do both, state what you did, and not cite the material for a sequence it does not settle.

What is "internal consistence"?

It is a bare label in the supplied material, misspelled, with no definition, no procedure and no statistic attached. It is the deck's only allusion to a reliability idea in the screening context, and the only definition of reliability anywhere in the supplied material sits in the research design week.

Can I really go back and ask a respondent about a strange answer?

Only if your design allows it, and the material is in two minds. This week instructs you to go back to the respondent or the original source; the data collection week instructs you to ensure anonymity. Both cannot hold in an anonymous survey, so the decision is made when you design the instrument and the consent wording rather than during screening.

Does this material tell me how to clean data in software?

No. All five key terms on the slide face a computer, and no software is named there. The week's only software content is a single sentence in its last content slide pointing at chapter 16 of the prescribed text - no menu path, no step, no screenshot and no output.

References and source attribution

  1. Veal, A. J. 2005, Business Research Methods: A Managerial Approach, Longman - the one work with complete bibliographic data cited anywhere in the supplied quantitative week.
  2. Bryman, A. 2016, Social Research Methods, 5th ed., Oxford University Press, Oxford - cited elsewhere in the supplied source; a place to obtain the screening and missing-data procedures this week does not supply.
  3. O'Leary, Z. 2017, The Essential Guide to Doing Your Research Project, 3rd ed., Sage Publications, London.
  4. Naoum, S. G. 2013, Dissertation Research & Writing for Construction Students, 3rd ed., Routledge.
  5. The supplied teaching source: the quantitative analysis and presentation slide deck, slides 4 and 5, read against the data collection week's study notes and slides. The quantitative week's study notes are listed in the upload manifest and are not present in the supplied files.

Suggested questions for Ask KEVOS

  • Give me a screening checklist that covers all nine defect labels this material names.
  • What missing-data rules would be defensible for a survey of about eighty project managers, and which text should I cite for them?
  • Draft the data preparation paragraphs of a methods chapter, given that the teaching material supplies no thresholds.
  • How do I design a survey so that I can trace a defective response back without breaking anonymity?
  • Show me what this material does and does not supply for handling missing data.

Related KEVOS knowledge

Coding Quantitative DataCore · quantitative analysisHandling OutliersCore · quantitative analysisDescriptive and Inferential StatisticsCore · quantitative analysisQuestionnaires and Response RatesCore · data collectionLevels of Measurement in Structured QuestionsCore · data collectionWhat This Material Does Not TeachCore · research practice
KEVOS® · Project Delivery · Research Projects Page KVS-PM-RES-0170 · v1.0.0 · content 2026.08 Last reviewed 2026-08-16

Continue learning

Designing a Data Collection Instrument: A Worked AttemptGuide · Research ProjectsNEXT LESSON →Coding Quantitative DataGuide · Research ProjectsConducting, Recording and Transcribing InterviewsGuide · Research ProjectsHandling OutliersGuide · Research Projects
KEVOS · Engineering, manufacturing and project improvement
ArticlesServicesCase studiesAboutContact
© 2026 KEVOS®