KEVOS
ArticlesServicesCase studiesAboutContact
ArticlesServicesCase studiesAboutContact
← ArticlesHandling OutliersProject Delivery · Research ProjectsLesson 172/216← PrevNext →
GuidePublished 16 Aug 202615 min readBy KEVOS Editorialoutliersextreme valuesoutlier detectiondata screening outliers
On this page

Ask about this page

KEVOS AIHandling Outliers

KEVOS knowledge first · trusted web sources when needed

KEVOS/Project Delivery/Research Projects/Quantitative Data Analysis
Project DeliveryResearch ProjectsCoreQuantitative Analysis

Handling Outliers

An extreme value is either the most informative observation in your dataset or the one that ruins it, and the supplied deck takes both positions - along with two others. This page prints every characterisation and every instruction in the order they arrive, and leaves the reader holding the decision the source leaves them.

Reading time16 minutes
LevelCore
Topic streamQuantitative Analysis
Source materialQuantitative Data Analysis
Updated2026-08-16

In brief

  • Four slides of one deck describe an outlier four different ways: a value that deviates markedly from other members of the sample, data different from the rest of the sample, an extreme value, and a point obviously different from the main pattern.
  • Three of those slides give three different instructions: remove obviously erroneous figures; treat wrongly including an extreme value as the more serious error; and investigate outliers because they may be valuable to understand.
  • The deck's final word on outliers reverses its first, and it never reconciles the two. Both positions are printed here and neither is adjudicated.
  • No detection rule of any kind is supplied - no cut-off, no interval rule, no score, no inspection procedure. The only method described is your own judgement that a value is implausible.
  • The slide headed "Outlier considerations" asks four questions, of which one concerns outliers, two concern statistical assumptions that are never named, and the last is never answered.

Four characterisations, in the order the deck gives them

An outlier decision is one of the few data-preparation choices that can move a headline finding on its own, so it matters that the material describing it is consistent. Here it is not, and the variation is not cosmetic: the four wordings pull towards two incompatible ideas, one about distance and one about error.

Note

Where this week's material comes from

The quantitative week of the supplied teaching source is a slide deck of 25 slides and nothing else. Its study notes are listed in the upload manifest and are not present in the supplied files, which makes it the only week in the subject without them. In every other week the notes carry the definitions and the slides carry the lists.

Everything this page records as missing is therefore missing from the supplied material. That is not a claim about the subject as taught - the notes that were not supplied may have contained a detection rule, a worked calculation or a decision procedure. It is a statement about what can honestly be drawn from the files, and about where you will have to go instead.

THE FOUR CHARACTERISATIONS, QUOTED AS THEY APPEAR

SlideWording as givenWhat it makes an outlier
4, key terms"an outlying observation, or outlier, is one that appears to deviate markedly or is very distant from other members of the sample in which it occurs"A distributional fact. Says nothing about correctness.
8, outliers"data that is different from the rest of your sample", then "an extreme value", then "Such an obviously erroneous figure"Moves from difference to extremity to error in three sentences, treating them as one thing.
9, outlier considerations"an extreme value"Extremity alone, in a sentence about which error is worse.
18, reading a graphical display"points which appear to be obviously different to the main pattern of the data"A visual fact about a chart, and a signal "which could be quite valuable to understand".

The key-terms definition and the treatment that follows it are not the same idea. Under the key-terms definition a genuine extreme observation is an outlier; under the following slide's treatment it is a candidate for deletion.

From the source

The outliers slide, quoted in full and in order

"If there is data that is different from the rest of your sample, you need to consider where its inclusion will distort your data. This data is usually referred to as an outlier. When reviewing collected data an extreme value may be noticed."

"Such an obviously erroneous figure will be removed from the data."

"However, when there is doubt whether a figure is an error or is genuine," - "to include an erroneous figure in the data would distort the results"; "to ignore a genuine figure would also cause distortion."

"The final decision is left to the researcher's judgement and depends on individual circumstances."

4characterisations of an outlier across one deck
3different instructions about what to do with one
0detection rules supplied anywhere in the week
Caution

A rule stated and withdrawn in consecutive sentences

"Such an obviously erroneous figure will be removed from the data" is unconditional. The next sentence introduces doubt and ends by leaving the final decision to the researcher's judgement. Nothing on the slide says where the boundary of "obviously" lies, so the rule applies exactly as far as your own confidence extends and no further.

Three instructions, and the reversal at the end of the deck

The three positions arrive ten slides apart and are never set beside one another in the source. Setting them beside one another is most of what this page can usefully do.

WHAT THE DECK TELLS YOU TO DO, BY SLIDE

SlideInstruction, in the source's wordsDirection it pushes
8"Such an obviously erroneous figure will be removed from the data" - qualified immediately by the risk that "to ignore a genuine figure would also cause distortion"Remove errors; then, on reflection, use judgement
9"Wrongly including an extreme value is usually regarded as a more serious error than wrongly ignoring one"Err towards exclusion
18Outliers "often point to some set of special circumstances which could be quite valuable to understand"Investigate; the value may be the finding

The deck states no priority among the three and never refers any one of them to the others.

Source gap

Three positions, no reconciliation, and this page does not choose either

The asymmetry on slide 9 - that wrongly including an extreme value is the more serious error - is stated with a hedge ("usually regarded"), with no attribution, no criterion, and no measure of how much more serious. Ten slides later the same deck describes outliers as valuable to understand.

Where the supplied material takes two positions, this library gives both and marks the conflict. It does not adjudicate, because adjudicating would mean supplying a rule the source does not contain and attributing it to the source. What the source does supply, and it is worth quoting when you write this up, is that "The final decision is left to the researcher's judgement and depends on individual circumstances."

The two obviously-wrong values

Source example — illustrative only

The deck's two illustrations of an obvious error

"For example, when in data about ages of children in year one at school, we find a figure of 34 years, it suggests the teacher's age was collected by mistake. Similarly, a height of 62 metres in data relating to men's heights suggests a misreading of the scale or a clerical error."

Both figures are illustrations of implausibility, not applications of a rule. Neither example states the size of the dataset, its true range, or what the correct value would have been. Nothing in them generalises to a threshold, and they should never be quoted as one.

The examples work because a reader already knows how old a year-one child is and how tall a man is. That is domain knowledge, not statistics, and it is the only detection method the week describes. It will find a data-entry slip in a variable you understand, and it will not find anything in a variable you are studying precisely because you do not yet understand it.

The one numerical illustration, and what it quietly demonstrates

Source example — illustrative only

The deck's only numeric illustration of an outlier's effect

"For example, the average amount of money brought onto the cruise ship by cruisers is averaged as $10,000 according to your survey. But in fact, if you take out a cashed up millionaire in their midst, nobody has more than $1,000."

Both figures are examples. No sample size is given, no individual values are given and no arithmetic is shown, so the example cannot be reproduced or checked. A reader cannot see from it how large the outlying value would have to be, or how small the sample, to produce the effect described.

Two things about this example repay attention. The first is that the person removed is presumably a genuine respondent rather than an error - which makes it an argument for care with genuine extremes, not for the asymmetry stated directly beneath it. The second is that the word "averaged" leaves it unstated which average is meant; the deck defines three of them four slides later.

Source gap

The lesson available in the deck's own materials, and never drawn

The example shows one respondent moving a reported average by an order of magnitude. The same deck defines a measure that "divides the distribution in half", and the two are never connected: the deck does not say what a single extreme value does to one measure of centre and not to another, and it does not say that the choice of average is a decision the cruise-ship problem forces.

Whether one of the three measures would have been a better summary here is not something the supplied material addresses, and this page will not supply the answer on its behalf. The three definitions, exactly as the deck gives them, are at Measures of Central Tendency.

No detection rule, anywhere in the week

This is the single most consequential absence on the topic. Every practical question a reader arrives with - how far from the centre is far enough, what to do with a value that is unusual but plausible, how to be consistent across variables - requires a rule, and the week contains none.

Source gap

What is not supplied

No cut-off expressed in standard deviations. No interquartile-range rule. No standardised score. No visual inspection procedure. No test, and no threshold of any kind. Naming these here says what is absent; it does not teach any of them, and none of them can be attributed to this material.

The absence compounds. A numeric detection rule would have to be expressed in some measure of spread, and no measure of spread is defined anywhere in the deck - see Variation and Spread. The one statistic the week's quantitative rule depends on is used four times and defined nowhere, which is set out at The Normal Model and the 68-95-99.7 Rule.

The deck's other route to spotting an extreme value is visual: its guidance on reading a chart asks whether any points appear obviously different from the main pattern. That is a real technique, and the week supplies no illustration of it - the two image slots in its presentation sequence are a production placeholder and an unlabelled picture with no recoverable content. See Reading a Graphical Display.

The six questions the deck asks before you enter data

Two slides carry questions to ask of a dataset. They are worth having: as a prompt list they are better than the deck's treatment of them. Read them as questions you now have to answer from somewhere else.

The six questions, quoted exactly and in order

  • Does the data accurately reflect the responses made by the participants of my study?
  • Is all the data in place and accounted for, or is some absent or missing?
  • Is there a pattern to the missing data
  • Are there any unusual or extreme responses present in the data set that may distort my understanding of the phenomena under study?
  • Does the data meet the statistical assumptions that underlie the multivariate technique I will be using?
  • What can I do if some of the statistical assumptions turn out to be violated?
Source gap

Three of the six rest on terms the week never supplies

"Statistical assumptions" is named twice on one slide and never defined, listed or exemplified. No assumption of any kind is named anywhere in the deck, so a reader is asked whether their data meet assumptions they have not been told exist.

"The multivariate technique I will be using" presupposes a technique the deck never names - no multivariate method appears anywhere in the week. And the last question, "What can I do if some of the statistical assumptions turn out to be violated?", is the final question of the outlier sequence: the deck moves straight on to a list of analysis issues and never returns to it. It is a question posed on a topic that was not introduced, about techniques that are not named, and it is not answered.

The pattern-in-missing-data question is in the same position - asked once, never answered - and is covered alongside the screening material at Data Screening and Cleaning.

Caution

The slide's title covers a quarter of its content

The slide is headed "Outlier considerations" and only the second of its four questions concerns outliers. The first concerns missing data; the last two concern distributional assumptions and multivariate methods. If you are searching this material for what it says about outliers, most of what is filed under that heading is about something else.

Making an outlier decision you can defend

The source leaves the decision to your judgement and says so explicitly. That makes the record of the decision the only thing an examiner or a client can assess, which changes what you have to do about it.

A defensible record - guidance beyond the source, built on what the source does say

  1. Decide the rule before you look at the results

    A rule chosen after you have seen which way it moves your finding is not a rule. The source does not require this; nothing else protects you from the objection.

  2. State the rule you used and where you took it from

    The supplied material contains no detection rule, so any rule you apply comes from elsewhere and should be cited to the text it came from.

  3. Separate errors from genuine extremes in your write-up

    The deck's own distinction is between a figure that is erroneous and one that is genuine, with different consequences stated for each. Report how many of each you found.

  4. Report the analysis both ways when it matters

    If dropping the value changes the finding, the reader needs to know. The source does not ask for this; the cruise-ship example is the argument for it.

  5. Say what an unusual case might mean

    The deck's last word is that outliers often point to special circumstances that could be valuable to understand. A sentence about what the extreme case was is often worth more than the mean it disturbed.

Check before you proceed

Before you delete anything

Can you name the rule, say where it came from, say how many observations it removed, and show what the result looks like with them and without them? If any answer is missing, you are not yet in a position to delete, and the source's own instruction - that the decision rests on your judgement in the circumstances - offers you no cover.

What to carry forward

  1. Four characterisations across four slides, and the distributional one on the key-terms slide is not the same idea as the error-focused one that follows it.
  2. Three instructions, ending with the opposite of where the deck began. Both ends stand; the source resolves neither.
  3. No detection rule of any kind is supplied. Whatever rule you use comes from outside this material and should be cited there.
  4. The cruise-ship figures are an illustration of one respondent's effect on a reported average. They are not a threshold, a benchmark or a worked calculation - no arithmetic is shown.
  5. "Statistical assumptions" is invoked twice and never defined, and the question about what to do when they are violated is asked and abandoned.

Frequently asked questions

What counts as an outlier according to this material?

It gives four answers. One slide defines it distributionally, as an observation that appears to deviate markedly or is very distant from other members of the sample. Another treats it as data different from the rest of the sample, then as an extreme value, then as an obviously erroneous figure. A third calls it an extreme value; a fourth calls it a point obviously different from the main pattern of the data. The source does not reconcile them and neither does this page.

Should I delete outliers?

The supplied material takes three positions: that an obviously erroneous figure will be removed, that wrongly including an extreme value is usually regarded as the more serious error, and that outliers often point to special circumstances that could be valuable to understand. Its own resolution is that the final decision is left to the researcher's judgement and depends on individual circumstances - so what you can defend is your reasoning and your record of it, not an appeal to the material.

How far from the mean does a value have to be before it is an outlier?

The supplied material states no distance, threshold or cut-off anywhere. It gives no rule expressed in standard deviations or in any other measure, and no measure of spread is defined anywhere in the week. Any numeric rule you use will come from a text outside this material and should be cited to it.

What are the statistical assumptions the deck asks about?

It never says. The phrase appears twice on one slide, in the question of whether your data meet the assumptions underlying the multivariate technique you will use, and no assumption is named, defined or exemplified anywhere in the week. No multivariate technique is named either, and the follow-up question about what to do when assumptions are violated is never answered.

Does removing one respondent really change an average that much?

The deck's cruise-ship illustration says a survey average of $10,000 becomes a maximum of $1,000 once one wealthy respondent is taken out. It shows no arithmetic and states no sample size, so the figures illustrate the direction of the effect rather than its size. Treat them as an example of what can happen, never as a benchmark.

Is an outlier always an error?

Not on the material's own key-terms definition, which is about distance from the rest of the sample and says nothing about correctness. The slide that follows nevertheless moves from difference to extremity to obvious error in three sentences. Its own worked case - a wealthy cruise passenger - is a genuine respondent rather than a mistake, which is the clearest evidence in the deck that the two ideas are not the same.

References and source attribution

  1. Veal, A. J. 2005, Business Research Methods: A Managerial Approach, Longman - the one work with complete bibliographic data cited in the supplied quantitative week.
  2. Bryman, A. 2016, Social Research Methods, 5th ed., Oxford University Press, Oxford - cited elsewhere in the supplied source; a place to obtain the outlier detection rules this week does not supply.
  3. O'Leary, Z. 2017, The Essential Guide to Doing Your Research Project, 3rd ed., Sage Publications, London.
  4. Naoum, S. G. 2013, Dissertation Research & Writing for Construction Students, 3rd ed., Routledge.
  5. The supplied teaching source: the quantitative analysis and presentation slide deck, slides 4, 8, 9, 10 and 18. The week's study notes are listed in the upload manifest and are not present in the supplied files.

Suggested questions for Ask KEVOS

  • Give me a cited outlier detection rule suitable for a survey with about sixty responses.
  • How should I report an analysis that changes when one extreme respondent is removed?
  • What are the usual assumptions behind common multivariate techniques, since this material names none?
  • Draft the outlier paragraph of a methods chapter that is honest about where the rule came from.
  • Help me decide whether an unusually long project duration in my dataset is an error or a finding.

Related KEVOS knowledge

Data Screening and CleaningCore · quantitative analysisMeasures of Central TendencyFoundation · quantitative analysisReading a Graphical DisplayCore · quantitative analysisVariation and SpreadCore · quantitative analysisDescriptive and Inferential StatisticsCore · quantitative analysisWhat This Material Does Not TeachCore · research practice
KEVOS® · Project Delivery · Research Projects Page KVS-PM-RES-0172 · v1.0.0 · content 2026.08 Last reviewed 2026-08-16

Continue learning

Coding Quantitative DataGuide · Research ProjectsNEXT LESSON →Measures of Central TendencyGuide · Research ProjectsData Screening and CleaningGuide · Research ProjectsDescriptive and Inferential StatisticsGuide · Research Projects
KEVOS · Engineering, manufacturing and project improvement
ArticlesServicesCase studiesAboutContact
© 2026 KEVOS®