Reading a Graphical Display
One slide gives a four-question procedure for looking at a chart; another gives the vocabulary of a distribution. Both are usable, neither is illustrated, and the deck's last word on outliers reverses its first.
The four key aspects, quoted in full
This is the most directly usable slide in the week. It gives a reader four questions to put to any chart, in order, each with a reason attached. It is a reading procedure rather than a list of terms, and it will improve how you look at a display immediately.
The stated count is four and the list beneath it holds four. That is worth recording, because it happens twice in the entire week — here, and on the slide defining a distribution. Everywhere else in this material an enumeration is left for the reader to derive.
The four aspects as a reading procedure
Read the overall shape
Is it reasonably symmetrical, or is there a hump on one side and a longer tail on the other? If there is, the graph is skewed — and the source is specific that the skew is named in the direction of the tail, not the hump.
Count the peaks
One clear peak, or more than one? More than one, says the source, may mean the data is a mixture from two or more sources — which is a statement about your sampling and your fieldwork, not only about the chart.
Look for points off the main pattern
Points obviously different from the main pattern are outliers. The source's reason for noticing them here is that they "often point to some set of special circumstances which could be quite valuable to understand".
Look for clusters
Are there clusters of data? The source offers the same explanation as for multiple peaks — that there may be different sources of data behind them.
Skew is defined by sight and by nothing else
The definition supplied is visual and, as far as it goes, correct and memorable: a graph with a hump on one side and a longer tail on the other "is said to be skewed in the direction of the tail". A reader who takes only that away from the week has something they can use on any chart they meet.
Peaks and clusters: two aspects, one explanation
Aspect 2 asks whether there is more than one clear peak and answers that the data may be a mixture from two or more sources. Aspect 4 asks whether there are clusters and answers that there may be different sources of data. The two aspects are given the same diagnosis and are never distinguished from each other.
The underlying instinct is the most practically valuable idea on the slide: a display with structure in it — two humps, separated groups — is telling you that the thing you measured may not be one population. In a delivery setting that is the difference between one process and two, or between one respondent group and two that should have been analysed separately.
The third aspect reverses the deck's own instruction
Here, an outlier is a point that "often point[s] to some set of special circumstances which could be quite valuable to understand" — something to investigate. Earlier in the same deck an obviously erroneous figure "will be removed from the data", and on the slide after that, wrongly including an extreme value is said to be usually regarded as a more serious error than wrongly ignoring one — an instruction that errs towards exclusion.
Three slides take three positions, and the deck never reconciles them. This page does not adjudicate between them either, because the source does not: the full record of the four characterisations and the three instructions is at Handling Outliers. What matters when you are reading a display is that the aspect you are being asked to notice here is framed as an opportunity, and the earlier framing was as a defect.
What a distribution is, and its three key features
The three definitions the slide supplies
- Distribution
- "the statistical name given to a collection of measurements taken on an aspect of interest within a particular situation"
- Bell shaped curve
- The symmetrical smooth curve typically drawn over a distribution to abstract its key properties; "the basis for the normal probability model".
- Parameter
- "represents a feature of the model such as location or spread"
The definition of a distribution is the useful one. It is deliberately plain, and it makes the point that a distribution is a property of a set of measurements, not of a chart — the chart is only how you look at it. Note the hedges in the curve statement as well: "typically" a symmetrical shape, "often" referred to as a bell-shaped curve. Preserve them if you quote it.
THE THREE KEY FEATURES AGAINST WHAT THE WEEK SUPPLIES
| Feature | What the source says about it | A way of measuring it |
|---|---|---|
| Shape | Named as a key feature and not developed on the slide. The only shape vocabulary in the week — symmetry, skew, peaks, clusters — sits on the reading slide above, presented as a way of reading a histogram rather than as an account of shape | None |
| Location | Glossed as "some measure of location of the distribution (perhaps an average or median)". The same idea is called central tendency seven slides earlier, with a different gloss | Three averages are defined elsewhere in the week; the slide names no measure of its own |
| Spread | Invoked twice on this slide — a key feature, and something you would want "some measure" of | None |
The count of three is stated by the source and the list beneath it holds three. This and the four key aspects are the only two places in the week where a stated count matches its list.
Reading a display when you have no measures to put beside it
The source does not prescribe the following. It is how to get honest value out of a reading procedure that is genuinely good and a vocabulary that is genuinely incomplete.
What to do with what each aspect tells you
A reading pass over any chart in your own report
- You have asked all four questions, in order, and written down the answer to each
- Any skew is described by the direction of its tail
- Multiple peaks or clusters have been chased back to a possible second source in the data
- Points off the main pattern have been investigated, not deleted by reflex
- Every claim about shape is stated as a description, not as a measurement
- Anything you have quantified came from a source you have cited, because this material quantifies none of it
What to carry forward
- Four questions, in order: overall shape, number of peaks, points off the pattern, clusters. It is the most immediately usable content in the week.
- Skew is named in the direction of the tail. That is the whole of what this material gives you on skew — no measure, no direction terminology, no consequence for your choice of average.
- More than one peak, or visible clusters, is a signal about your data sources before it is a fact about your chart.
- A distribution is a collection of measurements on an aspect of interest; shape, location and spread are its three key features; and the week supplies a way of measuring none of them.
- Location and central tendency are used for the same feature with incompatible glosses. Name the statistic you actually used and avoid both terms in your own writing.
Frequently asked questions
What should I look for when I read a histogram?
The source names four key aspects: the overall shape, whether there is one clear peak or more, whether any points sit obviously away from the main pattern, and whether there are clusters. It states the count as four and lists four, which is one of only two places in the week where a stated count matches its list.
Which way round is a skewed distribution named?
In the direction of the tail. The source's definition is a graph with a hump on one side and a longer tail on the other, said to be skewed in the direction of the tail. It gives no measure of skewness and does not use the terms positive, negative, left or right skew anywhere.
What does more than one peak mean?
The source says it may mean the data is a mixture of data from two or more sources. It offers the same explanation for clusters and never distinguishes the two situations, so treat both as a prompt to check whether you have combined groups that should be described separately.
What is a distribution, in this material?
"the statistical name given to a collection of measurements taken on an aspect of interest within a particular situation" The definition is plain and useful, and it makes clear that a distribution is a property of a set of measurements rather than of the chart you draw from them.
Why does the material use both location and central tendency?
It never says. One slide names location as a key feature of a distribution and glosses it as perhaps an average or median; an earlier slide defines central tendency as the most common value for variables measured at a nominal level, which is the mode. The two are the same feature with glosses pointing at different statistics, and the source reconciles them nowhere.
How do I report how spread out my data are?
Not from this material. It names spread as one of three key features of a distribution, says twice that you would be interested in some measure of it, and never names one anywhere in the week. Any measure you report has to come from, and be attributed to, a second source.
References and source attribution
- Veal, A. J. 2005, Business Research Methods: A Managerial Approach, 2nd ed., Longman. — the single entry on the week's reference slide; it is cited for the normal-distribution slide and for none of the material on this page.
- A publisher's companion website for a fourth-edition social research text, cited on the week's terminology slide as a bare address with no author, year or title stated. It is the source of the definitions of central tendency, mean, median and mode that this page compares against the distributions slide.
- The supplied teaching source: the week 11 slide deck on analysing data and the presentation of quantitative data, whose reading slide and distributions slide are reproduced in full on this page. No study notes for this week were supplied, and no chart image in the week is recoverable.
Suggested questions for Ask KEVOS
- Walk me through the four key aspects against the chart in my results chapter.
- My distribution has two peaks — what does the supplied material say that might mean?
- How should I describe a skewed distribution if I have no measure of skewness?
- What is the difference between location and central tendency in this material?
- Draft the paragraph that describes my histogram for a results chapter.
