Tracing a Figure Back to Its Source
A method for checking that a number in your write-up is still the number your data produced — built from six traces run on a real project, and from the two that came back changed.
Why a trace is possible here at all
Checking a published number normally stops at the chart, because the chart is usually all a reader gets. This project is different: the instrument, the raw response file, the analysis workbook and the thesis all survive together, so any figure in the document can be walked back to the question that produced it. The whole apparatus is described at a complete research project, end to end.
- Question on the instrument
- Stored response
- Tabulation
- Chart
- Sentence
Each arrow is a place where a number can change. Six figures were run through all five links, and the result is not the one most readers would predict.
The six traces
SIX FIGURES, TRACED FROM SENTENCE TO QUESTION
| Figure as published | What the data give | Verdict |
|---|---|---|
| "35 of the 50 respondents (70%)" chose the top two importance categories | 21 + 14 = 35; 35 ÷ 50 = 70% | Survives intact — count, sum, base and percentage all correct, and the aggregation is stated rather than assumed |
| "Only three participants (6%)" said change management is not important | 3 ÷ 50 = 6% | Survives intact, in both chapters that state it |
| Question 7 priorities: "36%, 22% and 26% … a total of 84%" | 15, 11 and 13 of 50 = 30%, 22% and 26%; total 78% (39 of 50) | Fails. One percentage wrong by six points, repeated in a second chapter, and a derived total built on it |
| Question 6 models: "36%, 22% and 22% … a combined 80%" | 18, 11 and 11 of 50 = 36%, 22%, 22%; 40 of 50 = 80% | Survives intact, including the derived total |
| Question 8 communication: "22%, 24% and 24% … 16% and 14%" | 11, 12, 12, 8 and 7 of 50; the five sum to 100% | Survives intact — the cleanest trace of the six |
| Age counts: 19, 17, 8 and 6 respondents | 19, 17, 8, 6 — totalling 50 | Numbers survive; three of four labels do not |
Every count above was recomputed from the fifty raw response rows for this extract. Verified. All figures belong to this study.
One honest limit on the method as run here. The raw response file holds only the four demographic questions, so for questions 5 to 8 — the ones the findings rest on — the chain has four links, not five. The stored responses exist in exactly one place, the first sheet of the workbook, and cannot be checked against an independent export.
The trace that fails, link by link
Question 7: top priority among five components of change management
The question
A required single-choice item with five options — leadership alignment, stakeholder engagement, employee communication and training, change readiness, organisational structure and design. One answer only, no ranking.
The stored responses
Fifty values in the dataset column, no blanks. Nothing ambiguous, nothing to interpret.
The tabulation
Stakeholder engagement 15, leadership alignment 13, employee communication and training 11, change readiness 6, organisational structure and design 5. Totals 50. Recomputed from the raw rows and correct.
The chart
A pie with percentage labels reading 30%, 26%, 22%, 12% and 10%. Every slice is right and the five sum to 100.
The sentence
The results chapter states the proportions identifying the three leading factors as "36%, 22% and 26%, respectively". Two are right. The first should be 30 per cent.
The repeat, and the total
The discussion chapter restates the same three percentages "for a total of 84% of the net responses". 36 + 22 + 26 does equal 84, so the total is internally consistent with the wrong number. The correct total is 78 per cent — 15 + 13 + 11 = 39 of 50.
There is a plausible account of where 36 came from, and it is worth stating carefully because it is an inference rather than something the material records. 36 per cent is the correct figure for a different question, two paragraphs earlier — the leading option on question 6 is 18 of 50, which is exactly 36 per cent, and that paragraph is where trace 4 succeeds. The most economical reading is that a percentage was carried across from the preceding paragraph. The source offers no account at all; this one is the library's.
The second defect: right numbers, wrong labels
The age paragraph fails differently, and the way it fails is more instructive than the arithmetic failure. Four counts are given for four age bands. All four counts are correct and they total fifty. Three of the four band labels are wrong.
Among the 50 participants: 19 were in the 18-30-year age group; 17 were in the 18-30-year age group; 8 were in the 18-30-year age group; 6 were in the 18-30-year age group.
The chart beside it has a correctly labelled category axis: the four band names appear in the right order, and the bar heights are approximately 19, 17, 8 and 6. But the chart carries no data labels. Its numbers can only be estimated off a gridline — the choice examined at presenting data in graphs.
One further drift runs the other way and is worth knowing about. An enterprise-size option printed on the instrument without brackets is stored with brackets in every file, and reaches the published figure with the brackets on its axis — while the prose silently normalises it back. The label is wrong on the figure and right in the prose: the exact inverse of the age paragraph. Where labels travel between files, see building an analysis workbook.
Running the trace on your own work
The method below is the library's, generalised from the six traces above. The supplied material sets out no procedure for checking a reported figure against the tabulation that produced it, and neither does this project. It takes a few minutes per figure, needs no tooling beyond what produced the numbers, and sits alongside the guidance at writing the results section.
Five moves per published figure
Start at the sentence, not at the data
List every number that appears in your prose — percentages, counts, totals, ranges. Work from the written claim backwards. Starting at the spreadsheet only tells you the spreadsheet is right, which in this project it always was.
Name the link that produced it
For each number, write down which tabulation it came from and which cell. If you cannot name the cell, that is the finding: the number has no traceable source and needs one before it is published.
Re-derive it from count and base
Recompute the percentage rather than reading it. 15 of 50 is 30 per cent whatever the sentence says. Do this even when — especially when — you are certain.
Check the base is the base you claim
Percentages of subsets are where bases drift. State the base beside the figure at least once: "39 of 50" travels better than "78 per cent" and cannot be silently rebased.
Check every other place the number appears
Search the document for the figure. Compare each occurrence against the tabulation, not against the first occurrence — otherwise you are checking a copy against its own copy, which is how this project's discussion chapter inherited an error.
The check that catches each defect, and where it belongs
SIX DEFECTS, THE CHECK THAT CATCHES EACH, AND WHEN TO RUN IT
| Defect | The check | Where it belongs |
|---|---|---|
| A percentage the counts do not produce | Re-derive from count and base, in writing, beside the sentence | As the results sentence is drafted |
| The same wrong figure repeated later | Search the document for each figure; compare every occurrence against the tabulation, never against another occurrence | After the discussion chapter is drafted |
| A derived total inheriting a wrong component | Recompute totals from components, then check the components against the base | Same pass as the check above |
| Prose contradicting the figure beside it | Read each figure and its adjacent paragraph as a pair, out of sequence from the rest of the chapter | Final read of the results chapter |
| Counts attached to the wrong category labels | Check labels and values as separate objects — read the labels down the list against the instrument's own option order | Same pass as the check above |
| A chart whose values cannot be read off it | Turn on data labels wherever the prose relies on the values | When the chart is made, not when the chapter is written |
This library's method, generalised from six traces on one project. No step of it is prescribed by the supplied material.
What a trace cannot tell you
- It cannot check what was never asked. Every figure here is a correct count of a question that was put. The trace says nothing about whether the question measured the thing the hypotheses were about — in this project, it did not.
- It cannot check a number with no surviving source. Where the raw export holds only some of the questions, the trace shortens by a link and the dataset becomes its own authority.
- It cannot see a scale problem. One item here places "Important" below "Fairly Important", so its middle categories cannot be ordered. The counts are still right, the trace still passes, and the measure is still weak — see the survey instrument, displayed.
- It cannot generalise. Five of six figures surviving is this thesis's result. It is not a rate, a benchmark or a statement about examined work — those findings are collected at research integrity in examined work.
The register matters as much as the method. This thesis passed examination, and every defect above is a defect in a document rather than in the person who wrote it. Two of them are the ordinary consequence of moving a verified number by hand, at the end of a long project, into a sentence. The value of the trace is that it finds exactly that class of error — which is the class a spreadsheet audit never sees.
What to carry forward
- Trace backwards from the sentence. Starting at the data only confirms the data, and in this project the data were never the problem.
- Re-derive every percentage from its count and base, including the ones you are sure of. 15 of 50 is 30 per cent however the sentence reads.
- Check repeated figures against the tabulation, never against their earlier appearance — that is how one error became three statements and a wrong total.
- Label the values on any chart whose numbers your prose relies on. When the chart is unlabelled and the prose is wrong, nothing in the document can correct it.
- Print the base. "39 of 50" survives being moved between chapters in a way that "78 per cent" does not.
Frequently asked questions
Does one failed trace mean the research is unreliable?
No, and the distinction is the point. The instrument was fielded, the responses were captured completely, all eight tabulations are arithmetically correct and every chart matches its tabulation. What failed was the transcription of one verified number into a sentence, and its repetition. That is a defect in the write-up of one project, not a verdict on the study or on the person who conducted it.
How do I know 36 per cent came from the paragraph above?
You do not know it, and this page does not claim you do. What is verifiable is that 36 per cent is the correct figure for the leading option on the previous question, that the figure appears two paragraphs earlier, and that 15 of 50 is 30 per cent. The account of how the two got swapped is the library's reading; the material offers no explanation.
Is five out of six a normal survival rate for published figures?
It is not a rate at all. Six figures from one thesis were traced because that thesis published everything needed to trace them. No comparison set exists, and the result cannot be projected onto other work — particularly work that publishes less, where the trace cannot be run.
Why does an unlabelled chart matter so much?
Because a labelled chart is a second, independent record of the values. When the prose carries the numbers and the figure carries only the categories, an error in the prose has nothing to correct it. In this project the two defects compounded precisely that way: correct axis, no values, and four counts attached to one label.
Where in a writing process should the trace sit?
Split it. Re-derive each percentage as the results sentence is drafted, and run the repeated-figure and figure-versus-prose checks after the discussion chapter exists, since that is where copies are made. Data labels are a decision at chart-making time, not at write-up time.
What if I cannot find the source of one of my own figures?
Treat it as a finding rather than an inconvenience. A number you cannot walk back to a cell is not yet checkable by anyone, including your examiner. Rebuild it from the stored responses before you decide whether it was right, and keep the tabulation it came from with the draft.
References and source attribution
- An examined master's thesis supplied as student work, with sixteen figures and the survey findings traced on this page. Researcher, supervisor, institution and protocol scrubbed. Used as observed practice, not as a model answer.
- The same project's survey instrument, raw response file and ten-sheet analysis workbook, from which all counts on this page were recomputed and verified.
- The supplied teaching source: weekly study notes, slide decks and assessment activities for a master's-level research methods subject in project management. Author, institution and year not stated in the supplied files.
Suggested questions for Ask KEVOS
- How do I check that a percentage in my thesis matches my data?
- What is the quickest way to find transcription errors in a results chapter?
- Should charts in a thesis carry data labels?
- How do I stop an error in my results chapter from spreading into my discussion?
- What should I keep so that someone else can verify my published figures?
