What your people measures are really measuring: individuals, systems and honest records

Most workforce measures look only at employees, rarely at managers or the system they work in. How to test each measure against a decision, add system measures and keep problem logs honest.

Ask a growing business for its monthly people report and it will hand over twenty or more numbers: turnover, absence, overtime, time to hire, training hours, tasks completed, complaints. Ask which of them changed a decision in the past year and the answer is usually two or three, followed by a pause.

That pause is not a sign of a weak manager. It is what happens when measures are added one at a time, each sensible when proposed, and almost never removed. Look more closely at a typical list and a second pattern appears. Nearly every measure is about the employee. Turnover is counted against the people who left, not the conditions they left. Task completion is counted against the person doing the task, not whoever planned it, sequenced it and cleared the obstacles. Hardly anything measures managers, decisions or the systems people work within.

This article explains what measures cost, why they act as instructions, why a set that only looks at individuals points at the wrong cause, why a problem log that feeds performance reviews stops being honest, and how to build a smaller set of measures that genuinely informs decisions. It is general information. Privacy, workplace surveillance, employment and work health and safety laws can apply to what you record about staff and how you use it, so check with the Office of the Australian Information Commissioner, Fair Work, your state safety regulator or an adviser before introducing new monitoring.

Measures cost more than they appear to

Every measure has three costs, and only the first usually reaches a budget:

  • Collection: the time to gather, check and report it, often spread across supervisors who never record it as a cost.
  • Distortion: once a number is published, people move it, and the cheapest way to move a number rarely improves the thing behind it.
  • Attention: twenty numbers a month do not get twenty careful judgements. They get a scan, and the scan settles on whichever number moved most, which is often the noisiest rather than the most important.

A useful rule of thumb from quality management is to express important measures in money where possible. “Unplanned departures in the warehouse cost us about $60,000 last year in recruitment, training and overtime” is easier to act on than “warehouse retention is 78%”.

Measures are instructions

Charles Goodhart’s observation that a measure tends to stop working once it becomes a target applies well beyond economics. Most distortion is not cheating. It is people sensibly responding to what they are measured on:

  • A task completion rate rises if tasks are split into smaller pieces, with no more work done.
  • Time to hire measured from first interview hides the weeks spent finding candidates. Time from the vacancy opening is the honest clock.
  • Absence targets can push sick people to work, moving the cost somewhere less visible and potentially creating health and safety problems.

For every measure you publish, ask out loud: how could this number improve without the business improving? Then watch for that route. The be careful what you reward article covers designing objectives that do not backfire.

Look at the system, not just the person

The quality pioneer W. Edwards Deming argued that most variation in performance comes from the system people work in rather than from the people themselves. Even if that is only partly true, a set of measures that looks only at individuals will prompt actions aimed at individuals for problems they cannot control.

Add a few measures that look at managers and the system:

  • Turnover by supervisor, compared across similar teams. If one supervisor’s teams lose staff at three times the rate of another’s doing the same work, that says more than the business-wide figure.
  • Time from a problem being raised to a decision, such as an equipment request or a roster issue.
  • Late changes imposed on staff, such as roster changes with less than 48 hours’ notice.
  • Rework caused by unclear instructions, recorded at the point it happens.

These are rarely in standard lists because the lists were built to look downwards.

Problem logs must not become report cards

Many businesses keep logs of incidents, near misses, risks or customer complaints. They are valuable only if people record things honestly. Problems start when the same log feeds performance reviews, for example by counting incidents against the supervisor who reported them.

Once a record can affect someone’s standing, every entry becomes a statement about its author as well as about the business, and it is written accordingly. Entries rarely disappear entirely. Instead, they are softened: the cause is recorded as external, the impact as minor, the ownership as vague. The log stays complete and becomes much less useful.

Saying “nobody will be penalised for reporting” is not enough. People do not judge by what is said; they judge by what happened to colleagues at the last review. The practical fix has three parts:

  1. Decide that the log is for learning, not appraisal, and say so in writing.
  2. Keep its content out of performance reviews, including indirectly through summaries.
  3. Reward the behaviours instead: incidents logged promptly, follow-up actions completed, reviews held on time. These can be checked without reading the entries, and they give no reason to understate.

Expect the log to look worse for a few months after the change, as problems that were previously unsaid appear. That is a sign it is working. The bad news early article covers the wider incentives that delay bad news.

The tool you score becomes your real model

The same effect applies to how leadership and performance are assessed. A business may teach supervisors that good leadership adapts to the person and situation, then assess them with a questionnaire that produces two fixed scores. Over time, people are promoted on the scores, because a number can be compared and cited while “it depends” cannot. The model in the assessment tool becomes the business’s real model of leadership, whatever the training said.

Before adopting any assessment tool, ask what theory of performance it measures, whether that matches what you want, and which decisions its results will be allowed to inform. A tool used for development conversations does not need the same rigour as one used for promotion or pay.

Measure behaviours that lead to results

Results arrive late. By the time a poor quarter shows up, the decisions that caused it were made months ago. Performance systems are more useful when they also look at the behaviours and conditions that produce results: whether customer follow-ups happen, whether handovers are completed, whether problems are raised early. The how measures reshape strategy article covers leading and lagging measures in more depth.

Seven tests for every measure

Apply these to each measure, and consider retiring anything that fails three:

TestKeep whenRetire when
DecisionIt changes a specific decision with a named ownerIt is reported “for awareness”
ThresholdA trigger for action is written down in advanceThe answer is “it depends”
ControllabilityWhoever is measured can act on the causeIt blends causes they cannot separate
CostCollection costs less than the decision is worthNobody has estimated the cost
DistortionThe way it could be gamed is known and watchedNobody has asked
ConsistencyDefinitions are written and stable over timeDefinitions vary by system or month
Review dateThere is a date to reconsider itIt has run unchanged for years

Consistency deserves a word of its own. Two systems often disagree about something as basic as headcount, because one counts casual staff and the other does not, or one counts at the start of the month and the other at the end. Customer and staff survey scores built on different scales or bands cannot be compared across periods or with other businesses. Write a one-line definition for each measure you keep, note when it changes, and treat figures from before and after a change with caution. Comparability comes from written definitions, not from the numbers themselves.

A cap helps too: perhaps ten measures for the owner’s monthly review. Every addition then requires a removal, which forces the question of what is worth knowing.

Handle personal information carefully

As measurement becomes more detailed, it starts to record a lot about named individuals: activity tracking, feedback from colleagues, attitude ratings. Treat these as one question rather than separate initiatives: what is collected about each person, who can see it, which decisions it may inform, how long it is kept, and how the person can see, correct or challenge it. Practices such as forced ranking or automatically removing the lowest-rated group carry legal risk under unfair dismissal, general protections and discrimination law, and also produce numbers that describe the policy rather than the people. Take advice before going near them.

A worked example

This is an illustration. A commercial cleaning business with about 40 staff across 15 sites produces a monthly people report with 22 measures. The office manager spends about three days a month compiling it and each of six supervisors spends a couple of hours supplying figures: roughly 35 hours a month, or more than 50 working days a year. When the owner marks which measures changed a decision in the past year, only three have an entry.

Supervisors’ reviews include the number of incidents on their sites, taken from the incident log. Incident numbers have been falling for a year, and the owner has taken that as good news.

The owner makes four changes:

  • Fewer measures. Applying the seven tests, the report shrinks to nine, each with a threshold and an owner.
  • System measures added. Turnover is shown by supervisor across comparable sites. Two supervisors running similar sites turn out to have very different staff turnover, which leads to a conversation about rostering practices rather than about individual cleaners. A measure of days from an equipment request to a decision is added, and shows requests waiting an average of three weeks.
  • The incident log separated from reviews. Supervisors are now assessed on whether incidents are logged within 24 hours and whether follow-up actions are completed, not on how many incidents occur.
  • Money where it matters. Turnover is also reported as an estimated annual cost of recruiting and training replacements.

In the first three months, logged incidents rise by about half. Most are minor, but two reveal a recurring chemical storage problem at one site that had not been reported before. Fixing equipment requests faster reduces overtime on the affected sites. The report now takes about a day to produce.

How this applies to a small Australian business

  • Ask which measures changed a decision in the past year.
  • Estimate what your reporting costs in hours.
  • Name how each measure could be gamed, and watch for it.
  • Add measures that look at managers and the system.
  • Keep problem logs out of performance reviews, and reward reporting behaviours.
  • Check what theory any assessment tool scores before relying on it.
  • Cap the number of measures the owner reviews.
  • Take advice before monitoring individuals more closely.

Signals worth watching

  • Reports nobody can link to a decision.
  • Falling incident counts with no change in practice.
  • Measures that rise while the work does not improve.
  • Every measure pointing at employees, none at managers.
  • Promotions explained by assessment scores alone.
  • Supervisors spending hours compiling figures nobody uses.

Common mistakes

  • Adding measures without removing any.
  • Treating measurement as free.
  • Blaming individuals for problems the system creates.
  • Using problem logs as report cards.
  • Letting an assessment tool define leadership by default.
  • Collecting personal data without thinking through its use.

Frequently asked questions

Should we stop measuring individuals? No. Individual measures are useful, especially for development. Add system and manager measures so the full picture is visible.

How many measures should a small business track? Enough to make the decisions that matter. For an owner’s monthly review, around ten is often plenty.

What if supervisors resist turnover-by-supervisor figures? Present them as information about conditions, not a verdict on people. Compare similar teams and talk about causes.

How do we reward safety reporting without encouraging false reports? Reward timely, complete reporting and follow-through, and review a sample of entries for quality.

Do privacy laws apply to employee data? Some do, depending on the business and the information. Check with the OAIC or an adviser, and with Fair Work on employment matters.

Questions to ask

  • Which of our measures changed a decision last year?
  • What does our reporting cost us each month?
  • How could each measure improve without the business improving?
  • Which measures look at managers and systems, not just people?
  • Do our problem logs feed into anyone’s review?
  • What does our assessment tool actually score?

Bringing it together

People measures cost more than they appear to and act as instructions whether or not anyone intends it. Test each one against a decision, a threshold and a known way it could be gamed, and cap the set. Add measures that look at managers and the system, because most performance problems start there. Keep problem logs for learning, not appraisal, and reward the behaviour of reporting instead. Check what any assessment tool really scores before it quietly becomes your model of good performance.


Source: KEVOS notes, drawing on teaching material on workforce metrics, risk registers and appraisal, leadership assessment and performance systems, and on ideas associated with W. E. Deming, P. Crosby and C. Goodhart. Examples and figures in this article are illustrations. This article is general information, not legal advice.

Need practical engineering, manufacturing or process support? KEVOS can help move the work forward.