Loading...
A measure of location answers "where is the middle?" — but there are three sensible answers, and they can differ sharply on the same data. Knowing to use, and being able to justify it, is worth more marks than the arithmetic.
The big picture
The mean uses every value, which is its strength and its weakness: it is efficient, it feeds into everything later (variance, the normal distribution, hypothesis tests on a mean), and a single extreme value can drag it somewhere unrepresentative. The median ignores everything except position, so an outlier cannot move it far — which is exactly why house prices and salaries are reported as medians. Recognising which of these behaviours a data set needs is the judgement being tested, and it recurs the moment you meet skewed distributions later in the chapter.
What you'll be able to do
The is the total divided by how many values there are. It uses every data point, which makes it sensitive to extremes.
The is the middle value once the data is ordered. With values it sits at position ; for even that lands between two values, so take their mean. Because it depends only on position, extremes barely affect it.
The is the most frequent value. It is the only average that works for categorical data, but a data set can have several modes or none at all.
In that example the mean of 14 is larger than four of the five values. The single 43 has pulled it clear of the data — a compact demonstration of why an outlier makes the mean a poor summary.
Grouped data tells you how many values fall in each class but not what they were. The standard assumption is that values are evenly spread within a class, so each class is represented by its .
The result is an , and the word matters — questions ask for "an estimate of the mean" and expect you to know why it cannot be exact.
Class boundaries need care with continuous data. A class recorded as – for a quantity measured to the nearest whole number actually spans to , so its midpoint is , not by accident — check the boundaries before halving.
Tip — Always call it an estimate and say why: the original values are unknown, so midpoints are assumed to represent each class.
For grouped data the median is found by linear interpolation. Locate the class containing the th value using cumulative frequency, then assume the values are evenly spread across that class and step the appropriate fraction of the way into it.
Note the position rule changes: for grouped data use rather than , because you are locating a position on a continuous scale rather than picking out a listed value.
Interpolation assumes an even spread inside the class, which is the same assumption the midpoint method makes for the mean. Both estimates inherit whatever error that assumption introduces.
Use the when the data is roughly symmetric with no extreme values, and when the result feeds into later calculations. Use the when the data is skewed or contains outliers. Use the for categorical data or when the most common value is what matters — a shop ordering shoe sizes wants the mode, not the mean.
simplifies awkward arithmetic. If , then — the mean transforms exactly as the data does, because it is a linear operation.
Tip — A justification must name the feature of the data. "The median, because the data contains an outlier at 43 which would distort the mean" scores; "the median, because it is better" does not.
Think like an examiner
Common misconceptions
Location
Stretch yourself
Nine employees earn (in £000s): 22, 24, 24, 26, 27, 28, 30, 31, 168. Calculate the mean and median, state which better represents typical pay, and justify your choice.
Hint — Compute both, then look at how many employees actually earn near each value.
Questions students ask
Key takeaways
How this fits the course
Test yourself
Ready to lock in Measures of Location? Pick a mode and earn XP & Dobloons.
Real past-paper questions on Measures of Location, marked mark-by-mark. How you do feeds straight into your weak-topic list, so your revision keeps targeting what actually needs work.