Loading...
An average summarises a whole data set with one representative value. You already know the mean, median and mode; A-Level adds the notation to calculate them efficiently from frequency tables and grouped data, the technique of interpolation for grouped medians, and — most importantly — the judgement to say which average suits a particular data set and why.
The big picture
Statistics questions at A-Level are rarely just calculations. OCR expects you to interpret: which average is appropriate, how an outlier distorts the mean but not the median, and what a summary statistic says about a real context such as the large data set. The mean uses every value and has the best mathematical properties, which is why it underpins variance, the normal distribution and hypothesis testing later on. The median is resistant to extreme values. The mode is the only average for categorical data. Sigma notation, , is the language for all of it, and coding shows how averages respond when data are shifted or rescaled.
What you'll be able to do
The is : the sum of the values divided by how many there are.
The is the middle value when the data are ordered. For values it is the th value; if is even, average the two middle values.
The is the most frequent value. Grouped data have a instead. Data can be bimodal, or have no mode.
When values repeat, a frequency table records each value with its frequency . The mean is .
For the median, find its position (for discrete data) and use cumulative frequencies to locate the value.
With grouped data the exact values are lost, so averages are . For the mean, use the class midpoint as in .
For the median of continuous grouped data, find the th position and use within the class that contains it, assuming values are spread evenly across the class.
Pay attention to class boundaries. Ages "20–29" in completed years run from 20 up to (but not including) 30.
Tip — State that grouped values are estimates, and why: "the midpoint is used because the exact values within each class are unknown".
If data are coded by , then . Adding a constant shifts the mean; multiplying scales it. The median and mode transform the same way.
An pulls the mean towards it but hardly affects the median. So for skewed data or data with outliers, the median is usually more representative.
The mean is preferred for roughly symmetric data because it uses every value. The mode is the only average possible for categorical (non-numerical) data, such as method of travel.
Average income is a classic trap: a few very high earners drag the mean well above what a typical person earns. Reported "typical" incomes normally use the median for exactly this reason.
Think like an examiner
Common misconceptions
Averages
Stretch yourself
The mean mass of 12 parcels is 3.5 kg. Two more parcels are added and the mean of all 14 becomes 3.8 kg. One of the new parcels weighs 5.1 kg. Find the mass of the other, and explain whether it is likely to be an outlier.
Hint — Convert each mean into a total mass.
Questions students ask
Key takeaways
How this fits the course
Test yourself
Ready to lock in Measures of Central Tendency? Pick a mode and earn XP & Dobloons.
Real past-paper questions on Measures of Central Tendency, marked mark-by-mark. How you do feeds straight into your weak-topic list, so your revision keeps targeting what actually needs work.