Loading...
A good diagram shows the shape of a data set in a way no single statistic can: whether it is symmetric or skewed, where it clusters, whether two variables move together. A-Level focuses on the diagrams that carry real information — histograms with unequal class widths, cumulative frequency graphs, box plots and scatter diagrams — and on reading them critically, starting with how the data were collected.
The big picture
Every statistical conclusion is only as good as the data behind it, which is why this topic starts with sampling: a biased sample makes even perfect calculations meaningless. From there, each diagram answers a particular question. A histogram shows a distribution’s shape, with area — not height — representing frequency. A cumulative frequency graph gives medians, quartiles and percentiles for large grouped data. Box plots make comparisons between groups fast and highlight outliers. Scatter diagrams show relationships between two variables, and the language of correlation — with its crucial caveat that correlation does not imply causation — carries straight into the regression and hypothesis tests later in the course. OCR regularly sets these questions in the context of its large data set, so interpretation in context is always part of the answer.
What you'll be able to do
A is the whole group of interest; a measures all of it, while a measures part of it. Samples are cheaper and quicker, but may not represent the population.
gives every member an equal chance of selection. sampling takes every th item from a list. sampling samples from each subgroup in proportion to its size. sampling fills targets for each group without random selection. sampling uses whoever is available.
Random methods reduce bias but need a sampling frame (a full list). Quota and opportunity sampling are convenient but prone to bias.
Tip — When criticising a sampling method, name the specific source of bias in context — "only people at the station at 8am are sampled, so non-commuters are excluded".
In a histogram of grouped continuous data, bars touch and . With unequal class widths, heights must be , not frequency.
Frequency density . To read a frequency back, multiply density by width — or, if the scale is not labelled, work out how many units of frequency one unit of area represents.
Histograms reveal shape: symmetric, positively skewed (a long tail to the right) or negatively skewed (a long tail to the left).
A graph plots running totals against the , joined by a smooth curve or straight lines. Read the median at , quartiles at and , and any percentile similarly.
A shows the minimum, , median, and maximum on a scale. Outliers are marked separately with crosses, and the whiskers then extend to the most extreme non-outlier values.
Box plots are ideal for comparing groups. The position of the median within the box, and the lengths of the whiskers, indicate skew.
When the median sits closer to than , the upper half of the middle 50% is more spread out — the data are positively skewed. Box plots let you see skew at a glance.
A plots paired data to show whether two variables are related. : tends to increase with . : tends to decrease. No clear pattern: no correlation.
The (independent) variable goes on the -axis and the (dependent) variable on the -axis.
: ice-cream sales and drownings are positively correlated because both rise in hot weather.
A regression line can be used to predict from within the range of the data (). Predicting outside that range () is unreliable.
Tip — Interpret a regression gradient as "for each one-unit increase in , changes by on average", with the actual units.
Think like an examiner
Common misconceptions
Representing data
Stretch yourself
In a histogram of reaction times, the bar for seconds is 4 cm wide and 6 cm tall and represents 36 people. The bar for seconds is 3.2 cm tall. How many people does it represent, and what is its frequency density?
Hint — Find how many people 1 cm² represents from the first bar. The horizontal scale is 4 cm per 0.05 s.
Questions students ask
Key takeaways
How this fits the course
Test yourself
Ready to lock in Representing Data? Pick a mode and earn XP & Dobloons.
Real past-paper questions on Representing Data, marked mark-by-mark. How you do feeds straight into your weak-topic list, so your revision keeps targeting what actually needs work.