Loading...
A well-chosen diagram makes the shape of a data set visible in a way a table never does. Three appear in the specification — histograms, box plots and cumulative frequency curves — and each answers a different question about the data.
The big picture
The one idea worth carrying away is that in a histogram it is , not height, that represents frequency. That is not a technicality invented to catch people out: it is what allows classes of different widths to be compared honestly on the same axes. Plotting frequency as height when the classes are unequal exaggerates the wide ones and is genuinely misleading — the same trick that appears in badly drawn newspaper graphics. Once frequency density is second nature, histograms stop being fiddly and the rest of the chapter, including estimating from a diagram, follows straightforwardly.
What you'll be able to do
A histogram represents frequency by . When class widths differ, plotting raw frequency as height would make a class twice as wide look twice as important for the same count — so the vertical axis carries instead.
Bars touch, because the underlying variable is continuous. That is the visible difference from a bar chart, which has gaps and is for categorical data.
To read a frequency back off a histogram, multiply the height by the width. To estimate the frequency in part of a class, take the corresponding fraction of that bar’s area — which assumes an even spread within the class, the same assumption used for grouped means.
Tip — Label the vertical axis "frequency density", not "frequency". Mislabelling loses a mark even when the bars are drawn correctly.
A box plot shows a five-figure summary: minimum, , median, , maximum. The box spans the interquartile range with the median marked inside it, and whiskers reach out to the extremes.
An is conventionally any value more than beyond a quartile. Outliers are plotted as separate points and the whisker is drawn only as far as the most extreme value that is not an outlier.
Because a box plot compresses a distribution into five numbers, two of them side by side make comparison very quick — which is exactly what exam questions ask for.
The is a convention, not a law of nature. Some questions specify a different multiplier, or define an outlier as more than two standard deviations from the mean — always use the rule the question gives rather than the one you remember.
A cumulative frequency curve plots the running total of frequencies against the — not the midpoint, which is the standard error here. Points are joined with a smooth curve.
Reading across from a cumulative frequency and down to the axis estimates a value: use for the median, for , for , and for the th percentile.
Reading the other way — up from a value and across — estimates how many observations fall below it, which is how "how many scored under 50?" questions are answered.
Tip — Plot cumulative frequency against the boundary of each class. Using midpoints shifts the whole curve and every estimate taken from it.
Diagrams reveal , which the averages alone can suggest but not confirm. In a positively skewed distribution the tail stretches to the right, and the mean is pulled above the median. Negative skew is the mirror image, with the tail and the mean to the left.
On a box plot, skew shows as an off-centre median: closer to means positive skew, closer to means negative. A roughly symmetric distribution has the median near the middle of the box with whiskers of similar length.
The quartile test states this numerically, and it is the version to quote when a question asks you to justify the skew rather than describe it.
Positive skew is common in real data bounded below by zero — incomes, waiting times, rainfall. There is a floor but no ceiling, so the tail can only run one way.
Think like an examiner
Common misconceptions
Diagrams
Stretch yourself
A histogram has classes (frequency density 1.2), (3.5) and (0.8). Find the total frequency, and estimate how many values lie between 15 and 35.
Hint — Frequency is area. For the partial classes, take the matching fraction of each bar.
Questions students ask
Key takeaways
How this fits the course
Test yourself
Ready to lock in Representing Data? Pick a mode and earn XP & Dobloons.
Real past-paper questions on Representing Data, marked mark-by-mark. How you do feeds straight into your weak-topic list, so your revision keeps targeting what actually needs work.