Loading...
AQA provides a that you are expected to have worked with before the exam. You do not take it in with you and you are not tested on remembering values — you are tested on being with its structure, its quirks and what conclusions it can and cannot support.
The big picture
The reason a real data set is in the specification is that clean textbook data teaches the wrong habits. Real data has gaps, inconsistent units, values recorded to different precisions, and variables that look related but are not. Meeting those problems in advance changes how you read an exam question: when a table has a blank cell you know that is a design feature rather than a printing error, and when a question asks whether a conclusion is justified you have a stock of concrete reasons to draw on. The marks come from that judgement, not from recall.
What you'll be able to do
You will not be asked to recall a number. Questions instead present an extract from the data set and ask you to calculate, interpret or criticise — and they are written assuming you already know what the variables mean and how the data behaves.
That familiarity shows up in small ways that earn marks: knowing which variables are categorical rather than numerical, knowing that some measurements are recorded only at certain locations or in certain months, and knowing which quantities are physically capable of being zero.
The practical consequence is that you should have spent time actually opening the data, not just read about it.
A question that seems to be missing information is often testing whether you know that the missing information does not exist in the data set. Recognising a genuine gap is itself the answer.
Real data sets contain blanks, and they are not all the same. A blank may mean the quantity was not measured, that the instrument failed, or that the value is genuinely zero and was left empty — and those are very different things.
The safe treatment is to exclude missing values from a calculation and , along with the reduced sample size. Substituting zero for a missing value is almost always wrong: it turns "we do not know" into a definite claim and drags the mean down.
If a large proportion of one variable is missing, that is itself worth commenting on, because it may mean the remaining values are not representative.
Tip — Dividing by the original sample size after removing blanks is a standard lost mark. The denominator must be the number of values you actually used.
Different variables in the data set are recorded in different units and to different precisions, and mixing them is a common error. A quantity in millimetres and one in metres cannot be compared or combined without conversion.
Precision matters for interpretation too. A variable recorded to the nearest whole unit cannot support a conclusion about differences of a tenth, and an answer quoted to more significant figures than the data justifies is misleading.
Distinguish variables (location, type, category) from ones. A mean of a categorical variable is meaningless, and questions sometimes offer exactly that as a trap.
Ask whether the arithmetic would still make sense if the categories were relabelled with letters. If replacing 1, 2, 3 with A, B, C destroys the calculation, the calculation was never legitimate.
Exam extracts are sub-samples — one location, one month, one year. A conclusion drawn from a sub-sample applies to that sub-sample first, and generalising further needs justification.
The standard criticisms are worth having ready: the period may be atypical, one location may not represent others, the sample may be too small to support the claim, and a relationship in the data does not establish that one variable causes the other.
Correlation and causation is examined explicitly. Two variables can move together because both respond to a third — a variable — or by coincidence in a small sample.
Tip — When asked to criticise a conclusion, attack the first — period, place, sample size — then the logic. Both are usually available marks.
Think like an examiner
Common misconceptions
Working with real data
Stretch yourself
A student uses one month of data from a single location to claim that a particular measurement is rising over time. The month contains 31 readings, of which 6 are missing. Give three separate reasons why the claim is not well supported.
Hint — Consider the time span, the geography, and what the missing readings might do to a trend.
Questions students ask
Key takeaways
How this fits the course
Test yourself
Ready to lock in The Large Data Set? Pick a mode and earn XP & Dobloons.
Real past-paper questions on The Large Data Set, marked mark-by-mark. How you do feeds straight into your weak-topic list, so your revision keeps targeting what actually needs work.