Averages, Measures of Spread and Outliers
Three averages, and the marks are for choosing the right one. Plus quartiles, the interquartile range, and how to prove a value is an outlier rather than just calling it one.
Get the method right under pressure
Free interactive practice on the steps that lose marks under exam pressure.
Start revising freeWhat you'll cover
Three averages, one decision
Most people can calculate a mean. Far fewer can say why the mean was the wrong thing to calculate, and that is where the interpretation marks live. An average is a choice. One dataset can honestly produce three different "averages", and picking the one that suits the data is a skill this specification tests directly. So this module does the arithmetic and the judgement together, on one small set of numbers you will get to know well.
What each average is good for
Know the definition and the reason you would reach for it:
Match each situation to the average you would use
- The favourite colour of 200 pupils
- Salaries in a company where one director earns far more than anyone else
- The heights of 30 pupils, with no unusual values
- Yearly growth rates that multiply on each other [Higher]
- the mode, since the data are categories
- the median, since one extreme value would distort the mean
- the mean, since it uses every value and none is extreme
- the geometric mean
Find the median
Seven pupils recorded how many books they read last month: 7, 3, 9, 4, 12, 6, 5. What is the median number of books?
Two ways to measure spread
Both describe how spread out the data are. They disagree about what to do with the extremes, and that disagreement is usually the point of the question.
Find the interquartile range
The same seven values: 7, 3, 9, 4, 12, 6, 5. Put them in order, find the lower quartile and the upper quartile, then work out the interquartile range.
Which tier needs what
This sub-section splits sharply between the tiers, so it is worth knowing which side of the line you are on. Both tiers: mode, median, mean, range, quartiles, percentiles, the interquartile range, and spotting an outlier by looking at the data. Higher tier only: weighted mean, geometric mean, interpercentile and interdecile range, standard deviation, identifying an outlier by calculation, and the standardised score, (x - mu) / sigma. One warning about that last one. The standardised score formula is not given in the assessment, so it has to be memorised. Most of the formulae you meet in this subject are provided; this is one that is not.
Why the median?
One pupil in that group actually read 60 books, not 12. What happens to the averages?
- The mean rises noticeably while the median stays where it is
- Both the mean and the median rise by the same amount
- The median rises and the mean is unaffected
- Neither changes, because it is only one value out of seven
The outlier boundary [Higher tier only]
Higher tier. For those seven values the lower quartile is 4, the upper quartile is 9 and the interquartile range is 5. Using the rule that an outlier lies more than 1.5 times the interquartile range beyond a quartile, what is the boundary above which a value counts as an outlier?
Which are true?
Select the TWO correct statements about outliers and averages.
- An outlier should be identified by a test, not just because a value looks unusual
- The interquartile range is unaffected by an outlier, while the range is not
- An outlier should always be deleted from the data
- An outlier changes the mode more than the mean
Choose the measure
Three datasets. Choose the measure that reports each one honestly.
- House prices on a street where one house sold for four times the price of any other. Which average should the report use?
- You want to describe how spread out exam marks are, and one pupil was absent for most of the course and scored very low. Which measure of spread?
- A value falls just outside your calculated outlier boundary. What should you do with it?
Order the method
Put the stages of finding the median and interquartile range of a list into a reliable order.
- Write the values in order, smallest first
- Count how many values there are
- Find the median: the middle value, or halfway between the two middle ones
- Split the data at the median into a lower half and an upper half
- Find Q1 as the median of the lower half, and Q3 as the median of the upper half
- Subtract to get the interquartile range: Q3 minus Q1
Mode, median or mean
The _____ is the only average that works for categories such as favourite colour. The _____ is the middle value once the data are in order, and one extreme value does not move it. The range is the largest minus the smallest, while the _____ range describes only the middle half, so an outlier cannot inflate it. At Higher tier an outlier can be identified by calculation, using _____ times the interquartile range beyond a quartile. And whatever the question, the first thing you do with a list is put it in _____.
Challenge the 480,000 pound headline
A local newspaper reports that "the average house price on Mill Street is 480,000 pounds", based on the mean of eight sales, one of which was a large house that sold for far more than the rest. Write a response explaining what is misleading and what should have been reported.
- Explain why the mean is the wrong average for this data, referring to the extreme sale
- Say which average should have been used instead, and why it reports the street more honestly
- Explain which measure of spread you would quote alongside it, and why not the range
- Say how you would show that the expensive sale really is an outlier rather than just unusual