DoRevision

Measures of Spread and Identifying Outliers

Range, quartiles, IQR, percentile ranges and standard deviation, and the outlier rule you have to carry into the exam when the harder-looking formula beside it will be printed for you. Why the median belongs with the IQR and the mean belongs with the standard deviation.

⏱️ 26 min 🎯 15 activities
Best used for
Intervention Mock preparation Cover lesson

Get the method right under pressure

Free interactive practice on the steps that lose marks under exam pressure.

Start revising free

What you'll cover

The line runs through this topic

An average tells you where a data set sits. A measure of spread tells you how tightly it is packed around that point, and two sets can share an identical mean while one is calm and the other is wild. This module covers the spread measures AQA asks for and the rules for deciding that a value is an outlier. Before any of it, one piece of exam technique that is worth marks on its own. Some formulae in this section will be printed in the question and some you must carry in your head, and the split is not the one you would guess. The most forbidding formula here, standard deviation, is given to you. A plain-looking rule about quartiles is not. Students lose marks in both directions: deriving something that was already on the page, and freezing over a rule nobody was going to print.

Five ways to measure spread

Each of these answers how spread out are the values, and each has a tier. Treat the labels as part of the content: a Foundation student does not need the last two.

Match the measure to its job

  • Range
  • Interquartile range
  • Interdecile range
  • Standard deviation
  • Uses only the largest and smallest values
  • Describes the middle half and ignores both tails
  • Trims a tenth off each end before measuring
  • Uses every value's distance from the mean

Find the range

Eleven students recorded the minutes they revised on one evening. In order: 12, 14, 15, 15, 16, 18, 19, 20, 21, 22, 48. Work out the range, in minutes.

Working out the quartiles

The same eleven values, already in order: 12, 14, 15, 15, 16, 18, 19, 20, 21, 22, 48. With eleven values the quartiles land neatly on data points, so there is nothing to interpolate. Take the position (n + 1) divided by 4, which is 12 divided by 4, giving 3. The lower quartile is therefore the 3rd value, which is 15. The median sits at position (n + 1) divided by 2, the 6th value, which is 18. The upper quartile sits at 3 times (n + 1) divided by 4, the 9th value, which is 21. You can check that against the other method you may have been taught: split the data either side of the median, and take the median of each half. The lower half is 12, 14, 15, 15, 16, whose middle value is 15; the upper half is 19, 20, 21, 22, 48, whose middle value is 21. Both methods agree, which is why a set of this size is a safe one to practise on.

Find the interquartile range

Higher tier. Using the quartiles just found for 12, 14, 15, 15, 16, 18, 19, 20, 21, 22, 48, where the lower quartile is 15 and the upper quartile is 21, work out the interquartile range, in minutes.

The formula trap

Here is the split for this section, and it is worth learning as carefully as the statistics itself. You must MEMORISE: the range, the interquartile range, the interpercentile and interdecile ranges, and the outlier rules. You will be GIVEN, printed in the question: the standard deviation formula, and the skewness formula that uses it. Read that again, because it runs the opposite way to instinct. Standard deviation looks like the hardest thing on the page and it is handed to you. The outlier rule is one short line about quartiles and nobody will print it. So do not spend revision time drilling the standard deviation formula into memory, and do not walk into the exam assuming the outlier rule will appear when you need it.

Which must you memorise?

Select the TWO formulae you must carry into the exam because they will NOT be printed in the question.

  • The outlier rule using 1.5 times the interquartile range
  • Interquartile range = upper quartile minus lower quartile
  • The standard deviation formula
  • The skewness formula, 3 times mean minus median, divided by standard deviation

Find the upper fence

Higher tier. A large outlier is any value greater than the upper quartile plus 1.5 times the interquartile range. For our data the upper quartile is 21 and the interquartile range is 6. Work out the value a reading must exceed to count as a large outlier.

Is it an outlier?

The fences for our data are 6 and 30. Which statement about the eleven values 12, 14, 15, 15, 16, 18, 19, 20, 21, 22, 48 is correct?

  • 48 is an outlier and 12 is not
  • Both 48 and 12 are outliers, since they are the two extremes
  • Neither is an outlier, because all the values are plausible revision times
  • 48 is an outlier because it is more than double the mean

What the outlier moved

Suppose the 48 was a recording error and the true value was 26. Nothing else changes and there are still eleven values. Look at which measures notice and which do not.

Which claim is true?

Based on the two columns you have just compared, select the ONE statement that is correct.

  • The median and the interquartile range did not move at all, because neither depends on how extreme the extreme values are.
  • The interquartile range did not move because the interquartile range never changes when data changes.
  • The standard deviation fell because there were fewer values in the second set.
  • The mean is the better summary here because it moved, which shows it is more sensitive to the data.

Pairing the average with the spread

When you compare two data sets, pair the average with a matching measure of spread. The median goes with the _____, because both are found by _____ and neither is dragged by an extreme value. The mean goes with the _____, because both are calculated from _____ value in the set. If a set contains an outlier you have decided to keep, the _____ pairing describes it more honestly.

interquartile range position standard deviation every median and IQR range mode mean and standard deviation counting one

Compare two data sets

A question gives you two sets of journey times, one shown as a table and one as a box plot, and asks you to compare them. Work through the decisions.

  • You notice the table has one very large value. Which pair of measures should you compare on?
  • The box plot gives you quartiles directly. What does that let you do?
  • You want to check whether the large value is formally an outlier. What do you need?
  • You confirm it is an outlier. The question does not tell you to remove it. What do you write?

Justify your choice

A charity records how many minutes each of its volunteers worked last Saturday. Most worked between two and four hours, but one volunteer stayed all day. The charity wants to publish a single summary of a typical shift, with a measure of spread. Write your recommendation.

  • Say which average and which measure of spread you would publish, as a pair
  • Explain why those two belong together, in terms of how each is calculated
  • Explain what the long shift would do to the mean, the range and the standard deviation
  • Explain how you would test formally whether that shift counts as an outlier, and say which part of that test you must remember rather than being given
  • Say whether you would remove the value, and justify your decision either way