Measures of Dispersion

QMI1500 - Elementary Quantitative Methods · Descriptive Statistics

Measures of Dispersion

Measures of dispersion describe the spread or variability of a dataset. They provide insights into how much the data points differ from each other and from the central measure, such as the mean or median. Understanding dispersion helps you to assess the reliability and consistency of your data.

Range

The range is the simplest measure of dispersion. It is calculated by subtracting the smallest value in the dataset from the largest value.

Formula:

Range = Maximum value - Minimum value

Example

Consider the following dataset representing the ages of five employees at a company: 25, 30, 35, 40, 50.

  1. Identify the maximum value: 50
  2. Identify the minimum value: 25
  3. Calculate the range: 50 - 25 = 25

The range of this dataset is 25 years.

Watch out: The range does not provide information about the distribution of values between the minimum and maximum. It can be misleading if there are outliers (values significantly higher or lower than the rest).

Variance

Variance measures how far each number in the dataset is from the mean (average) and thus from every other number in the dataset. It is calculated as the average of the squared differences from the mean.

Formula:

Variance (σ²) = Σ (xi - μ)² / N

Where:

  • σ² = variance
  • Σ = summation symbol (sum of all values)
  • xi = each individual value in the dataset
  • μ = mean of the dataset
  • N = number of values in the dataset

Example

Using the same dataset of ages: 25, 30, 35, 40, 50.

  1. Calculate the mean:
Mean (μ) = (25 + 30 + 35 + 40 + 50) / 5 = 180 / 5 = 36
  1. Calculate each squared difference from the mean:
(25 - 36)² = 121
(30 - 36)² = 36
(35 - 36)² = 1
(40 - 36)² = 16
(50 - 36)² = 196
  1. Add the squared differences:
121 + 36 + 1 + 16 + 196 = 370
  1. Divide by the number of values (N = 5):
Variance (σ²) = 370 / 5 = 74

The variance of the dataset is 74.

Watch out: Variance is expressed in squared units. This can make interpretation difficult. For example, if you are measuring ages, variance will be in years².

Standard Deviation

The standard deviation is the square root of the variance. It provides a measure of dispersion in the same units as the original data, making it easier to interpret.

Formula:

Standard Deviation (σ) = √Variance

Example

Continuing with our previous example:

Standard Deviation (σ) = √74 ≈ 8.6

The standard deviation of the dataset is approximately 8.6 years.

Tip: A low standard deviation indicates that the data points are close to the mean, while a high standard deviation indicates that the data points are spread out over a larger range of values.

Interquartile Range (IQR)

The interquartile range (IQR) measures the spread of the middle 50% of the data. It is calculated by subtracting the first quartile (Q1) from the third quartile (Q3).

Formula:

IQR = Q3 - Q1

Quartiles are values that divide the dataset into four equal parts. Q1 is the median of the lower half of the dataset, and Q3 is the median of the upper half.

Example

Using the dataset: 25, 30, 35, 40, 50, we first need to find Q1 and Q3.

  1. Order the data (already ordered): 25, 30, 35, 40, 50.
  2. Identify Q1 (the median of the first half: 25, 30): Q1 = 27.5.
  3. Identify Q3 (the median of the second half: 40, 50): Q3 = 45.
  4. Calculate the IQR:
IQR = Q3 - Q1 = 45 - 27.5 = 17.5

The interquartile range of the dataset is 17.5 years.

Watch out: The IQR is less affected by outliers than the range. It focuses only on the central portion of the data.

Conclusion

Measures of dispersion are essential for understanding the variability in a dataset. The range, variance, standard deviation, and interquartile range each provide different insights into the spread of data. When analysing data, consider using multiple measures to gain a comprehensive view of its dispersion.

Summary

  • Range is the difference between the maximum and minimum values.
  • Variance measures the average of the squared differences from the mean.
  • Standard deviation is the square root of the variance.
  • Interquartile range measures the spread of the middle 50% of the data.

Check your understanding

  1. What is the formula for calculating the range?
  2. How do you calculate the variance of a dataset?
  3. What does a low standard deviation indicate about a dataset?
  4. Why is the interquartile range preferred in some cases over the range?