Measures of Dispersion

STA1505 - Statistics for Beginners · Descriptive Statistics

Measures of Dispersion

Measures of dispersion describe the spread or variability of a dataset. They provide information on how much the data points differ from the average value. Understanding dispersion is crucial for interpreting data accurately.

Range

The range is the simplest measure of dispersion. It is calculated by subtracting the smallest value in the dataset from the largest value.

Formula: Range = Maximum value - Minimum value

Example: Consider the following dataset of exam scores: 55, 70, 65, 80, 90.

  1. Identify the maximum value: 90
  2. Identify the minimum value: 55
  3. Calculate the range: 90 - 55 = 35

The range of the exam scores is 35.

Watch out: The range only considers the extreme values and does not reflect how data points cluster around the mean.

Variance

Variance measures the average squared deviation of each data point from the mean. It gives a sense of how spread out the data is around the mean.

Formula: Variance (σ²) = Σ (x - μ)² / N

Where:

  • σ² = variance
  • Σ = sum of
  • x = each data point
  • μ = mean of the dataset
  • N = number of data points

Example: Using the same dataset: 55, 70, 65, 80, 90.

  1. Calculate the mean (μ):
μ = (55 + 70 + 65 + 80 + 90) / 5 = 72
  1. Calculate each deviation from the mean:
55 - 72 = -17
70 - 72 = -2
65 - 72 = -7
80 - 72 = 8
90 - 72 = 18
  1. Square each deviation:
(-17)² = 289
(-2)² = 4
(-7)² = 49
(8)² = 64
(18)² = 324
  1. Sum the squared deviations:
289 + 4 + 49 + 64 + 324 = 730
  1. Divide by the number of data points (N = 5):
Variance = 730 / 5 = 146

The variance of the exam scores is 146.

Watch out: Variance is expressed in squared units, which can make it difficult to interpret directly.

Standard Deviation

The standard deviation is the square root of the variance. It provides a measure of dispersion in the same units as the original data, making it easier to interpret.

Formula: Standard Deviation (σ) = √Variance

Example: Continuing from the previous example, we calculated the variance to be 146.

Standard Deviation = √146 ≈ 12.08

The standard deviation of the exam scores is approximately 12.08.

Watch out: A low standard deviation indicates that the data points are close to the mean, while a high standard deviation indicates that they are spread out over a larger range of values.

Interquartile Range (IQR)

The interquartile range (IQR) measures the spread of the middle 50% of a dataset. It is calculated by finding the difference between the first quartile (Q1) and the third quartile (Q3).

Formula: IQR = Q3 - Q1

Example: For the dataset: 55, 70, 65, 80, 90, first, we need to order the data (already ordered). Then, we find Q1 and Q3.

  1. Determine Q1 (the median of the first half of the data):
Q1 = 65
  1. Determine Q3 (the median of the second half of the data):
Q3 = 80
  1. Calculate IQR:
IQR = 80 - 65 = 15

The interquartile range of the exam scores is 15.

Watch out: The IQR is less affected by outliers than the range, making it a more robust measure of dispersion.

Summary of Measures of Dispersion

  • Range: Difference between the maximum and minimum values.
  • Variance: Average of the squared deviations from the mean.
  • Standard Deviation: Square root of the variance, in the same units as the data.
  • Interquartile Range: Difference between the third quartile and the first quartile.

Check your understanding

  • What is the range of the dataset: 45, 60, 75, 90, 100?
  • How do you calculate variance for a dataset?
  • What does a high standard deviation indicate about a dataset?
  • Explain the significance of the interquartile range in understanding data dispersion.