Measures of Central Tendency
STA1505 - Statistics for Beginners · Descriptive Statistics
Measures of Central Tendency
Measures of central tendency are statistical values that represent the centre point or typical value of a dataset. The three main measures of central tendency are the mean, median, and mode. Each measure provides different insights into the dataset.
The Mean
The mean is the average of a set of numbers. To calculate the mean, you add all the numbers together and then divide by the count of the numbers.
Formula:Mean (μ) = (Σx) / n
Where:
- Σx = the sum of all values
- n = the number of values
Example of Calculating the Mean
Consider the following dataset representing the ages of a group of five people: 22, 25, 30, 28, and 35.
- Add the ages together: 22 + 25 + 30 + 28 + 35 = 140
- Count the number of ages: There are 5 ages.
- Divide the total by the count: 140 / 5 = 28
The mean age is 28 years.
Remember: The mean can be affected by extreme values (outliers).
The Median
The median is the middle value of a dataset when the numbers are arranged in ascending order. If there is an even number of values, the median is the average of the two middle numbers.
Steps to Find the Median:- Arrange the numbers in ascending order.
- Identify the middle number. If there are two middle numbers, calculate their average.
Example of Calculating the Median
Using the same dataset of ages: 22, 25, 30, 28, and 35, first arrange them in order: 22, 25, 28, 30, 35.
- There are 5 numbers, so the middle number is the third one: 28.
The median age is 28 years.
Now, consider a different dataset with an even number of values: 22, 25, 30, and 28.
- Arrange the numbers: 22, 25, 28, 30.
- There are 4 numbers, so the two middle numbers are 25 and 28.
- Calculate their average: (25 + 28) / 2 = 26.5.
The median for this dataset is 26.5 years.
Watch out: Always arrange the numbers in order before finding the median.
The Mode
The mode is the value that appears most frequently in a dataset. A dataset may have one mode, more than one mode (bimodal or multimodal), or no mode at all.
Steps to Find the Mode:- Count the frequency of each value in the dataset.
- Identify the value(s) with the highest frequency.
Example of Calculating the Mode
Consider the dataset: 22, 25, 22, 30, 28, 35, 22.
- Count the frequency: 22 appears 3 times, 25 appears 1 time, 30 appears 1 time, 28 appears 1 time, and 35 appears 1 time.
The mode is 22 because it appears most frequently.
Now, consider another dataset: 22, 25, 30, 28, 35.
- Count the frequency: Each number appears once.
This dataset has no mode because no number repeats.
Tip: Always check the frequency of each value to determine the mode accurately.
Comparing the Measures
- The mean considers all values and is useful for normally distributed data.
- The median is less affected by outliers and provides a better measure for skewed distributions.
- The mode shows the most common value and is useful for categorical data.
Choosing the Right Measure
- If your data is normally distributed, the mean is appropriate.
- If your data has outliers, the median is a better choice.
- If you are dealing with categorical data, use the mode.
Remember: Understanding the nature of your data helps in selecting the right measure of central tendency.
Summary
- The mean is the average of a dataset.
- The median is the middle value of a dataset.
- The mode is the most frequently occurring value in a dataset.
- Choose the appropriate measure based on the characteristics of your data.
Check your understanding
- Calculate the mean of the following dataset: 10, 15, 20, 25, 30.
- What is the median of the following dataset: 7, 3, 9, 5, 6?
- Identify the mode in the dataset: 4, 1, 2, 4, 3, 4, 5.
- When would you choose the median over the mean?