Descriptive Statistics
STA1506 - Basic Statistical Computing · Exploring Data
Descriptive Statistics
Descriptive statistics summarise and describe the main features of a dataset. They provide simple summaries about the sample and the measures. Descriptive statistics can include measures of central tendency, measures of variability, and graphical representations of data.
Measures of Central Tendency
Measures of central tendency indicate where the centre of a dataset lies. The three most common measures are the mean, median, and mode.
Mean
The mean is the average of a set of numbers. To calculate the mean, sum all the values and divide by the number of values.
mean = (x1 + x2 + ... + xn) / nFor example, consider the following dataset of exam scores: 70, 85, 90, 75, 80.
- Sum the scores: 70 + 85 + 90 + 75 + 80 = 400
- Count the number of scores: 5
- Calculate the mean: 400 / 5 = 80
The mean score is 80.
Remember: The mean can be affected by extreme values (outliers).
Median
The median is the middle value when the data is arranged in ascending order. If there is an even number of values, the median is the average of the two middle numbers.
Using the same dataset: 70, 75, 80, 85, 90, first arrange the data in order:
- Arrange: 70, 75, 80, 85, 90
- Count the number of scores: 5 (odd number)
- Find the middle value: 80
The median score is 80.
For an even number of scores, consider the dataset: 70, 75, 80, 85.
- Arrange: 70, 75, 80, 85
- Count: 4 (even number)
- Find the average of the two middle values: (75 + 80) / 2 = 77.5
The median score is 77.5.
Watch out: Ensure you arrange the data in order before finding the median.
Mode
The mode is the value that appears most frequently in a dataset. A dataset may have one mode, more than one mode, or no mode at all.
For example, in the dataset 70, 85, 90, 75, 80, the mode is not applicable because all values occur only once.
In the dataset 70, 75, 70, 85, the mode is 70 because it appears most frequently.
Tip: If you have two values that appear with the same highest frequency, the dataset is bimodal.
Measures of Variability
Measures of variability indicate how spread out the values in a dataset are. The common measures of variability are the range, variance, and standard deviation.
Range
The range is the difference between the highest and lowest values in a dataset.
For example, in the dataset 70, 85, 90, 75, 80:
- Find the highest value: 90
- Find the lowest value: 70
- Calculate the range: 90 - 70 = 20
The range is 20.
Variance
Variance measures how far each number in the dataset is from the mean and, therefore, from every other number in the dataset. The formula for variance is:
variance = Σ(xi - mean)² / nWhere Σ is the summation symbol, xi represents each value, and n is the number of values.
Using the previous dataset: 70, 85, 90, 75, 80, we first calculate the mean, which is 80.
- Calculate each deviation from the mean: 70 - 80 = -10, 85 - 80 = 5, 90 - 80 = 10, 75 - 80 = -5, 80 - 80 = 0.
- Square each deviation: (-10)² = 100, 5² = 25, 10² = 100, (-5)² = 25, 0² = 0.
- Sum the squared deviations: 100 + 25 + 100 + 25 + 0 = 250.
- Divide by the number of values: 250 / 5 = 50.
The variance is 50.
Standard Deviation
The standard deviation is the square root of the variance. It provides a measure of the average distance of each data point from the mean.
standard deviation = √varianceContinuing from the previous example:
- Calculate the standard deviation: √50 ≈ 7.07.
The standard deviation is approximately 7.07.
Tip: A low standard deviation indicates that the data points tend to be close to the mean, while a high standard deviation indicates that the data points are spread out over a wider range.
Graphical Representation of Data
Graphical representations help to visualise data and understand its distribution. Common types of graphs include histograms, box plots, and bar charts.
Histograms
A histogram is a graphical representation of the distribution of numerical data. It displays the number of data points that fall within specified ranges (bins).
To create a histogram:
- Determine the range of the data.
- Divide the range into intervals (bins).
- Count how many data points fall into each bin.
- Draw bars for each bin, where the height of the bar represents the count of data points.
Box Plots
A box plot (or whisker plot) displays the distribution of a dataset based on a five-number summary: minimum, first quartile (Q1), median, third quartile (Q3), and maximum.
To create a box plot:
- Calculate the five-number summary.
- Draw a box from Q1 to Q3.
- Draw a line at the median inside the box.
- Extend lines (whiskers) from the box to the minimum and maximum values.
Summary of Descriptive Statistics
- Measures of central tendency include mean, median, and mode.
- Measures of variability include range, variance, and standard deviation.
- Graphical representations include histograms and box plots.
Check your understanding
- What is the mean of the following dataset: 60, 70, 80, 90, 100?
- How do you calculate the median for an even number of values?
- What does a high standard deviation indicate about a dataset?
- Describe how you would create a histogram for a given dataset.