Data Distribution
STA1505 - Statistics for Beginners · Descriptive Statistics
Data Distribution
Data distribution refers to how values in a dataset are spread or arranged. Understanding data distribution is crucial in statistics because it provides insights into the characteristics of the data. This topic will cover types of data distributions, their shapes, and how to describe them.
Types of Data Distributions
There are several types of data distributions, but the most common ones are:
- Normal Distribution: This is a symmetric, bell-shaped distribution where most of the observations cluster around the central peak, and probabilities for values further away from the mean taper off equally in both directions.
- Skewed Distribution: This distribution is not symmetric and can be either positively skewed (tail on the right) or negatively skewed (tail on the left).
- Bimodal Distribution: This distribution has two different modes or peaks, indicating that there are two prevalent values in the dataset.
- Uniform Distribution: In this distribution, all outcomes are equally likely. The graph of a uniform distribution is a rectangle.
Normal Distribution
The normal distribution is significant in statistics because many statistical methods assume that data follows a normal distribution. The characteristics of a normal distribution include:
- The mean, median, and mode are all equal.
- It is symmetric around the mean.
- Approximately 68% of the data falls within one standard deviation of the mean, about 95% within two standard deviations, and about 99.7% within three standard deviations.
Remember: The empirical rule applies to normal distributions. Use it to understand how data is spread around the mean.
Example of Normal Distribution
Consider a dataset of students' test scores in a statistics exam:
Scores: 55, 60, 65, 70, 70, 75, 80, 85, 90, 95To determine if this dataset follows a normal distribution, we can calculate the mean and standard deviation.
Step 1: Calculate the Mean
The mean is calculated by summing all the scores and dividing by the number of scores.
Mean = (55 + 60 + 65 + 70 + 70 + 75 + 80 + 85 + 90 + 95) / 10Calculating this gives:
Mean = 75Step 2: Calculate the Standard Deviation
The standard deviation measures the dispersion of the dataset. First, we find the variance:
Variance = [(55-75)^2 + (60-75)^2 + (65-75)^2 + (70-75)^2 + (70-75)^2 + (75-75)^2 + (80-75)^2 + (85-75)^2 + (90-75)^2 + (95-75)^2] / 10Calculating each squared difference:
(-20)^2 = 400, (-15)^2 = 225, (-10)^2 = 100, (-5)^2 = 25, (-5)^2 = 25, (0)^2 = 0, (5)^2 = 25, (10)^2 = 100, (15)^2 = 225, (20)^2 = 400Now, summing those values:
400 + 225 + 100 + 25 + 25 + 0 + 25 + 100 + 225 + 400 = 1025Now, divide by the number of scores:
Variance = 1025 / 10 = 102.5The standard deviation is the square root of the variance:
Standard Deviation = √102.5 ≈ 10.12Step 3: Assess the Distribution
Using the mean (75) and the standard deviation (10.12), we can see that:
- 68% of scores will fall between 64.88 (75 - 10.12) and 85.12 (75 + 10.12).
- 95% will fall between 54.76 (75 - 2*10.12) and 95.24 (75 + 2*10.12).
- 99.7% will fall between 44.64 (75 - 3*10.12) and 105.36 (75 + 3*10.12).
This confirms a normal distribution shape.
Skewed Distribution
In a skewed distribution, data points are not symmetrically distributed around the mean. Skewness indicates the direction of the tail:
- Positively Skewed: The tail on the right side is longer or fatter. The mean is greater than the median.
- Negatively Skewed: The tail on the left side is longer or fatter. The mean is less than the median.
Example of Skewed Distribution
Consider the following dataset of household incomes in thousands of rand:
Incomes: 20, 25, 30, 35, 40, 45, 100, 120, 150, 200To check for skewness, calculate the mean and median.
Step 1: Calculate the Mean
Mean = (20 + 25 + 30 + 35 + 40 + 45 + 100 + 120 + 150 + 200) / 10Calculating gives:
Mean = 73.5Step 2: Calculate the Median
First, order the incomes:
Ordered Incomes: 20, 25, 30, 35, 40, 45, 100, 120, 150, 200Since there are 10 values, the median is the average of the 5th and 6th values:
Median = (40 + 45) / 2 = 42.5Step 3: Assess Skewness
Here, the mean (73.5) is greater than the median (42.5). This indicates a positive skew.
Watch out: Do not confuse the mean and median when assessing skewness. The mean is affected by extreme values, while the median is not.
Bimodal Distribution
A bimodal distribution has two distinct peaks. This can occur when data comes from two different groups. For example, the heights of adult males and females in a population could create a bimodal distribution.
Example of Bimodal Distribution
Consider the following dataset of heights in cm:
Heights: 150, 152, 154, 155, 160, 162, 170, 172, 180, 182, 185, 200In this case, we can see two groups of heights. To identify the modes, we can group the data:
Group 1 (150-160 cm): 5 valuesGroup 2 (170-200 cm): 7 valuesBoth groups show peaks, indicating a bimodal distribution.
Uniform Distribution
In a uniform distribution, all outcomes are equally likely. This means that every value within a certain range has the same probability of occurring.
Example of Uniform Distribution
Consider the roll of a fair six-sided die. Each side has an equal chance of landing face up:
Outcomes: 1, 2, 3, 4, 5, 6Each outcome has a probability of 1/6.
Remember: In a uniform distribution, the probability of each outcome is the same.
Graphical Representation of Data Distributions
Visualising data distributions is essential in statistics. Common graphical representations include:
- Histograms: These show the frequency of data points within specified ranges (bins). Each bin represents a range of values.
- Box Plots: These provide a visual summary of the minimum, first quartile, median, third quartile, and maximum of a dataset.
- Density Plots: These smooth out the histogram to show the probability density of the variable.
Example of a Histogram
To create a histogram for the following dataset of exam scores:
Scores: 55, 60, 65, 70, 70, 75, 80, 85, 90, 95You can group the scores into bins:
Bin 1: 50-60, Bin 2: 61-70, Bin 3: 71-80, Bin 4: 81-90, Bin 5: 91-100Count how many scores fall into each bin:
Bin 1: 1, Bin 2: 4, Bin 3: 3, Bin 4: 2, Bin 5: 0Now, you can plot these frequencies on the histogram.
Summary
- Data distribution shows how values are spread in a dataset.
- Types of distributions include normal, skewed, bimodal, and uniform.
- Normal distribution is symmetric, while skewed distribution has tails on one side.
- Graphical representations like histograms and box plots help visualize data distributions.
Check your understanding
- What are the characteristics of a normal distribution?
- How can you determine if a dataset is positively skewed?
- What is a bimodal distribution, and how can it occur?
- What is the difference between a histogram and a box plot?