Data Distribution

STA1505 - Statistics for Beginners · Descriptive Statistics

Data Distribution

Data distribution refers to how values in a dataset are spread or arranged. Understanding data distribution is crucial in statistics because it provides insights into the characteristics of the data. This topic will cover types of data distributions, their shapes, and how to describe them.

Types of Data Distributions

There are several types of data distributions, but the most common ones are:

  • Normal Distribution: This is a symmetric, bell-shaped distribution where most of the observations cluster around the central peak, and probabilities for values further away from the mean taper off equally in both directions.
  • Skewed Distribution: This distribution is not symmetric and can be either positively skewed (tail on the right) or negatively skewed (tail on the left).
  • Bimodal Distribution: This distribution has two different modes or peaks, indicating that there are two prevalent values in the dataset.
  • Uniform Distribution: In this distribution, all outcomes are equally likely. The graph of a uniform distribution is a rectangle.

Normal Distribution

The normal distribution is significant in statistics because many statistical methods assume that data follows a normal distribution. The characteristics of a normal distribution include:

  • The mean, median, and mode are all equal.
  • It is symmetric around the mean.
  • Approximately 68% of the data falls within one standard deviation of the mean, about 95% within two standard deviations, and about 99.7% within three standard deviations.

Remember: The empirical rule applies to normal distributions. Use it to understand how data is spread around the mean.

Example of Normal Distribution

Consider a dataset of students' test scores in a statistics exam:

Scores: 55, 60, 65, 70, 70, 75, 80, 85, 90, 95

To determine if this dataset follows a normal distribution, we can calculate the mean and standard deviation.

Step 1: Calculate the Mean

The mean is calculated by summing all the scores and dividing by the number of scores.

Mean = (55 + 60 + 65 + 70 + 70 + 75 + 80 + 85 + 90 + 95) / 10

Calculating this gives:

Mean = 75

Step 2: Calculate the Standard Deviation

The standard deviation measures the dispersion of the dataset. First, we find the variance:

Variance = [(55-75)^2 + (60-75)^2 + (65-75)^2 + (70-75)^2 + (70-75)^2 + (75-75)^2 + (80-75)^2 + (85-75)^2 + (90-75)^2 + (95-75)^2] / 10

Calculating each squared difference:

(-20)^2 = 400, (-15)^2 = 225, (-10)^2 = 100, (-5)^2 = 25, (-5)^2 = 25, (0)^2 = 0, (5)^2 = 25, (10)^2 = 100, (15)^2 = 225, (20)^2 = 400

Now, summing those values:

400 + 225 + 100 + 25 + 25 + 0 + 25 + 100 + 225 + 400 = 1025

Now, divide by the number of scores:

Variance = 1025 / 10 = 102.5

The standard deviation is the square root of the variance:

Standard Deviation = √102.5 ≈ 10.12

Step 3: Assess the Distribution

Using the mean (75) and the standard deviation (10.12), we can see that:

  • 68% of scores will fall between 64.88 (75 - 10.12) and 85.12 (75 + 10.12).
  • 95% will fall between 54.76 (75 - 2*10.12) and 95.24 (75 + 2*10.12).
  • 99.7% will fall between 44.64 (75 - 3*10.12) and 105.36 (75 + 3*10.12).

This confirms a normal distribution shape.

Skewed Distribution

In a skewed distribution, data points are not symmetrically distributed around the mean. Skewness indicates the direction of the tail:

  • Positively Skewed: The tail on the right side is longer or fatter. The mean is greater than the median.
  • Negatively Skewed: The tail on the left side is longer or fatter. The mean is less than the median.

Example of Skewed Distribution

Consider the following dataset of household incomes in thousands of rand:

Incomes: 20, 25, 30, 35, 40, 45, 100, 120, 150, 200

To check for skewness, calculate the mean and median.

Step 1: Calculate the Mean

Mean = (20 + 25 + 30 + 35 + 40 + 45 + 100 + 120 + 150 + 200) / 10

Calculating gives:

Mean = 73.5

Step 2: Calculate the Median

First, order the incomes:

Ordered Incomes: 20, 25, 30, 35, 40, 45, 100, 120, 150, 200

Since there are 10 values, the median is the average of the 5th and 6th values:

Median = (40 + 45) / 2 = 42.5

Step 3: Assess Skewness

Here, the mean (73.5) is greater than the median (42.5). This indicates a positive skew.

Watch out: Do not confuse the mean and median when assessing skewness. The mean is affected by extreme values, while the median is not.

Bimodal Distribution

A bimodal distribution has two distinct peaks. This can occur when data comes from two different groups. For example, the heights of adult males and females in a population could create a bimodal distribution.

Example of Bimodal Distribution

Consider the following dataset of heights in cm:

Heights: 150, 152, 154, 155, 160, 162, 170, 172, 180, 182, 185, 200

In this case, we can see two groups of heights. To identify the modes, we can group the data:

Group 1 (150-160 cm): 5 values
Group 2 (170-200 cm): 7 values

Both groups show peaks, indicating a bimodal distribution.

Uniform Distribution

In a uniform distribution, all outcomes are equally likely. This means that every value within a certain range has the same probability of occurring.

Example of Uniform Distribution

Consider the roll of a fair six-sided die. Each side has an equal chance of landing face up:

Outcomes: 1, 2, 3, 4, 5, 6

Each outcome has a probability of 1/6.

Remember: In a uniform distribution, the probability of each outcome is the same.

Graphical Representation of Data Distributions

Visualising data distributions is essential in statistics. Common graphical representations include:

  • Histograms: These show the frequency of data points within specified ranges (bins). Each bin represents a range of values.
  • Box Plots: These provide a visual summary of the minimum, first quartile, median, third quartile, and maximum of a dataset.
  • Density Plots: These smooth out the histogram to show the probability density of the variable.

Example of a Histogram

To create a histogram for the following dataset of exam scores:

Scores: 55, 60, 65, 70, 70, 75, 80, 85, 90, 95

You can group the scores into bins:

Bin 1: 50-60, Bin 2: 61-70, Bin 3: 71-80, Bin 4: 81-90, Bin 5: 91-100

Count how many scores fall into each bin:

Bin 1: 1, Bin 2: 4, Bin 3: 3, Bin 4: 2, Bin 5: 0

Now, you can plot these frequencies on the histogram.

Summary

  • Data distribution shows how values are spread in a dataset.
  • Types of distributions include normal, skewed, bimodal, and uniform.
  • Normal distribution is symmetric, while skewed distribution has tails on one side.
  • Graphical representations like histograms and box plots help visualize data distributions.

Check your understanding

  1. What are the characteristics of a normal distribution?
  2. How can you determine if a dataset is positively skewed?
  3. What is a bimodal distribution, and how can it occur?
  4. What is the difference between a histogram and a box plot?
    Data Distribution – STA1505 - Statistics for Beginners notes | Tyro Study