Confidence Intervals

STA1505 - Statistics for Beginners · Inferential Statistics

Confidence Intervals

Confidence intervals are a key concept in inferential statistics. They provide a range of values that is likely to contain the population parameter, such as the mean or proportion, based on a sample statistic. Understanding confidence intervals allows you to make inferences about a population from a sample.

What is a Confidence Interval?

A confidence interval is an estimate of a range of values for a population parameter. It is constructed from sample data and is associated with a confidence level, usually expressed as a percentage (e.g., 95% confidence level). This percentage indicates how confident you are that the interval contains the true population parameter.

Remember: A 95% confidence interval means that if you were to take 100 different samples and compute a confidence interval for each sample, approximately 95 of the intervals would contain the true population parameter.

Components of a Confidence Interval

A confidence interval consists of three main components:

  • The sample statistic (e.g., sample mean)
  • The margin of error
  • The confidence level

Calculating a Confidence Interval for the Mean

To calculate a confidence interval for the mean, follow these steps:

  1. Determine the sample mean (B5) and the sample standard deviation (s).
  2. Decide on the confidence level (e.g., 95%) and find the corresponding critical value (z or t) from statistical tables.
  3. Calculate the margin of error (E) using the formula:

E = critical value × (s / √n)

where n is the sample size.

  1. Construct the confidence interval using the formula:

Confidence Interval = (B5 - E, B5 + E)

Example: Confidence Interval for the Mean

Suppose you have a sample of 30 students from a class, and their test scores are as follows:

78, 82, 85, 90, 76, 88, 94, 79, 84, 91, 75, 87, 83, 80, 92, 86, 89, 77, 81, 93, 95, 74, 97, 96, 70, 72, 73, 71, 69, 68

First, calculate the sample mean (B5) and sample standard deviation (s):

Mean (B5) = (78 + 82 + 85 + 90 + 76 + 88 + 94 + 79 + 84 + 91 + 75 + 87 + 83 + 80 + 92 + 86 + 89 + 77 + 81 + 93 + 95 + 74 + 97 + 96 + 70 + 72 + 73 + 71 + 69 + 68) / 30 = 81.5
Standard Deviation (s) = √[(Σ(x - B5)²) / (n - 1)] = 7.45

Next, for a 95% confidence level, use the critical value of 1.96 (for large samples, z-distribution can be used). Now, calculate the margin of error (E):

E = 1.96 × (7.45 / √30) = 1.96 × 1.36 = 2.67

Finally, construct the confidence interval:

Confidence Interval = (81.5 - 2.67, 81.5 + 2.67) = (78.83, 84.17)

This means you can be 95% confident that the true mean test score of all students in the class is between 78.83 and 84.17.

Calculating a Confidence Interval for Proportions

To calculate a confidence interval for a population proportion, follow these steps:

  1. Determine the sample proportion (p) and the sample size (n).
  2. Decide on the confidence level (e.g., 95%) and find the corresponding critical value (z).
  3. Calculate the margin of error (E) using the formula:

E = critical value × √[p(1 - p) / n]

  1. Construct the confidence interval using the formula:

Confidence Interval = (p - E, p + E)

Example: Confidence Interval for Proportions

Suppose you surveyed 200 people about their preference for a new product, and 120 said they liked it. The sample proportion (p) is:

p = 120 / 200 = 0.6

For a 95% confidence level, the critical value is 1.96. Now, calculate the margin of error (E):

E = 1.96 × √[0.6(1 - 0.6) / 200] = 1.96 × √[0.6 × 0.4 / 200] = 1.96 × √[0.0012] = 1.96 × 0.03464 = 0.0679

Finally, construct the confidence interval:

Confidence Interval = (0.6 - 0.0679, 0.6 + 0.0679) = (0.5321, 0.6679)

This means you can be 95% confident that the true proportion of people who like the product is between 53.21% and 66.79%.

Watch out: Ensure that the sample size is large enough for the normal approximation to be valid. A common rule is that both np and n(1 - p) should be greater than 5.

Interpreting Confidence Intervals

Interpreting a confidence interval correctly is crucial. A confidence interval does not guarantee that the true population parameter lies within the interval for any specific sample. Instead, it reflects the uncertainty of the estimate based on the sample data.

Example of Interpretation

If a 95% confidence interval for the mean test score is (78.83, 84.17), you can say that you are 95% confident that the true mean test score for all students is between 78.83 and 84.17. However, if you take a different sample, the confidence interval might change.

Factors Affecting Confidence Intervals

Several factors can affect the width of a confidence interval:

  • Sample Size (n): Larger sample sizes result in smaller margins of error, leading to narrower confidence intervals.
  • Confidence Level: Higher confidence levels lead to wider confidence intervals.
  • Variability in Data: Greater variability in the data results in wider confidence intervals.

Example of Sample Size Effect

Suppose you have a sample mean of 81.5 with a standard deviation of 7.45. If you increase the sample size from 30 to 100, the new margin of error will be:

E = 1.96 × (7.45 / √100) = 1.96 × 0.745 = 1.46

The new confidence interval will be:

Confidence Interval = (81.5 - 1.46, 81.5 + 1.46) = (80.04, 82.96)

This interval is narrower than the previous one, showing that increasing the sample size reduces uncertainty.

Tip: Always consider the trade-off between sample size and resources. Increasing sample size can improve estimates but may require more time and money.

Common Misconceptions

Many students misunderstand confidence intervals. Here are some common misconceptions:

  • Believing that a confidence interval guarantees that the population parameter is within the interval.
  • Thinking that a wider interval is always better; wider intervals indicate more uncertainty.
  • Assuming that the confidence level applies to individual samples rather than to the method of constructing intervals.

Summary

  • A confidence interval estimates a range of values for a population parameter.
  • It consists of a sample statistic, margin of error, and confidence level.
  • The formula for a confidence interval for the mean is: Confidence Interval = (B5 - E, B5 + E).
  • The formula for a confidence interval for proportions is: Confidence Interval = (p - E, p + E).
  • Factors affecting confidence intervals include sample size, confidence level, and data variability.

Check your understanding

  • What is a confidence interval?
  • How do you calculate the margin of error for a sample mean?
  • What factors affect the width of a confidence interval?
  • Why is it important to interpret confidence intervals correctly?