Inferential Statistics

STA1507 - Introduction to Research Skills · Data Analysis

Inferential Statistics

Inferential statistics is a branch of statistics that allows you to make conclusions about a population based on a sample. A population is the entire group you want to draw conclusions about, while a sample is a subset of that population. The goal of inferential statistics is to use sample data to estimate population parameters and test hypotheses.

Population and Sample

A population includes all members of a specified group, while a sample is a smaller group selected from that population. For example, if you want to study the average height of adult men in South Africa, your population would be all adult men in the country, and your sample could be 100 men selected from different regions.

Remember: Always ensure your sample is representative of the population to avoid bias.

Sampling Methods

There are several methods for selecting a sample:

  • Simple Random Sampling: Every member of the population has an equal chance of being selected. For example, randomly selecting names from a hat.
  • Stratified Sampling: The population is divided into subgroups (strata) and samples are taken from each stratum. For example, if you divide the population by age groups and then randomly select individuals from each age group.
  • Systematic Sampling: Members are selected at regular intervals. For example, selecting every 10th person from a list.
  • Cluster Sampling: The population is divided into clusters, and entire clusters are randomly selected. For example, selecting entire schools from a district to study student performance.

Estimation

Estimation involves using sample data to estimate population parameters. There are two types of estimates:

  1. Point Estimate: A single value that serves as an estimate of a population parameter. For example, if the average height of your sample of 100 men is 1.75 m, this is a point estimate of the average height of all adult men in South Africa.
  2. Interval Estimate: A range of values that is likely to contain the population parameter. This is often expressed as a confidence interval.

Confidence Intervals

A confidence interval gives an estimated range of values which is likely to include an unknown population parameter, with a certain level of confidence. The most common confidence levels are 90%, 95%, and 99%.

The formula for a confidence interval for a population mean is:

CI = x̄ ± Z * (σ/√n)

Where:

  • x̄ = sample mean
  • Z = Z-value (from Z-table based on confidence level)
  • σ = population standard deviation
  • n = sample size

Example of a Confidence Interval

Suppose you have a sample mean height of 1.75 m, a population standard deviation of 0.1 m, and a sample size of 100. For a 95% confidence level, the Z-value is 1.96.

Using the formula:

CI = 1.75 ± 1.96 * (0.1/√100)

Calculate the margin of error:

Margin of Error = 1.96 * (0.1/10) = 0.0196

Now, calculate the confidence interval:

CI = 1.75 ± 0.0196

This results in:

CI = (1.7304, 1.7696)

You can say with 95% confidence that the average height of all adult men in South Africa is between 1.7304 m and 1.7696 m.

Watch out: Ensure you use the correct Z-value for your chosen confidence level. Using the wrong Z-value will lead to incorrect confidence intervals.

Hypothesis Testing

Hypothesis testing is a method used to decide whether there is enough evidence to reject a null hypothesis (H0). The null hypothesis usually states that there is no effect or no difference. The alternative hypothesis (H1) states the opposite.

The steps in hypothesis testing are:

  1. State the null and alternative hypotheses.
  2. Select a significance level (α), commonly set at 0.05.
  3. Calculate the test statistic using sample data.
  4. Determine the critical value from statistical tables.
  5. Make a decision: reject or fail to reject the null hypothesis based on the test statistic and critical value.

Example of Hypothesis Testing

  • H0: μ = 1.75 m (the average height is 1.75 m)
  • H1: μ ≠ 1.75 m (the average height is not 1.75 m)

Assume you collect a sample of 100 men, and the sample mean height is 1.78 m with a standard deviation of 0.1 m. To test the hypothesis, you can use a Z-test:

The Z-test statistic is calculated as follows:

Z = (x̄ - μ) / (σ/√n)

Substituting the values:

Z = (1.78 - 1.75) / (0.1/√100) = 3

Now, compare the Z-value to the critical Z-value at α = 0.05. The critical values are -1.96 and 1.96.

Since 3 is greater than 1.96, you reject the null hypothesis. There is enough evidence to conclude that the average height of adult men in South Africa is different from 1.75 m.

Tip: Always check the assumptions of the test you are using. For a Z-test, the sample should be normally distributed, or the sample size should be large (n > 30).

P-Values

The p-value is the probability of obtaining a test statistic as extreme as, or more extreme than, the one observed, assuming the null hypothesis is true. A smaller p-value indicates stronger evidence against the null hypothesis.

In hypothesis testing, if the p-value is less than the significance level (α), you reject the null hypothesis. If it is greater, you fail to reject the null hypothesis.

Example of P-Value Calculation

Since 0.001 < 0.05, you reject the null hypothesis. This confirms your earlier conclusion that the average height of adult men in South Africa is different from 1.75 m.

Conclusion

Inferential statistics provides powerful tools for making predictions and decisions based on sample data. By understanding how to estimate population parameters and conduct hypothesis tests, you can apply these techniques to real-world research problems.

Remember: Always ensure your sample is representative of the population and check the assumptions of the statistical tests you use.

Summary

  • Inferential statistics allows conclusions about populations based on samples.
  • Sampling methods include simple random, stratified, systematic, and cluster sampling.
  • Estimation can be point estimates or interval estimates (confidence intervals).
  • Hypothesis testing involves rejecting or failing to reject a null hypothesis based on sample data.
  • P-values provide a measure of evidence against the null hypothesis.

Check your understanding

  1. What is the difference between a population and a sample?
  2. Explain the concept of a confidence interval.
  3. What are the steps involved in hypothesis testing?
  4. How do you interpret a p-value in hypothesis testing?