Interpreting Statistical Outputs

STA1506 - Basic Statistical Computing · Reporting Results

Interpreting Statistical Outputs

Statistical outputs are the results generated by statistical software after performing analyses on data. Understanding these outputs is crucial for making informed decisions based on data. In this topic, you will learn how to interpret common statistical outputs, including descriptive statistics, inferential statistics, and regression analysis.

Descriptive Statistics

Descriptive statistics provide a summary of the data. They help you understand the basic features of your dataset. Common descriptive statistics include measures of central tendency and measures of variability.

Measures of Central Tendency

Measures of central tendency indicate where the centre of a dataset lies. The most common measures are:

  • Mean: The average of all data points.
  • Median: The middle value when data points are arranged in order.
  • Mode: The value that appears most frequently.

For example, consider the following dataset of test scores: 56, 78, 78, 85, 90, 92.

Calculating the Mean

To calculate the mean:

Mean = (56 + 78 + 78 + 85 + 90 + 92) / 6

Calculating the sum:

Mean = 479 / 6 = 79.83

The mean score is approximately 79.83.

Calculating the Median

To find the median, arrange the scores in order:

56, 78, 78, 85, 90, 92

Since there are six scores (an even number), the median is the average of the two middle values (78 and 85):

Median = (78 + 85) / 2 = 81.5

The median score is 81.5.

Calculating the Mode

In this dataset, the mode is 78 as it appears most frequently.

Remember: The mean can be influenced by extreme values (outliers), while the median is more robust in such cases.

Measures of Variability

Measures of variability indicate how spread out the data points are. Common measures include:

  • Range: The difference between the highest and lowest values.
  • Variance: The average of the squared differences from the mean.
  • Standard Deviation: The square root of the variance.

Using the previous dataset, let's calculate the range, variance, and standard deviation.

Calculating the Range

The range is calculated as follows:

Range = Highest score - Lowest score = 92 - 56 = 36

The range is 36.

Calculating the Variance

First, find the differences from the mean:

Differences: (56 - 79.83), (78 - 79.83), (78 - 79.83), (85 - 79.83), (90 - 79.83), (92 - 79.83)

This gives us:

-23.83, -1.83, -1.83, 5.17, 10.17, 12.17

Next, square these differences:

Squared Differences: 567.43, 3.35, 3.35, 26.73, 103.43, 148.43

Now, calculate the variance:

Variance = (567.43 + 3.35 + 3.35 + 26.73 + 103.43 + 148.43) / 6 = 141.60

The variance is approximately 141.60.

Calculating the Standard Deviation

The standard deviation is the square root of the variance:

Standard Deviation = √141.60 ≈ 11.87

The standard deviation is approximately 11.87.

Tip: Use statistical software to calculate descriptive statistics quickly, especially for large datasets.

Inferential Statistics

Inferential statistics allow you to make conclusions about a population based on a sample. Common inferential statistics include hypothesis testing and confidence intervals.

Hypothesis Testing

Hypothesis testing involves making an assumption (the hypothesis) and testing whether the data supports it. There are two types of hypotheses:

  • Null Hypothesis (H0): The hypothesis that there is no effect or difference.
  • Alternative Hypothesis (H1): The hypothesis that there is an effect or difference.

For example, suppose you want to test whether a new teaching method improves student performance. Your null hypothesis might be that the new method has no effect on scores.

Conducting a Hypothesis Test

1. Define your null and alternative hypotheses.

2. Choose a significance level (α), usually set at 0.05.

3. Calculate the test statistic using your data.

4. Compare the test statistic to a critical value or use a p-value to determine whether to reject H0.

Suppose you have a sample mean score of 82, a population mean score of 78, a standard deviation of 10, and a sample size of 30. You can calculate the test statistic using the formula:

Test Statistic (z) = (Sample Mean - Population Mean) / (Standard Deviation / √Sample Size)

z = (82 - 78) / (10 / √30) = 4 / (10 / 5.477) = 4 / 1.825 = 2.19

Now, compare this z-value with a critical z-value from the z-table for α = 0.05. The critical z-value is approximately 1.96. Since 2.19 > 1.96, you reject the null hypothesis.

Watch out: Ensure that you understand the difference between Type I and Type II errors. A Type I error occurs when you reject a true null hypothesis, while a Type II error occurs when you fail to reject a false null hypothesis.

Confidence Intervals

A confidence interval provides a range of values that is likely to contain the population parameter. For example, a 95% confidence interval means you can be 95% confident that the true population parameter lies within this range.

Calculating a Confidence Interval

The formula for a confidence interval for the mean is:

Confidence Interval = Sample Mean ± (Critical Value × Standard Error)

The standard error (SE) is calculated as:

SE = Standard Deviation / √Sample Size

Continuing with the previous example, calculate the standard error:

SE = 10 / √30 ≈ 1.83

Now, find the critical value for a 95% confidence level, which is approximately 1.96. The confidence interval is:

Confidence Interval = 82 ± (1.96 × 1.83)

This gives us:

82 ± 3.59 = (78.41, 85.59)

The 95% confidence interval for the mean score is (78.41, 85.59).

Regression Analysis

Regression analysis examines the relationship between variables. It helps you understand how the dependent variable changes when one or more independent variables change.

Simple Linear Regression

In simple linear regression, you have one dependent variable and one independent variable. The goal is to find the best-fitting line through the data points.

Equation of the Regression Line

The equation of a simple linear regression line is:

y = b0 + b1x

Where:

  • y = dependent variable
  • b0 = y-intercept
  • b1 = slope of the line
  • x = independent variable

For example, suppose you want to predict students' test scores based on hours studied. You conduct a regression analysis and find:

y = 50 + 10x

This means that for every additional hour studied, the test score increases by 10 points.

Interpreting the Coefficients

The coefficient b0 (50) indicates the expected test score for a student who did not study at all. The coefficient b1 (10) indicates the increase in score for each additional hour studied.

Remember: Always check the p-values associated with the coefficients to determine if they are statistically significant.

Summary

  • Descriptive statistics summarise data with measures of central tendency and variability.
  • Inferential statistics allow you to make conclusions about a population based on sample data.
  • Regression analysis helps understand relationships between variables.

Check your understanding

  1. What is the difference between the mean and the median?
  2. How do you calculate the standard deviation from the variance?
  3. What is the purpose of hypothesis testing?
  4. Explain the significance of the confidence interval.