Types of Statistics
STA1505 - Statistics for Beginners · Introduction to Statistics
Types of Statistics
Statistics can be broadly classified into two main categories: descriptive statistics and inferential statistics. Each type serves a different purpose in the field of statistics.
Descriptive Statistics
Descriptive statistics involves summarising and presenting data in a meaningful way. This type of statistics helps to describe the basic features of the data in a study. It provides simple summaries about the sample and the measures. Descriptive statistics can include measures of central tendency and measures of variability.
Measures of Central Tendency
Measures of central tendency describe the centre of a data set. The three most common measures are the mean, median, and mode.
- Mean: The mean is the average of a data set. To calculate the mean, you add all the values together and divide by the number of values.
- Median: The median is the middle value when the data set is ordered from smallest to largest. If there is an even number of observations, the median is the average of the two middle numbers.
- Mode: The mode is the value that appears most frequently in the data set.
Remember: The mean can be affected by extreme values (outliers), while the median is a better measure of central tendency when outliers are present.
Example of Calculating Mean, Median, and Mode
Consider the following data set representing the ages of a group of five people: 22, 25, 25, 30, 35.
Step 1: Calculate the Mean
Mean = (22 + 25 + 25 + 30 + 35) / 5Mean = 107 / 5 = 21.4
Step 2: Calculate the Median
First, order the data: 22, 25, 25, 30, 35. The median is the middle value, which is 25.
Step 3: Calculate the Mode
The mode is the value that appears most frequently. In this case, the mode is 25.
Measures of Variability
Measures of variability describe how spread out the values in a data set are. Common measures include the range, variance, and standard deviation.
- Range: The range is the difference between the highest and lowest values in a data set.
- Variance: Variance measures the average squared deviation from the mean. It indicates how much the values in a data set differ from the mean.
- Standard Deviation: The standard deviation is the square root of the variance. It provides a measure of the average distance of each data point from the mean.
Remember: A low standard deviation indicates that the data points tend to be close to the mean, while a high standard deviation indicates that the data points are spread out over a wider range of values.
Example of Calculating Range, Variance, and Standard Deviation
Using the same data set (22, 25, 25, 30, 35), let us calculate the range, variance, and standard deviation.
Step 1: Calculate the Range
Range = Highest value - Lowest valueRange = 35 - 22 = 13
Step 2: Calculate the Variance
First, find the mean (which we previously calculated as 25). Then, calculate each deviation from the mean:
Deviation from mean: (22 - 25), (25 - 25), (25 - 25), (30 - 25), (35 - 25)Which gives us: -3, 0, 0, 5, 10. Now, square each deviation:
Squared deviations: 9, 0, 0, 25, 100Now, calculate the variance:
Variance = (9 + 0 + 0 + 25 + 100) / 5Variance = 134 / 5 = 26.8
Step 3: Calculate the Standard Deviation
Standard Deviation = √VarianceStandard Deviation = √26.8 ≈ 5.18
Inferential Statistics
Inferential statistics involves making predictions or inferences about a population based on a sample of data. This type of statistics helps to draw conclusions and make decisions based on data analysis.
Sampling
A sample is a subset of a population. Using a sample allows researchers to make inferences about the entire population without needing to collect data from every individual. There are different methods of sampling, including:
- Random Sampling: Every member of the population has an equal chance of being selected.
- Stratified Sampling: The population is divided into subgroups (strata) and samples are taken from each subgroup.
- Systematic Sampling: Members are selected at regular intervals from a sorted list.
Tip: Random sampling is often the best method to ensure that the sample is representative of the population.
Hypothesis Testing
Hypothesis testing is a method used to determine if there is enough statistical evidence in a sample to infer that a certain condition is true for the entire population. A hypothesis is a statement that can be tested. There are two types of hypotheses:
- Null Hypothesis (H0): This is a statement of no effect or no difference. It is the hypothesis that the researcher tries to disprove.
- Alternative Hypothesis (H1): This represents the statement that there is an effect or a difference.
Watch out: Students often confuse the null and alternative hypotheses. Remember that the null hypothesis is what you aim to test against.
Example of Hypothesis Testing
Suppose a researcher wants to test if a new teaching method is more effective than the traditional method. The null hypothesis (H0) could be that there is no difference in effectiveness between the two methods, while the alternative hypothesis (H1) could be that the new method is more effective.
Conclusion
Understanding the types of statistics is essential for data analysis. Descriptive statistics allows you to summarise and describe data, while inferential statistics enables you to make predictions and test hypotheses. Both types are crucial for making informed decisions based on data.
Summary
- Descriptive statistics summarise and present data.
- Measures of central tendency include mean, median, and mode.
- Measures of variability include range, variance, and standard deviation.
- Inferential statistics involve making predictions about a population based on sample data.
- Hypothesis testing is used to determine if there is enough evidence to support a hypothesis.
Check your understanding
- What is the difference between descriptive and inferential statistics?
- How do you calculate the median of a data set?
- What are the two types of hypotheses in hypothesis testing?
- Why is random sampling important in research?