Statistics is the branch of mathematics that deals with the collection, organisation, presentation, analysis and interpretation of numerical data. In our daily lives we constantly deal with data: the marks of students in a class, the temperatures of a city across a week, the runs scored by cricketers, the heights of plants in a garden. Statistics helps us make sense of such numbers, find patterns and take decisions.
In earlier classes we learnt how to collect and represent data using bar graphs, pictographs and tally marks. In this chapter we will go further. We will learn how to present data in a systematic way using frequency distributions, both for individual observations and for grouped data with class intervals. We will learn to draw bar graphs, histograms and frequency polygons, and finally we will study the three measures of central tendency: the mean, the median and the mode.
Raw data are the original, unorganised numbers collected from a survey or an observation. To make the data meaningful, we must organise it. The two main types of data are:
When we arrange data by counting how many times each value occurs, we get a frequency distribution. The number of times a particular value appears is called its frequency. For example, if the marks 45 appears 5 times, its frequency is 5.
Ungrouped (individual) frequency distribution: each distinct value is listed with its frequency. Grouped frequency distribution: when the data has many distinct values, we group them into class intervals like 0-10, 10-20, 20-30, etc., and count the frequency of each class.
For grouped data, the following terms are important:
When data are continuous, we use exclusive classes like 10-20, 20-30, where the upper limit of one class equals the lower limit of the next. When data are discrete (like marks out of 50 which are integers), we may use inclusive classes like 10-19, 20-29, and convert them to exclusive form by adjusting the limits before drawing histograms.
A bar graph represents the data with rectangular bars of equal width. The height of each bar is proportional to the frequency of the corresponding category. The bars are equally spaced. Bar graphs are used for discrete data like the number of students in each house.
A histogram is used for grouped data. In a histogram, the rectangles are drawn with widths equal to the class sizes and heights proportional to the class frequencies. There are no gaps between the rectangles. If the class intervals are unequal, the heights must be adjusted so that the area of each rectangle is proportional to the frequency.
A frequency polygon is drawn by plotting the class marks against the frequencies, joining the points with straight lines, and connecting the ends to the base line. It can also be drawn from a histogram by joining the midpoints of the top of each rectangle.
The mean of a set of observations is their sum divided by the number of observations. Mean = (sum of all observations)/(number of observations)
For grouped data, the mean is calculated as the sum of (class mark x frequency) divided by the sum of frequencies: Mean = sum (f x x) / sum f, where x is the class mark and f is the frequency.
The median is the middle value of a set of observations arranged in ascending or descending order. - If the number of observations n is odd, the median is the ((n + 1)/2)-th observation. - If n is even, the median is the average of the (n/2)-th and the ((n/2) + 1)-th observations.
The mode is the value that occurs most frequently in a data set. A data set may have one mode, more than one mode, or no mode at all (when every value occurs equally often). For grouped data, the mode lies in the class with the highest frequency, called the modal class.
Statistics is used in every field: economics uses it for calculating national income and prices, the government uses census data for planning, science uses it for analysing experiments, and sports use it to measure performance. The mean gives a typical value of the data, the median is not affected by extreme values, and the mode tells us the most common value. Choosing the right measure depends on the question we want to answer.
| Term | Meaning |
|---|---|
| Class interval | Group of values like 10-20 |
| Lower limit | Smallest value in a class |
| Upper limit | Largest value in a class |
| Class size | Upper limit - lower limit |
| Class mark | (Lower limit + upper limit)/2 |
| Range | Highest value - lowest value |
| Frequency | Number of times a value occurs |
| Measure | Definition |
|---|---|
| Mean | Sum of observations divided by number of observations |
| Median | Middle value when data is arranged in order |
| Mode | Value that occurs most frequently |
| Mean (grouped) | sum (f x x) / sum f |
In this chapter we studied statistics, the science of handling numerical data. We learnt how to collect primary and secondary data, organise raw data into frequency distributions, and use class intervals, class marks and ranges for grouped data. We learnt to represent data graphically using bar graphs, histograms and frequency polygons, and to interpret these graphs. Finally, we studied the three measures of central tendency: the mean, the median and the mode, including their computation for ungrouped and grouped data. Statistics connects mathematics with everyday life, and the ability to summarise data with a single representative value is one of the most useful mathematical skills for the future.