Comprehensive theory, key formulas, diagrams, and memory aids for Organisation of Data.
Once data have been collected, they are usually in the form of raw, unorganised numbers called raw data. A raw mass of figures is difficult to read, compare or interpret. The next step in a statistical study is therefore to organise the data - to classify and arrange them in a systematic way so that their essential features become clear. This process of organisation converts raw data into a frequency distribution, which is the basis for all further statistical analysis.
Organisation of data involves three related operations: editing, classification and tabulation. Editing removes errors and inconsistencies in the raw data. Classification groups the data into classes according to their common characteristics. Tabulation presents the classified data in the form of statistical tables with rows and columns. Together these operations make the data concise, comparable and ready for analysis.
This chapter deals with the meaning and types of classification, the variable and its types, the construction of a frequency distribution - including continuous frequency distributions, class limits, class intervals, class size and mid-values - and the ways of preparing discrete and continuous series from raw data.
Organisation of data is the process of arranging raw data in a systematic order to make them understandable and usable. Raw data are the data in their original, unprocessed form as collected from the field. For example, the daily collection of a milk dairy may be: 2, 3, 2, 4, 5, 2, 3, 4, 5, 2 ... In this form the figures convey very little.
The objectives of organising data are:
The two main steps of organisation are classification and tabulation. Classification is the grouping of data into classes or categories on the basis of common characteristics; tabulation is the presentation of this classified data in a table.
Classification may be of the following types, depending on the basis adopted:
A characteristic that takes different values in different individuals or units is called a variable. The value of the variable changes from unit to unit; for example, marks scored by students in a class, height of plants, or income of families.
Variables are of two types:
Because a continuous variable can take any value, its data are grouped into class intervals. For example, heights are expressed in ranges such as 150-155 cm, 155-160 cm, etc.
The frequency of a particular value of a variable is the number of times that value occurs in the data. If the value 3 occurs 5 times in a set of observations, its frequency is 5. The total of all frequencies is equal to the total number of observations.
A frequency distribution is a table that presents the values of a variable (or class intervals) along with their corresponding frequencies. Frequency distributions may be of two types:
Discrete (ungrouped) frequency distribution: In this distribution the individual values of a discrete variable are listed with their frequencies. For example, the number of members in 20 families may be presented as: 1, 2, 3, 4, 5 with frequencies 2, 5, 8, 3, 2.
Continuous (grouped) frequency distribution: In this distribution the range of a continuous variable is divided into class intervals, and the number of observations falling in each class is counted. For example, marks of students may be grouped as 0-10, 10-20, 20-30, and so on, with their frequencies.
A continuous frequency distribution is constructed through the following steps:
The important terms used in a continuous frequency distribution are:
A discrete series shows the individual values of a variable with their frequencies. A continuous series shows class intervals with their frequencies.
Class intervals may be of two types:
Exclusive classes: In exclusive classes, the upper limit of one class is the lower limit of the next class, and the upper limit of each class is excluded from that class but included in the next class. For example, in classes 0-10, 10-20, 20-30, the value 10 belongs to the class 10-20, not to 0-10. This type removes ambiguity about where a value belongs, and the class interval equals the difference between successive lower limits. The mid-value is correctly given by (lower + upper)/2.
Inclusive classes: In inclusive classes, both the lower and upper limits are included in the class, e.g. classes 0-9, 10-19, 20-29. Here there are no gaps between classes. When converting inclusive classes to exclusive classes for computation of mid-values, 0.5 is subtracted from the lower limits and 0.5 added to the upper limits.
The following series are commonly used:
| Type of Classification | Basis | Example |
|---|---|---|
| Chronological | Time | Production of wheat 2015-2025 |
| Geographical | Place | State-wise production of rice |
| Qualitative | Attributes | Sex, literacy, occupation |
| Quantitative | Numerical values | Marks, income, age |
| Term | Definition | Formula |
|---|---|---|
| Frequency | Number of times a value occurs | f |
| Range | Difference of max and min | Range = Max - Min |
| Class interval | Width of a class | Upper limit - Lower limit |
| Mid-value | Midpoint of a class | (L + U) / 2 |
| Number of classes | Sturges' rule | 1 + 3.322 log N |
graph TD
A["ORGANISATION OF DATA"] --> B["Classification"]
A --> C["Tabulation"]
A --> D["Variables"]
A --> E["Frequency distribution"]
B --> B1["Chronological - time"]
B --> B2["Geographical - place"]
B --> B3["Qualitative - attributes"]
B --> B4["Quantitative - numerical"]
D --> D1["Discrete - counting, whole numbers"]
D --> D2["Continuous - measurement, any value"]
E --> E1["Discrete frequency distribution"]
E --> E2["Continuous frequency distribution"]
E2 --> F["Class limits, interval, mid-value, frequency"]
E2 --> G["Exclusive vs inclusive classes"]
E --> H["Cumulative frequency - less than / more than"]
In this chapter we learned how to organise raw data so that it becomes meaningful. We studied the objectives of organisation, the four types of classification - chronological, geographical, qualitative and quantitative - and the concept of the variable with its two forms, discrete and continuous. We defined frequency and constructed discrete and continuous frequency distributions, learning the important terms of class limits, class interval, mid-value and cumulative frequency. We also understood the difference between exclusive and inclusive classes and how to convert inclusive classes into exclusive ones. This organised data, in the form of frequency distributions, is the raw material for the next stage of statistical analysis - the presentation of data through tables and diagrams - and for the computation of the statistical measures studied in the later chapters.