Comprehensive theory, key formulas, diagrams, and memory aids for Collection of Data.
Economic analysis begins with data. To study prices, incomes, employment, output or consumption, the economist needs reliable numerical facts. Data are collected from various sources using different methods, and the quality of any statistical conclusion depends on the quality of the data on which it is based. This chapter explains what data are, how they can be collected, what instruments and methods are used, and the important concepts of census, sampling and the sources of error in data collection.
Statistical data may be classified in two broad ways. First, on the basis of their source, data may be primary or secondary. Second, on the basis of how they are gathered, data may be collected through a census or through a sample. The choice of the appropriate source and method depends on the nature of the problem, the time available, the cost involved and the degree of accuracy required. Poorly designed data collection produces misleading results, however sophisticated the later analysis may be.
In this chapter we study the meaning and types of data, the instruments used in collecting primary data, the choice between census and sample, the essentials of a good questionnaire, the concepts of sample, pilot survey, sampling error and non-sampling error, and finally the sources of secondary data in India such as government publications, journals, and international agencies.
Primary data are those that are collected afresh and for the first time by the investigator directly from the field, for a specific purpose. For example, when the NSSO interviews households to know their monthly consumption expenditure, the responses obtained are primary data. Primary data are original in character and are more reliable, but their collection is expensive, time-consuming and requires trained investigators.
Secondary data are those that have already been collected and processed by some other agency for its own purpose, and are used by a new investigator for a different purpose. Government publications, reports of the Reserve Bank of India, census reports, journals and international sources like the World Bank are typical sources of secondary data. Secondary data are cheaper and quicker to obtain, but they may not be fully relevant or reliable for the investigator's specific purpose, and the investigator cannot control their quality.
The choice between primary and secondary data depends on the nature and objective of the enquiry, the availability of time and money, and the accuracy required. An investigator should prefer primary data when the required information is not available from existing sources or is outdated, and secondary data when reliable and relevant published data already exist.
Primary data can be collected by several methods. The important methods are:
The questionnaire is the most common instrument for collecting primary data. It is a set of questions, printed or typed in a definite order, relating to the problem under enquiry. A schedule is a set of questions printed on a sheet that is filled in by the investigator who personally interviews the respondents, whereas a questionnaire is generally filled by the respondents themselves.
The essentials of a good questionnaire are:
Pilot survey: Before the actual survey, a small-scale trial of the questionnaire is carried out on a few respondents to detect defects such as ambiguous questions, missing alternatives or excessive length. This trial run is called a pilot survey.
Census (complete enumeration): In a census, every unit of the universe or population is enumerated. The universe is the entire set of items from which data are collected; for example, all households of India. A census gives complete and accurate information about the whole population, as in the case of the decennial Census of India. But a census is expensive, time-consuming and requires a large trained staff, and in many cases it is impossible to reach every unit.
Sample survey: In a sample survey, only a part of the universe - called the sample - is studied, and the results are generalised to the whole population. Sampling saves time, money and effort, gives quicker results, and can achieve greater accuracy because a small number of units can be studied more intensively. In India, the NSSO uses sample surveys to estimate consumption expenditure, employment and poverty.
The choice between census and sample depends on:
The important methods of sampling are:
Random or probability sampling: In this method every unit of the universe has an equal chance of being selected, and units are chosen by the principle of chance without any bias. The lottery method and the use of random number tables are the common techniques. Random sampling may be simple random sampling, systematic random sampling, stratified random sampling or cluster sampling.
Simple random sampling: Units are selected directly from the whole universe, each having an equal chance of selection.
Stratified sampling: The universe is first divided into homogeneous strata or groups, and then a random sample is drawn from each stratum in proportion to its size. This improves the representativeness of the sample.
Non-random sampling: In this method units are selected not by chance but by the judgement of the investigator or by convenience. Judgement sampling, convenience sampling and quota sampling are non-random methods. Such samples are easy to draw but may be biased.
Errors in statistical data may arise at any stage of the enquiry. They are broadly of two types:
Sampling errors: These are the differences between the value of a parameter of the universe and the value obtained from the sample, arising purely because a part, and not the whole, of the universe is studied. Sampling errors arise only in sample surveys, are not due to carelessness, and can be reduced by increasing the size of the sample, or by adopting a more refined method of sampling, such as stratified sampling. A larger and more representative sample gives results closer to the universe values.
Non-sampling errors: These are errors that arise during collection, recording, tabulation and analysis of data, and they occur both in census and sample surveys. They may be due to the faulty selection of the sample, the carelessness of the investigator, incorrect answers by respondents, non-response, measurement errors or mistakes in tabulation and processing. Non-sampling errors cannot be eliminated by increasing the sample size; they can be reduced only by careful planning and execution of the survey.
The important sources of secondary data in India are:
| Basis | Primary Data | Secondary Data |
|---|---|---|
| Meaning | Collected afresh by the investigator | Already collected by someone else |
| Originality | Original in character | Not original |
| Cost | Expensive | Cheap |
| Time | Time-consuming | Quick |
| Reliability | More reliable | Less reliable |
| Control over quality | Investigator controls | Investigator cannot control |
| Method of Sampling | Meaning | Remark |
|---|---|---|
| Simple random | Every unit has equal chance | Unbiased |
| Systematic | First unit random, then fixed interval | Every kth unit |
| Stratified | Universe divided into strata, random from each | Most representative |
| Cluster | Whole groups chosen at random | Useful for large areas |
| Judgement | Selected by investigator's judgement | Non-random, biased |
graph TD
A["COLLECTION OF DATA"] --> B["Sources"]
A --> C["Methods of collection"]
A --> D["Census vs Sample"]
A --> E["Errors"]
B --> B1["Primary data"]
B --> B2["Secondary data"]
C --> C1["Direct personal investigation"]
C --> C2["Indirect oral investigation"]
C --> C3["Correspondents"]
C --> C4["Questionnaire / schedule"]
D --> D1["Census - complete enumeration"]
D --> D2["Sample - part of universe"]
D --> D3["Random sampling - equal chance"]
D --> D4["Non-random sampling - judgement"]
E --> E1["Sampling error - only in sample, reduced by larger sample"]
E --> E2["Non-sampling error - in both, reduced by care"]
This chapter explained the entire process of data collection, which is the foundation of all statistical work. We distinguished primary data, collected afresh for a specific purpose, from secondary data, which are already available from existing sources. We examined the methods of collecting primary data - direct and indirect investigation, correspondents, and questionnaires - and the essentials of a good questionnaire, including the role of the pilot survey. We compared census with sample surveys and saw when each is appropriate, studied the principal sampling methods, and distinguished sampling errors, which arise only in samples, from non-sampling errors, which can occur in any survey. Finally, we listed the important sources of secondary data in India. Together with the next chapter, where we learn to organise the collected data into frequency distributions, this chapter gives the student a complete picture of how reliable economic information is produced.