Showing posts with label Sampling Distribution. Show all posts
Showing posts with label Sampling Distribution. Show all posts

Friday, November 23, 2012

Sampling Distribution of Proportion

Introduction to Sampling and Sampling Distributions:

Often in food stalls, the shoppers often taste a small piece of an item and based on their experience with the small piece they decide to buy or not to buy. This is an example of sampling. This is done in many factories as well. Consider a tyre manufacturing factory. The quality inspection engineers pull out a few manufactured tyres and test the tyres thoroughly and in that process these tyres are destroyed.

Note in the previous two cases why was the sampling done? If the shoppers in the food stall had tasted the entire quantity available, there would have been nothing left for sale. Similarly if the quality inspector had tested all the tyres, all the tyres would have been destroyed and there would have been nothing available to sell. So sampling for inspection was unavoidable in these cases.

Consider another case where you want to study about the average income of graduates in a particular region. If we try to get information from all the graduates in the region, the process will be time consuming and would be elaborate. However, if we collect data from a representative group of people, the same can be done quickly and easily

In the above example we deduce information about a bigger set based on a sample. The bigger set is called the population. We infer the data of the population based on the sample data. In the above cases, the total amount of the food in the food stall, total tyres produced by the factory, total graduates in the region are the population. The total tyres produced by

If the above samples are taken repeatedly and the mean of each sample is plotted is plotted against its probability then the distribution obtained is called sampling distribution.

Sampling Distribution of Proportions:

A sampling distribution described above can be partially described by the mean and the standard deviation. If we collect samples repeatedly from a population and calculate the mean, there will be difference in means for different samples. The means will be different. This is because each sample has different sampling elements from the population. This variation is called the standard deviation of the distribution of sample means or simply as standard error of the means. Similarly, the standard deviation of each sample might vary and this variation in the standard deviation is called the standard deviation of the sampling distribution or simply as standard errors.

The aim of taking samples is to estimate the population characteristics like mean and standard deviations. If we take samples repeatedly then the values of mean and standard deviation for each sample will be different. If so, which one will we consider as an accurate representation of the population data?  The standard error tells us how reliable it will be if we predict the population values from the sample statistics. There is an important theorem correlating the characteristics of the sample and the population and this is called the central limit theorem

Central Limit Theorem says the mean of the sampling distribution will be equal to the population mean regardless of the sample size. The sampling distribution of the means will approach normality as the sample size increases. This theorem helps to make inference about the population without knowing much about the population or the distribution of population.Understanding need help with math problems is always challenging for me but thanks to all math help websites to help me out.

This means the sampling distribution means will be the population means and the standard deviation of the population is accurately estimated if we use a large sample size. This is illustrated in the figure below. As we can see form the graph, as the sample size increase the mean is the same but the standard deviation decreases.



The process of evaluating the population data from the sample is called estimation. Central limit theorem is the basis on which the theories of estimations have been propounded. The theory of estimation helps us to find the population values of mean and standard deviation form the sample values.

Statisticians often use a sample to estimate the proportion of occurrences in a population. For example if we want to measure the unemployment rate or the proportion of unemployed people in the population. When we make try to estimate the proportion of a population from the sample we call it as a estimate of the population proportional and the sampling distribution is the sampling distribution of the proportion.

Exercises on Sampling Distribution:

Q:1 A machine is supposed to fill on an average 125 grams of a liquid in a bottle with a standard deviation of 20 grams. A random quality inspection shows a mean of 130. The inspector concludes that the sample is wrong. Is he correct?

Ans: No. The sample data need not accurately represent the mean. So just because the sample value was 130 grams does not mean it is a wrong sample.

Q:2 A sample of patients with tooth diseases wanted to be carried out regarding their brushing and eating habits. The statistician approaches a group of dentists and asks them to submit data. Each dentist takes data from 50 patients and submits the average value to the statistician. The statistician draws inference based on this. Was this a sampling of the patients?

Ans: No. The statistician used the data from the dentists, which was a mean of the 50 patients and not individual patient. So the statistician had plotted a sampling distribution and not a sample value.

Prob 3: What sample size will give you mean value equal to the population mean and without any error?

Ans: The sample size should be the same as population. This means that we need to carry out the data collection for the entire population.