Showing posts with label Statistics. Show all posts
Showing posts with label Statistics. Show all posts

February 26, 2015

Scale parameter in probability distributions

A parameter s is called a scale parameter of a probability distribution when it follows the following property:

F(x|s,θ) = F(x/s | 1,θ)

Here, F is the cumulative distribution function (cdf) and θ represents other parameters in the distribution.

When the probability distribution function (pdf) is defined for all values, then s must follow the following property:

fs(x) = f(x/s)/s

Here, f is pdf of the standard distribution, and  fs is pdf of the scaled distributions.

Intuitively,  higher scale parameter means higher spread of the distribution. Standard normal distribution with different scales are shown in the following plot.


And, here is matlab code to generate the above plot.

scales = [5,2,1,0.5,0.3];
x = -10:.001:10;
for s = scales
    y = normpdf(x./s) ./ s ;
    plot(x,y);
    hold on
end
legend('scale=5', 'scale=2', 'scale=1', 'scale=0.5', 'scale=0.3')
title('Standard normal distribution with different scale parameters')

November 15, 2014

Useful links describing statistical concepts

July 19, 2012

Preprocessing of microarray data

Normalization:

When microarray data is obtained from multiple arrays, it is necessary to normalize the dataset to avoid variation due to different environments. There are several normalization techniques available in the literature. For example, Lowess normalization, Quantile normalization etc. Among these, quantile normalization is the current favorite method applied on microarray analysis.

Transformation:

Besides normalization, it is also beneficial to transform the data to correctly treat both up- and down-regulated data. The most widely used transformation technique is the logarithmic base 2. Notably, logarithms treat numbers and their reciprocals symmetrically. For example: log2(1) = 0, log2(2) = 1, log2(1/2) = -1, log2(4) = 2, log2(1/4) = -2.

Filtering:

If the intensity of hybridization in microarray is low (close to the background), then usually relative error becomes high. The common practice is to filter out (discard) the array elements which are statistically significantly different from the background.

References:
1. Slonim DK, Yanai I (2009) Getting Started in Gene Expression Microarray Analysis. PLoS Comput Biol 5(10): e1000543. doi:10.1371/journal.pcbi.1000543
2. Quackenbush, J. (2002) Microarray data normalization and transformation. Nature Genetics. Vol.32 supplement pp496-501.

March 21, 2012

Multiple testing correction in genomics

For hypothesis testing, P-value is a highly used metric. The target is to have a very low (less than significance level) P-value i.e. to lower the probability of getting the observation by chance. Few demonstrations are available in the web.



One important point about P-value is that it is statistically valid when single score is computed. But in genomics, usually thousands of genes or millions of SNP or other scores are tested, which means that the calculated P-value is the probability of observation by chance using large number of scores. So, P-value threshold has to be justified, as it is valid only for one score.

The most widely used method for multiple testing correction is Bonferroni correction, which divides the significance threshold (α) by the number of tests (n). From a Bonferroni adjusted significance threshold α=0.01, we can be sure that none of the scores would be observed by chance from the null hypothesis. This is a usually a too strict adjustment.

Rather than saying that we want to be 99% sure that none of the observed scores is drawn from the null hypothesis, it is frequently sufficient to get a set of scores a little percentage of which may be drawn from the null hypothesis. This is actually the basis of False Discovery Rate (FDR) estimation. For some score threshold t, let Sobs is the number of observed score >= t, and Snull is that of null score >= t, then FDR is defined as

FDR = Sobs / Snull.

A limitation of FDR is further addressed in another metric, q-value which is defined as the minimum FDR attained at or above a given score.

Then the question arises, is Bonferroni correction, which is most widely used, of any use in any circumstance? The answer actually depends on the tradeoff between the costs and benefits associated with false positive and false negative. The guideline is: if
 follow-up analyses depend upon group of scores and a little fixed percentage of error is tolerable, then FDR analysis is appropriate. Otherwise, when if follow-up focus on a single example, then the Bonferroni adjustment is more appropriate.

Reference: Noble, W.S. How does multiple testing correction work? Nature Biotechnology 27, 1135-1137 (2009).

March 16, 2012

z-statictics vs t-statistics

z-statistics use z-score which defines z-score as how many standard deviation away from the mean.

z-score = (µ-x)/σ

But, as µ and σ are essentially mean and standard deviation of population, not of sample, those have to be estimated. If, the sample size is large enough (usually at least greater than or equal to 30) µ is estimated as sample mean while σ is estimated with denominator (n-1) instead of n in the mathematical definition of standard deviation to avoid biasing. As I cannot write mathematical equations here, I am giving the link of wiki http://en.wikipedia.org/wiki/Unbiased_estimation_of_standard_deviation.

According to central limit theorem, standard deviation of the sample distribution of samples is estimated as s/√n, where s is the sample standard deviation. And then using this σ z-score is calculated and 78-95.4-99.7 rule is applied to calculate the desired probability.

But, when sample size is not large enough (< 30), then the estimation of σ is under-estimated. Consequently, z-score (or z-statistics) does not work. In this case, t-statistics is used.

In a summary, if sample size is large enough (>= 30), use z-statistics, otherwise use t-statistics.

Here is the video from Khan Academy.


Central Limit Theorem

A little formally, probability distribution of the sum (or average) of i.i.d. variables with finite variance approaches a normal distribution.

informally, say we have a probability distribution of integers. now, if we randomly take N numbers and then calculate then mean of those taken numbers (say it "sample mean"). Then the distribution of "sample mean" will be a normal distribution (with same mean µ and standard deviation σ/√N ).

As N increases, the standard deviation will decrease and the distribution will be more like a true normal distribution.

A nice explanation from Khan Academy.