A group of SNPs in a region of a genome may be in high linkage disequilibrium (LD). In such case, one SNP, called a tag SNP, represents the whole group. As sequencing SNPs is costly, often only the tag SNP, instead of the all the SNPs in the group, is sequenced to find genetic variation that may be associated with a phenotype.
Singleton SNP:
Sometimes one tag SNP represents only itself i.e., it is not in high LD with any other SNPs in that region. Such tag SNP is called a singleton SNP.
Homologs, Orthologs, and Paralogs - these 3 terms are conceptually related. It is necessary to understand the distinction among them.
Homology means that two genes are related by descent i.e. they have a common ancestral DNA sequence. Homology can be divided into two parts - Orthology and Paralogy. Orthologs are results of speciation, while Paralogs are results of gene duplication.
Orthology and Paralogy can easily be determined from the ancestral tree. You have to track along the vertical line of descent and find the place where the pair of genes join. If they join at an upside-down 'Y' node, then they are orthologs. In contrast, if they join at a horizontally connected node, then they are paralogs.
Here is an example [1]. In Fig. (a):
A1 has 5 Orthologs - B1, B2, C1, C2, C3, as all five orthologs join with A1 at an inverted 'Y' node, where speciation occurred.
B1 & B2 are paralogs, as they meet horizontally, where gene duplication took place.
B1 & C1 are orthologs.
C1, C2 & C3 are paralogs to each other.
Fig. (b) and Fig. (c) are actually same. (c) is actually a detailed illustration of (b). Here:
A1 & A2 (and also B1 & B2) are orthologs.
A1 & B1; A1 & B2; A2 & B1; A2 & B2 are all paralogs.
Identification of orthologs can play a significant role to determine evolutionary history. Usually, orthologs have the same function as their common ancestor, while paralogs do not. Bioinformaticians often take advantage of this behavior to differentiate orthologs from paralogs. But, functional similarity (or dissimilarity) does not necessarily imply orthologs (or paralogs).
Reference:
Jensen,R.A. (2001) Orthologs and paralogs - we need to get it right. Genome Biol., 2(8). [link]
রেস্ট্রিকশন এনজাইম (Restriction Enzyme) নামে কিছু প্রোটিন আছে যা দিয়ে ডিএনএ কে কিছু নির্দিষ্ট জায়গায় কাটা যায়। ব্যাক্টেরিয়া থেকে খুব কম খরচেই রেস্ট্রিকশন এনজাইম তৈরি করা যায়।
একেকটা ব্যাক্টেরিয়া একেকটা নির্দিষ্ট ক্রমে (সিকোয়েন্সে) ডিএনএ কাটতে পারে, একে বলে রিকগনিশন সাইট (Recognition site) বা ক্লিভেজ সাইট (Cleavage site)। রিকগনিশন সাইট সাধারণত ৪, ৬ কিংবা ৮ টা ক্ষার নিয়ে তৈরি।
ডিএনএ কাটা যেমন একটা নির্দিষ্ট ক্রম থেকে শুরু হয়, তেমনি একটা নির্দিষ্ট ক্রমে শেষ হয়। তবে বেশির ভাগ ক্ষেত্রেই ডিএনএর দুইটা সুতা শেষ প্রান্তে অসমানভাবে কাটে। অসমান দুই সুতার শেষ প্রান্তে একে অপরের পরিপূরক ক্রম থাকে, যা উপযুক্ত পরিবেশে হাইড্রোজেন বন্ধন তৈরির মাধ্যমে খাপে-খাপে মিলে যেতে পারে। এই দুই প্রান্তকে তাই বলে স্টিকি প্রান্ত (Sticky end)।
(সৌজন্যে: Essentials of Medical Genomics by Stuart M. Brown)
দুইটা স্টিকি প্রান্তের মধ্যে ফসফেট বন্ধন তৈরির মাধ্যমে জোড়া লাগানোর জন্য ডিএনএ লাইগেজ (DNA Ligase) নামের ব্যাক্টেরিয়াল এনজাইমও আছে। এই জোড়া লাগানোকে বলে ডিএনএ পেস্টিং।
এভাবে রেস্ট্রিকশন এনজাইম ও ডিএনএ লাইগেজ ব্যবহার করে মাধ্যমে দুইটা ভিন্ন প্রজাতির ডিএনএ কাট-পেস্টের মাধ্যমে নতুন কৃত্রিম কম্বিনেশন তৈরি করা যেতে পারে।
R is a powerful statistical tool which is heavily used in bioinformatics. Bioconductor, a very useful tool for high throughput genomic data analysis, has been developed using R. Here are two important links for learning R and Bioconductor.
When microarray data is obtained from multiple arrays, it is necessary to normalize the dataset to avoid variation due to different environments. There are several normalization techniques available in the literature. For example, Lowess normalization, Quantile normalization etc. Among these, quantile normalization is the current favorite method applied on microarray analysis.
Transformation:
Besides normalization, it is also beneficial to transform the data to correctly treat both up- and down-regulated data. The most widely used transformation technique is the logarithmic base 2. Notably, logarithms treat numbers and their reciprocals symmetrically. For example: log2(1) = 0, log2(2) = 1, log2(1/2) = -1, log2(4) = 2, log2(1/4) = -2.
Filtering:
If the intensity of hybridization in microarray is low (close to the background), then usually relative error becomes high. The common practice is to filter out (discard) the array elements which are statistically significantly different from the background.
References:
1. Slonim
DK,
Yanai
I
(2009)Getting Started in Gene Expression Microarray Analysis.PLoS Comput Biol 5(10):e1000543.doi:10.1371/journal.pcbi.1000543
2. Quackenbush, J. (2002) Microarray data normalization and transformation. Nature Genetics. Vol.32 supplement pp496-501.
For hypothesis testing, P-value is a highly used metric. The target is to have a very low (less than significance level) P-value i.e. to lower the probability of getting the observation by chance. Few demonstrations are available in the web.
One important point about P-value is that it is statistically valid when single score is computed. But in genomics, usually thousands of genes or millions of SNP or other scores are tested, which means that the calculated P-value is the probability of observation by chance using large number of scores. So, P-value threshold has to be justified, as it is valid only for one score.
The most widely used method for multiple testing correction is Bonferroni correction, which divides the significance threshold (α) by the number of tests (n). From a Bonferroni adjusted significance threshold α=0.01, we can be sure that none of the scores would be observed by chance from the null hypothesis. This is a usually a too strict adjustment.
Rather than saying that we want to be 99% sure that none of the observed scores is drawn from the null hypothesis, it is frequently sufficient to get a set of scores a little percentage of which may be drawn from the null hypothesis. This is actually the basis of False Discovery Rate (FDR) estimation. For some score threshold t, let Sobs is the number of observed score >= t, and Snull is that of null score >= t, then FDR is defined as
FDR = Sobs / Snull.
A limitation of FDR is further addressed in another metric, q-value which is defined as the minimum FDR attained at or above a given score.
Then the question arises, is Bonferroni correction, which is most widely used, of any use in any circumstance? The answer actually depends on the tradeoff between the costs and benefits associated with false positive and false negative. The guideline is: if
follow-up analyses depend upon group of scores and a little fixed percentage of error is tolerable, then FDR analysis is appropriate. Otherwise, when if follow-up focus on a single example, then the Bonferroni adjustment is more appropriate.
Reference: Noble, W.S. How does multiple testing correction work? Nature Biotechnology27, 1135-1137 (2009).