Showing posts with label Genomics. Show all posts
Showing posts with label Genomics. Show all posts

March 30, 2015

Tag SNP and Singleton SNP

Tag SNP:

A group of SNPs in a region of a genome may be in high linkage disequilibrium (LD). In such case, one SNP, called a tag SNP, represents the whole group. As sequencing SNPs is costly, often only the tag SNP, instead of the all the SNPs in the group, is sequenced to find genetic variation that may be associated with a phenotype.

Singleton SNP:

Sometimes one tag SNP represents only itself i.e., it is not in high LD with any other SNPs in that region. Such tag SNP is called a singleton SNP.

References:

1) Wikipedia, Tag SNP, accessed on 29 March 2015.
2) Xiayi Ke et al. (2008), Singleton SNPs in the human genome and implications for genome-wide association studies, European Journal of Human Genetics, 16, 506–515.

May 1, 2014

Homologs, Orthologs, and Paralogs

Homologs, Orthologs, and Paralogs - these 3 terms are conceptually related. It is necessary to understand the distinction among them.

Homology means that two genes are related by descent i.e. they have a common ancestral DNA sequence. Homology can be divided into two parts - Orthology and Paralogy. Orthologs are results of speciation, while Paralogs are results of gene duplication.

Orthology and Paralogy can easily be determined from the ancestral tree. You have to track along the vertical line of descent and find the place where the pair of genes join. If they join at an upside-down 'Y' node, then they are orthologs. In contrast, if they join at a horizontally connected node, then they are paralogs.





Here is an example [1]. In Fig. (a):

  • A1 has 5 Orthologs - B1, B2, C1, C2, C3, as all five orthologs join with A1 at an inverted 'Y' node, where speciation occurred.
  • B1 & B2 are paralogs, as they meet horizontally, where gene duplication took place.
  • B1 & C1 are orthologs.
  • C1, C2 & C3 are paralogs to each other.

Fig. (b) and Fig. (c) are actually same. (c) is actually a detailed illustration of (b). Here:

  • A1 & A2 (and also B1 & B2) are orthologs.
  • A1 & B1; A1 & B2; A2 & B1; A2 & B2 are all paralogs.
Identification of orthologs can play a significant role to determine evolutionary history. Usually, orthologs have the same function as their common ancestor, while paralogs do not. Bioinformaticians often take advantage of this behavior to differentiate orthologs from paralogs. But, functional similarity (or dissimilarity) does not necessarily imply orthologs (or paralogs).

Reference:
  1. Jensen,R.A. (2001) Orthologs and paralogs - we need to get it right. Genome Biol., 2(8). [link]
Other Sources:

June 13, 2013

Computational Biology Problems and Challenges

Link to Few Computational Biology Problems and Challenges:
  1. https://genomeinterpretation.org/
  2. http://www.ncbi.nlm.nih.gov/books/NBK25461/#_a2000dec5ddd00344_
  3. http://www.youtube.com/watch?v=bVhOntMCmnQ
  4. http://www.frontiersin.org/Bioinformatics_and_Computational_Biology/10.3389/fgene.2011.00060/full
  5.  http://www.stats.ox.ac.uk/research/genome/teaching2/topics_in_computational_biology

Research articles on automatic pathway construction


Research articles on automatic pathway construction:
  1. Inferring gene regulatory networks from gene expression data by path consistency algorithm based on conditional mutual information (2011, citation:8)
  2. Biological Pathway Extension Using Microarray Gene Expression Data (2008, citation:1)
  3. Microarray analysis of gene expression: considerations in data mining and statistical treatment (2006, citation: 66)
  4. Genomic analysis of metabolic pathway gene expression in mice (2005, citation: 67)
  5. PathMAPA: a tool for displaying gene expression and performing statistical tests on metabolic pathways at multiple levels for Arabidopsis (2003, citation: 54)
  6. Bayesian Consensus Pathway Construction and Expansion Using Microarray Gene Expression Data (From NCBI)
  7. Biological Networks (Software from UCSD)
  8. Reconstructing dynamic gene regulatory networks from sample-based transcriptional data (2012, citation:3)
  9. Reconstructing regulatory networks from the dynamic plasticity of gene expression by mutual information (2013, citation:0) 
  10. A Gaussian graphical model for identifying significantly responsive regulatory networks from time series gene expression data (2012, citation:1) 
  11. An integrative genomics approach to the reconstruction of gene networks in segregating populations (2004, citation:135) 
  12. Integrating genetic and gene expression data: application to cardiovascular and metabolic traits in mice () 
  13. Uncovering regulatory pathways that affect hematopoietic stem cell function using 'genetical genomics' (Nature Genetics, 2005, citation:333) 
  14. Inferring gene transcriptional modulatory relations: a genetical genomics approach (2005, citation:77) 
  15. Complex trait analysis of gene expression uncovers polygenic and pleiotropic networks that modulate nervous system function (Nature Genetics, 2005, citation:515) 
  16. Integrated transcriptional profiling and linkage analysis for identification of genes underlying disease (Nature Genetics, 2005, citation:410) 
  17. An integrative genomics approach to infer causal associations between gene expression and disease (Nature Genetics, 2005, citation:544)

August 31, 2012

DNA Cut-Paste

  • রেস্ট্রিকশন এনজাইম (Restriction Enzyme) নামে কিছু প্রোটিন আছে যা দিয়ে ডিএনএ কে কিছু নির্দিষ্ট জায়গায় কাটা যায়। ব্যাক্টেরিয়া থেকে খুব কম খরচেই রেস্ট্রিকশন এনজাইম তৈরি করা যায়। 
  • একেকটা ব্যাক্টেরিয়া একেকটা নির্দিষ্ট ক্রমে (সিকোয়েন্সে) ডিএনএ কাটতে পারে, একে বলে রিকগনিশন সাইট (Recognition site) বা ক্লিভেজ সাইট (Cleavage site)। রিকগনিশন সাইট সাধারণত ৪, ৬ কিংবা ৮ টা ক্ষার নিয়ে তৈরি।
  • ডিএনএ কাটা যেমন একটা নির্দিষ্ট ক্রম থেকে শুরু হয়, তেমনি একটা নির্দিষ্ট ক্রমে শেষ হয়। তবে বেশির ভাগ ক্ষেত্রেই ডিএনএর দুইটা সুতা শেষ প্রান্তে অসমানভাবে কাটে। অসমান দুই সুতার শেষ প্রান্তে একে অপরের পরিপূরক ক্রম থাকে, যা উপযুক্ত পরিবেশে হাইড্রোজেন বন্ধন তৈরির মাধ্যমে খাপে-খাপে মিলে যেতে পারে। এই দুই প্রান্তকে তাই বলে স্টিকি প্রান্ত (Sticky end)।
 (সৌজন্যে: Essentials of Medical Genomics by Stuart M. Brown)
  • দুইটা স্টিকি প্রান্তের মধ্যে ফসফেট বন্ধন তৈরির মাধ্যমে জোড়া লাগানোর জন্য ডিএনএ লাইগেজ (DNA Ligase) নামের ব্যাক্টেরিয়াল এনজাইমও আছে। এই জোড়া লাগানোকে বলে ডিএনএ পেস্টিং।
  • এভাবে রেস্ট্রিকশন এনজাইম ও ডিএনএ লাইগেজ ব্যবহার করে মাধ্যমে দুইটা ভিন্ন প্রজাতির ডিএনএ কাট-পেস্টের মাধ্যমে নতুন কৃত্রিম কম্বিনেশন তৈরি করা যেতে পারে।

July 20, 2012

Learning R and Bioconductor

R is a powerful statistical tool which is heavily used in bioinformatics. Bioconductor, a very useful tool for high throughput genomic data analysis, has been developed using R. Here are two important links for learning R and Bioconductor.
  1. http://www.cyclismo.org/tutorial/R/index.html
  2. http://manuals.bioinformatics.ucr.edu/home/R_BioCondManual

Besides, the lectures by Dr. Roger D Peng, Associate Professor, Johns Hopkins University in the Computing for Data Analysis course are very helpful. The lectures are available in youtube - http://www.youtube.com/watch?v=EiKxy5IecUw&list=PL7Tw2kQ2edvr2lv8FTvg9msf8YHQz0MS0

Happy learning!

July 19, 2012

Preprocessing of microarray data

Normalization:

When microarray data is obtained from multiple arrays, it is necessary to normalize the dataset to avoid variation due to different environments. There are several normalization techniques available in the literature. For example, Lowess normalization, Quantile normalization etc. Among these, quantile normalization is the current favorite method applied on microarray analysis.

Transformation:

Besides normalization, it is also beneficial to transform the data to correctly treat both up- and down-regulated data. The most widely used transformation technique is the logarithmic base 2. Notably, logarithms treat numbers and their reciprocals symmetrically. For example: log2(1) = 0, log2(2) = 1, log2(1/2) = -1, log2(4) = 2, log2(1/4) = -2.

Filtering:

If the intensity of hybridization in microarray is low (close to the background), then usually relative error becomes high. The common practice is to filter out (discard) the array elements which are statistically significantly different from the background.

References:
1. Slonim DK, Yanai I (2009) Getting Started in Gene Expression Microarray Analysis. PLoS Comput Biol 5(10): e1000543. doi:10.1371/journal.pcbi.1000543
2. Quackenbush, J. (2002) Microarray data normalization and transformation. Nature Genetics. Vol.32 supplement pp496-501.

March 21, 2012

Multiple testing correction in genomics

For hypothesis testing, P-value is a highly used metric. The target is to have a very low (less than significance level) P-value i.e. to lower the probability of getting the observation by chance. Few demonstrations are available in the web.



One important point about P-value is that it is statistically valid when single score is computed. But in genomics, usually thousands of genes or millions of SNP or other scores are tested, which means that the calculated P-value is the probability of observation by chance using large number of scores. So, P-value threshold has to be justified, as it is valid only for one score.

The most widely used method for multiple testing correction is Bonferroni correction, which divides the significance threshold (α) by the number of tests (n). From a Bonferroni adjusted significance threshold α=0.01, we can be sure that none of the scores would be observed by chance from the null hypothesis. This is a usually a too strict adjustment.

Rather than saying that we want to be 99% sure that none of the observed scores is drawn from the null hypothesis, it is frequently sufficient to get a set of scores a little percentage of which may be drawn from the null hypothesis. This is actually the basis of False Discovery Rate (FDR) estimation. For some score threshold t, let Sobs is the number of observed score >= t, and Snull is that of null score >= t, then FDR is defined as

FDR = Sobs / Snull.

A limitation of FDR is further addressed in another metric, q-value which is defined as the minimum FDR attained at or above a given score.

Then the question arises, is Bonferroni correction, which is most widely used, of any use in any circumstance? The answer actually depends on the tradeoff between the costs and benefits associated with false positive and false negative. The guideline is: if
 follow-up analyses depend upon group of scores and a little fixed percentage of error is tolerable, then FDR analysis is appropriate. Otherwise, when if follow-up focus on a single example, then the Bonferroni adjustment is more appropriate.

Reference: Noble, W.S. How does multiple testing correction work? Nature Biotechnology 27, 1135-1137 (2009).