Showing posts with label Sequencing. Show all posts
Showing posts with label Sequencing. Show all posts

May 13, 2014

Sequence Alignments and Seed Design

Sequence alignment is an important task in genome analysis. It searches for similarity in DNA, RNA or Protein that may be a result of functional, structural or evolutionary relationships between sequences. Although the scientific community is researching on this topic for decades, the problem is not yet  solved. Because, sequences are extremely large - millions or billions of bases or amino acids. Due to run-time complexity, researchers typically uses heuristics to search for similarity. There is no way to measure if the current heuristics can find all or most of the true alignments. Frith & Noe showed that improved search heuristics can find new alignments [1]. What heuristics do they use?

Before going to the heuristics, proposed by Frith and Noe, let us briefly look at the existing approaches. Existing search methods generally follow a seed-and-extend approach. They first find short matches (seed), and then search for high-scoring alignments near each seed. Here, the sensitivity of alignment (quality to find all alignments) depends on seed design. Better seed design can result into better sensitivity.

Several kinds of seeds exist in current literature. The simplest seed is an exact match of a given length. A spaced seeds allow mismatches at certain positions. For example, 1110111101 is a seed that looks for a match of length 10, with possible mismatch at position 4 and 9. Transition-constrained seeds allow constrained mismatches at certain positions. Here unlike spaced seeds, they do not allow any kind of mismatch, rather they allow only transitions (A↔G, or T↔C), not transversions. For example, 11T011110T is a seed that looks for a 10-length match with possible mismatch at position 4 & 9 and possible transitions at position 5 & 10. There is a biological explanation behind allowing transitions. Transitions do not change the chemical structure drastically (no. of rings are same), and are less likely to substitute amino acids [2]. Transitions thus do not change the functionality drastically, while transversions usually do. As a result, transitions stay in the sequence as silent substitutions. In fact, transitions are more frequent than transversions, even though the random probability of transversions are double than that of transitions.

Both spaced and transition-constrained seeds look for a fixed length match. In contrast, adaptive seeds can have variable length. Adaptive seeds are lengthened until the frequency of matches in the target sequence becomes less than or equal to a threshold [3]. Let, a fixed length seed matched at 1 million positions of the target sequence. Then we have to perform the expensive alignment steps 1 million times. Adaptive seeds lengthen the seed size to reduce match frequency, and thus improve runtime. Sparse seeds can also be used to reduce match frequency. Sparse seeds do not look for a match at every position; instead they put a regular interval between each starting points (for e.g., search at every second or third positions).

In general, each parameter of seed design has a reciprocal effect on sensitivity and runtime. For example, small seeds have higher sensitivity but longer runtime than those of long seeds. Thus, it can be a good idea to observe the effect of different parameters. Frith & Noe did exactly the same thing in [1]; they analyzed the performance of different parameters using sensitivity and runtime. The following figure summarizes theirs results.



Frith & Noe found that transition-constrained seeds can be useful to find new alignments (see part-C in the above figure). They carefully designed seeds using different transition-transversion ratios. The transition-transversion ratio between human and dog is 3:2. While previous researches mostly used 1:1 ratio, Firth et. al. used 3:2. This generally reduces the error, keeping the runtime same. The authors found about 20,000 new alignments in this approach between human and mouse. All of these definitely do not signify functional or evolutionary relationships. The authors speculate a high probability to find new significant alignments. As most of the new alignments were unaligned, they speculate high rate of orthologs. This claim needs more support.


References:

[1] Frith,M.C. and Noé,L. (2014) Improved search heuristics find 20 000 new alignments between human and mouse genomes. Nucleic Acids Res., 42, e59.

[2] Carr, S.M. (2013), Transition versus Transversion mutations, https://www.mun.ca/biology/scarr/Transitions_vs_Transversions.html, Accessed on May 12, 2014.

[3] Kiełbasa,S.M. et al. (2011) Adaptive seeds tame genomic sequence comparison. Genome Res., 21, 487–93.






April 2, 2014

Short read sequencing or Long read sequencing?

Genome sequencing is a hot topic these days. Currently, the popular method of sequencing is to generate millions of short reads, typically 50 to 150 nucleotides long, and then assemble the reads in computational approach. Illumina, almost having a monopoly in sequencing business, follows this strategy. However, this strategy has some drawbacks. For example, it reads genome from multiple cells, and the biological signals in those cells are averaged to generate a consensus sequence. Consequently, it cannot identify the molecular-level biological differences. Moreover, this strategy does not work well with repetitive sequences or heterozygous sequences.

In contrast, long reads can be used for sequencing. These reads can be 100 times longer that short reads. Thus the long reads have fundamentally more information than short ones. Long reads can help uniquely map the reads in complex regions including repetitive elements. However, long reads currently suffers from an elevated error rate, about 15%. That means, one in every 7 or 8 bases is incorrect. Due to this limitation, long reads alone are yet not suitable for sequencing. However, a combination of short and long reads can perform much substantially better than any of the two methods.

Pacific Biosciences, a biotechnology company, focuses on long reads. They are trying to improve the error correction algorithm so that sequencing can be performed only from the long reads, without using the short reads. That would be a great achievement, as it would reduce the cost, and also enable identification of heterozygous and repetitive elements. Thus, we may expect that the the monopoly of Illumina would be reduced.

Another biotech company, Oxford Nanopore, is also in the race. They follow a different technology. They use the characteristic conductance change when single-stranded DNA passes through or near the nanopore, a small hole of the order of 1 nanometer in internal diameter. This strategy also produces long reads from single cell. Although this approach suffers from a high error rate, it has been shown in an experiment that more than 80% of the reads had perfect 50-nucleotides sections. This is impressive. If a proper error correction algorithm can be devised, Oxford Nanopore can be beat the dominance of Pacific Biosciences in the long-read field.

To make the scenario more interesting, GynapSys, another biotech company, aims at developing a small all-electronic instrument, like an iPad, that will perform all the sequencing steps, and thus reduce the sequencing time and cost.

Let's see which technology (or company) dominates the rest.

References:
1) Greenleaf,W.J. and Sidow,A. (2014) The future of sequencing: convergence of intelligent design and market Darwinism. Genome Biol., 15, 303.
2) http://www.genengnews.com/insight-and-intelligenceand153/the-long-and-the-short-of-dna-sequencing/77899725/
3) http://www.fluidigm.com/december-31-2013.html
4) http://allseq.com/knowledgebank/emerging-technologies/genapsys

April 1, 2014

Sequencing GWAS

A nice article on the difference between GWAS with SNPs and that with Sequencing.
http://massgenomics.org/2014/03/gwas-sequencing-realities.html