Wednesday, August 5, 2009

ADAM: Array-based Discovery of Adaptive Mutations

Source: Goodarzi H, Hottes AK, Tavazoie, 2009. Global discovery of adaptive mutations. Nature Methods, 581 - 583.

I had to wait for a long time to actually write about a project of my own. I had been thinking about this problem for at least 2 years before finding the
solution. The problem I wanted to address was a classic question: how can you map a mutation? If you do a selection and find a mutant which exerts your phenotype of interest, how would you go around finding it? The classic approach involves painstakingly going through a set of known markers around the genome and calculate the linkage between the known marker and your site of mutation. Then choose the closest marker and redo the linkage analysis with more markers close to that site, until the location becomes small enough for direct PCR amplification and sequencing. But this whole protocol, takes weeks if not months and it is quite labor-intensive. Especially, if your mutant carries multiple effective mutations, finding them becomes more complicated.

With the advent of whole-genome sequencing or re-sequencing of bacterial genomes can help us find all the mutations in the genome. But in many cases, it is not obvious whether a mutation is responsible in eliciting a certain phenotype or not. In other words, every single mutation should then be tested for its effect on the phenotype of interest.

To tackle this problem, we developed ADAM as a whole-genome
approach that enables you to find all the adaptive mutations (and not silent mutations) in a matter of days. ADAM, uses parallel, genome-wide linkage analysis to simultaneously identify all mutated loci with direct fitness contributions. ADAM has three components:
1. A library of selectable markers embedded in the DNA of the parental strain.
2. A mechanism for transferring markers from the parental strain's library into the evolved strain in such a way that DNA from the parental strain adjacent to the marker replaces the corresponding DNA in the evolved strain.
3. A method for measuring the frequency of markers throughout the genome.

As the first component, we used a high coverage library of kanamycin-marked transposon insertions all across the wild-type genome. We then trasnfer the markers from the WT genome to the evolved background using P1 transduction. In this step, if a transposon is close to a site of mutation, the recombination results in the correction of this mutation back to the WT genotype. Now, in this secondary transposon library in the mutant background, the markers whose introduction to the genome has been accompanied by the loss of an adaptive mutation will be at a relative disadvantage under the selective conditions. To find the location of these transposons, we compare the frequency of transposon insertion events in each locus under selective and non-selective conditions. Adaptive mutations can then be discovered in locations in which a stretch of loci show marker depletion upon selection. For this step, we use a microarray-based whole genome footprinting.


We used ADAM to find a known CmlR cassette as a proof-of-principle. We then used this method to identify adaptive mutations in lab-evolved strains (growth on Asn and ethanol tolerance).


Friday, July 3, 2009

We're all Iranians

Source: Nature 460, 11-12 (2 July 2009) | doi:10.1038/460011a; Published online 1 July 2009

Well, I'm an Iranian and the recent unfoldings in Iran has engaged me to the extent that I don't have time to write about science anymore. But a few days ago, I stumbled upon this editorial from Nature, which truly impressed me both as a scientist and as an Iranian.

I've never been a fan of politics. Politicians hold a reactionary mindset, aiming at simplifying the matters down to meaningless words; however, a scientist is bent towards evidence and proof. There is no question that any public movement in any country should be both acknowledged and encouraged. But, what can we, in the scientific community accomplish?
1. Universities are very well equipped to speak loud and in clear... making clear that Iranians are supported by non-governmental entities across the world.
2. The recent events accompanied by massive crackdowns on academics followed by mass resignations and discontent, would urge many young Iranian academics to leave Iran. The colleges, universities and research institutions all across the world can accomplish a great deal by prioritizing these researchers for a while or at least create a more active network to find suitable jobs for them (something similar to Scholars at Risk Network).

Wednesday, May 20, 2009

PNAG-driven mode of biofilm formation in E. coli

Source: Amini S, Goodarzi H, Tavazoie S. (2009). Genetic dissection of an exogenously induced biofilm in laboratory and clinical isolates of E. coli. PLoS Pathog. 2009 May;5(5):e1000432.

This is the first experimental project the progress of which I've observed from conception to completion. I witnessed the myriad of challenges that Sasan had to deal with to make the case for his hypotheses. And amazingly enough, he pulled it through. In experimental biology, it's difficult to claim victory at any point. Every given experiment prompts more experiments that are both labor-intensive and time-consuming. In computational biology, it is the model, the experiment or the approach that matters... it has to be novel. In experimental biology, on the other hand, what you need to do is clear (most of the time), it's just that matter of doing it. E.g. for testing the function of a gene you need to knock it out. Every one knows that. But dong it is the difficult part. This study was one of those cases were many corroborative experiments needed to be done, from microarray-based fitness profiling of mutant libraries to peptidoglycan extraction and visualization (polyacrylamide gels) to PNAG purification and validation (mass spectrometry). And remember that one lab does not actively use all these methodologies and sometimes for doing one experiment you have to set-up the whole experiment: ordering the reagents, making the solutions, optimizing conditions and ...

At this point, I actually feel ashamed. The fact that I reduce thousands of hours of work into a couple of sentences, passing judgement from the safety of my laptop bothers me (just a little though). I'm not going to do this for this particular study. If you're interested, read the whole paper (it is open access). I would just tell you that this is an example of a complete story, from start to finish and I congratulate Sasan for accomplishing this painstaking task.

Monday, April 20, 2009

Ribosome Profiling

Source: Ingolia et al (2009). Genome-Wide Analysis in Vivo of Translation with Nucleotide
Resolution Using Ribosome Profiling. Science 324: 218-223.

Proteomics is the game, but due to technical difficulties in the direct quantification of protein types in the cell, we have instead been using transcriptome measurements as a proxy for protein expression. This is a decent proxy but far from perfect. In this paper, the authors bring us a step closer to protein quantification through measuring the portion of transcriptome that is actually being translated. They combine the ribosome-mediated protection of traslated RNA molecules with the power of deep sequencing to determine the ribosome positionings at a single nt resolution.

What the authors found?
1. There is an excess of ribosomes bound to the first 30-40 codons; a quantity which drops substantially in the later codons.
2. uORFs are quite widespread, resulting in a high ribosome presence in the 5'UTR.

Tuesday, March 31, 2009

Revealing Genetic Interactions in E. coli

Source: Typas et al, (2008). High-throughput, quantitative analyses of genetic interactions in E. coli. Nature Methods 5, 781 - 787.

Genetic interaction studies in bacteria are quite challenging. Before this study, we had data for hundreds of interactions in
E. coli in comparison with thousands known for yeast. In this paper, the authors introduce a method termed GIANT which relies on massive Hfr conversion and double mutant generation in E. coli. The method is presented below as a figure from the original paper:
Here, we rely on conjugation to select for double markers (double mutants) and assay they growth on 384 or 1536 colony arrays. The authors make the case for their method through several validation steps. In the end, this method has the ability to generalize to other organisms for which deletion collections are available.

Friday, March 13, 2009

Proliferation-resistant biotechnology

Source: Nouri A., Chyba C., 2009. Proliferation-resistant biotechnology: an approach to improve biological security. Nature Biotechnology 27, 234 - 236.

A dear friend of mine Ali Nouri (also a AAAS congressional fellow) has published an interesting commentary in the current issue of Nature Biotech. Generally, we (the scientists) sometimes fail to grasp the security implications of science. Now, while the uninhibited progression of science is essential for our prosperity and avoiding regression back to another dark age, we should also try to find inexpensive and applivable ways to boost security.

Making an entire organism from scratch is not a dream anymore. This has been done in case of many viruses. The resurrection of the 1918 influenza virus caused a turmoil in our field. While we learned alot about the virus (e.g. how close it actually is to avian flu), many questioned whether thi
s type of research should be prohibited. Personally, I don't think any type of basic research should be prohibited because any thing may simply revolutionize our lives, but I agree that this information should be protected against misuse and abuse.
In this commentary, the authors have simply requetsed the companies to screen their bulk requests and raise a red flag if the requested gene or genome belongs to a list of dangerous oragnsisms or toxins. Steps as simple as this are very cheap to implement. And I'm sure many of you are already coming up with solutions for potential bypass of this problem. But if we put enough obstacles in the way of misusing these technologies, the accumulative security would actually synergistically increase and may very well pass the threshold for many ill-willed individuals.

Monday, February 9, 2009

Leading into my work...

Source: Lisec, J., Meyer, R.C., Steinfath, M., Redestig, H., Becher, M., Witucka-Wall, H., Fiehn, O., Torjek, O., Selbig, J., Altmann, T., and Willmitzer, L. Identification of metabolic and biomass QTL in Arabidopsis thaliana in a parallel analysis of RIL and IL populations. 2008. The Plant Journal, 53: 960-72

The authors created a number of RIL and IL lines in Arabidopsis and then ran targeted GC-MS on them. They were able to measure 181 compounds and find QTLs for 84, for a total of 157 QTLs. The contribution of these loci was between 1.7 and 52.1%. They found that many of these metabolites co-mapped, and that in nearly all of them a good candidate gene could be found that might explain the effect. They defined a candidate gene as a gene within the support interval in the direct pathway of the metabolite. None of these metabolite linkages showed a strong ability to change biomass.



Other notes:
-permutation test for candidate gene: randomly assign linkage to metabolite, sort through interval and see if any genes overlap with metabolite in AraCyc
=most metabolites showed no significance
=only 13 metabolites showed a higher than permutation-average number of candidate genes
-near impossible to find epistasis, found it only explained 2.72% of phenotypic variation on average
-nonrandom distribution of mQTLs, does not correlate with distribution of metabolic genes

Monday, February 2, 2009

Speed-genotyping

Source: Lai, C-Q., Leips, J., Zou, W., Roberts, J.F., Wollenberg, K.R., Parnell, L.D., Zeng, Z-B., Ordovas, J.M., and Mackay, T.F.C. Speed-mapping quantitative trait loci using microarrays. 2007. Nature Methods, 4(10): 839-41

The authors used microarrays to genotype a large number of individuals for a QTL study into longevity. Instead of individually genotyping and measuring the phenotype, the authors instead selected a subset of the population based on their phenotype (longevity). Then they pooled this subset’s DNA and ran it across a microarray that had oligos from both parents. They compared each marker hybridization with a young group that should be equally mixed for the alleles at each marker. A simple t-test was computed for each marker (with FDR correcting) to determine whether that marker had a skewed allele ratio between samples. Multiple QTLs were found, more so than using previous genotyping methods.

Wednesday, January 28, 2009

Noise propagation in transcription networks

Source: Dunlop et al (2008). Regulatory activity revealed by dynamic correlations in gene expression noise. Nature Genetics 40(12):1493-1498.

Biological events are stochastic in nature. Random fluctuations in protein concentration, expression and etc relays noise through the transcription network via the regulatory links. For example, a random decrease in the concentration of a repressor results in an increase in the expression of its target gene; however, only if the concentration of the repressor falls within an "active" range in which the expression of the target genes is sensitive to small chanages in the repressor content (see Fig. below).
Thus, observed correlations between the expression of different genes may be the result of a direct or indirect regulatory process. However, in addition to intrinsic noise (fluctuations in the expression of a given gene), we should also consider the extrinsic noise in which all the genes are uniformly affected by a given change (e.g. a random increase in the ribosome content of the cell increases the expression of all the genes). Extrinsic noise causes false positive correlation (see Fig. below).
Thus, any measurement of correlations must be normalized by the effect of extrinsic noises. In this paper, the authors use both stochastic modeling and experimental validation to make the case for this phenomenon.

Tuesday, January 13, 2009

MISSING: ATP!!

Source: Kresnowati, M.T.A.P., van Winden, W.A., Almering, M.J.H., ten Pierick, A., Ras, C., Knijnenburg, T.A., Daran-Lapujade, P., Pronk, J.T., Heijnen, J.J., and Daran, J.M. When transcriptome meets metabolome: fast cellular responses of yeast to sudden relief of glucose limitation. 2006. Molecular Systems Biology, 49

The authors subjected yeast held at steady-state low-glucose levels to a pulse of glucose and recorded their transcriptional and metabolic differences five minutes after the pulse. The most shocking discovery was the remarkable drop in AXP levels, led mainly by ATP. ATP was not simply converted to ADP, nor were AXPs converted for RNA incorporation, over 80% of AXP was unaccounted for after the pulse. Additionally, early-glycolytic metabolites climbed after the pulse but later-glycolytic metabolites sharply dropped. This was explained by the observed jump in NADH/NAD which would inhibit glyceraldehyde-3-phosphate dehydrogenase. With the switch from gluconeogenesis to glycolysis, these later compounds would flush into TCA or ethanol production but not be replenished until redox equilibrium in the cell was returned. On the transcriptome front, over 1000 genes were found to differ between at least two time points, differences didn’t begin until after 120s, though most until after 210s. The upregulated genes were enriched for ribosome biogenesis, amino acid metabolism and purine synthesis, all of the genes leading to adenine production through de novo synthesis, RNA degradation, sulfur metabolism, and conversion. The downregulated genes were enriched for C1-metabolism, energy reserves, and TCA. Additionally a number of genes in those pathways were found to have an order of magnitude lower half-lives for transcripts, from ~30 minutes to four! Looking at 3’, post-stop codon regions, the degraded genes nearly all shared in at least one of four regions that were abundantly found compared to chance.



Other notes:
-1154 genes significantly change
=K-means clustering into 5 groups
-CXP, UXP, and GXP levels also dipped but not on the same magnitude of AXP
-TCA intermediates increased, except citrate
=probably two separate branches: TCA and glyoxylate cycle
=TCA genes downregulated, glyoxylate genes upregulated


So yeah, it's cool that 1/6th of the genome changes its transcription. And yeah, it's interesting that there's an 8-fold difference in transcript half-lives. But WHERE DOES ALL THE ATP GO?!?! In case you're new to biology: ATP is one of the top 10 most used molecules (by number of reactions). This is like saying that upon the introduction to oxygen, humans lose 80% of their red blood cells and no one can see any dead red blood cells, they just vanish. If anyone knows any follow up studies that solved this conundrum, please send my way!

Tuesday, December 23, 2008

Divergent Initiation of Transcription

Source: Core et al. (2008). Nascent RNA sequencing reveals widespread pausing and divergent initiation at human promoters. Science 322:1845-1848.

In the current issue of science (Vol. 322), two back-to-back articles are published both reporting divergent transcription close to TSS. The authors have used global run-on sequencing to determine the site, amount and orientation of active RNApols. One of their main finding is this divergence in transcription.
How is this helpful?
1. Transcription leads to chromatin modifications that may be essential for dynamic expression (see my previous post)
2. Transcription may expose the binding sites that are otherwise engaged in nucleosomes.
3. The resulting negative supercoiling may benefit transcription in the region.

Monday, December 15, 2008

Gene Expresion Regulation: Chromatin Remodelling

Source: Hirota et al. (2008). Stepwise chromatin remodelling by a cascade of transcription initiatoin of non-coding RNAs. Nature 456:130-134.

The RNA-seq strategy has revolutionized our way of doing biology, but it has also complicated the way we used to look at gene expression regulation. First, it has been shown that a huge number of RNAs are produced without ever being translated. Many of these species are envisioned to participate in some sort of expression regulation... In this paper, the authors make the case for one such mechanism: firing from upstream promoters results in chromatin modifications that leads to the activation of the main promoter.


While studying the regulation on fbp 1+ in yeast, the authors observed that upon starvation it takes around 60 min for the main RNA to show up; however, during this period 3 other longer RNAs show up suggesting active upstream promoters (a, b and c in figure below). Using chromatin-IP for RNApolII, they confirmed the occupation of these upstream promoters upon activation. They also assyed chromatin remodelling using MNase assay to show that the chromatin is in fact modified upon activation.
The key point here, however, was the fact that upon cloning a transcription termination site between the upstream promoters and the main promoter inhibits activation... which means transcription is required for the observed chromatin remodelling. The figure below, from the original paper, shows the details of this mechanism.

Friday, December 12, 2008

Correlating Transcription and Cell Cycle

Source: Klevecz, R.R., Bolen, J., Forrest, G., and Murray, D.B. A genomewide oscillation in transcription gates DNA replication and cell cycle. 2004. PNAS, 101(5): 1200-5

The authors measured transcript abundance as it fluctuated with changes in dissolved oxygen content for yeast. They found that there were three timepoints were gene expression peaked: two peaks with >2,000 genes reaching their maximum expression when oxygen levels were high (cells nonrespiring) and one peak where 650 genes reached their maximum expression when oxygen levels were low (cells respiring). Compared transcripts to states, and found that mitochondrial genes are expressed during reductive phase when mitochondrial function is minimal; while sulfur metabolism genes are expressed in respiratory phase right before they are needed for DNA replication in beginning of reductive phase. Most periods were ~40 minutes, and other studies showed that on a variety of media the doubling times of yeast were some multiple of 40 minutes.



•Other notes:
-cell-to-cell synchronization involved through respiratory inhibition by H2S and phase shifts due to acetaldehyde
-87% of genes expressed maximally in reductive phase
=2400 early, 2200 late
-650 genes maximum expression in oxidative phase
-4-12 minute lag between transcript peak and maximum gene product function
-DNA replication begins abruptly at end of respiration, H2S levels rise
-separation in time between oxidative and reductive phases goes to transcript levels and is coordinated with DNA replication
=prevents oxidative stress

Tuesday, December 2, 2008

Metabogenome: Discovering compounds a genes!

Source: Keurentjes, J.J.B., Fu, J., Ric de Vos, C.H., Lommen, A., Hall, R.D., Bino, R.J., van der Plas, L.H.W., Jansen, R.C., Vreugdenhil, D., and Koornneef, M. The genetics of plant metabolism. 2006. Nature Genetics, 38(7): 842-9

The authors used LC-QTOF MS to create metabolite profiles for two parental strains of Arabidopsis thaliana and 160 RIL descendants. They then genotyped these strains and ran QTL analysis on the >2000 metabolites they measured to come up with a staggering number of potential QTLs. Interestingly, a large number of compounds were not detected in either parent strain but only in the RILs. QTL hotspots were found and a study on a specific hotspot and its linked metabolites’ pathway carried out. This analysis was able to find the relative position in the pathway between two loci. Additionally, an analysis on a hotspot with unknown metabolites gave a set of metabolites to classify. Once discovered, the distinction in phenotypes revealed the presence of a previously unrealized enzyme in one of the parents. The authors close by stating that pathway elucidation and identification are possible through this high-throughput analysis, as well as metabolite grouping for identification.

Other notes:
-75% of compounds were assigned a QTL
-853 of 2129 metabolites not detected in either parent
-QTL for 1592 metabolites, roughly 2 QTLs per compound
-all AOP-related metabolites also map to MAM, while few MAM-metabolites map to AOP suggests AOP is downstream of MAM
-correlation between masses were calculated based on QTL profiles: vectors of P-values associated with markers
-co-occurrence of well-known or unknown metabolites may reveal pathway information

Monday, December 1, 2008

Yeast Epistasis Map

Source: Roguev et al (2008). Conservation and Rewiring of Functional Modules Revealed by an Epistasis Map in Fission Yeast. Science 322:405.

Epistasis analysis is one of the most direct methods for defining functional relationships between genes and proteins. These interactions can be negative (synthetic lethality) or positive (suppression). Whole-genome high-throughput epistatic maps (E-MAP) were peviously published for S. cerevisiae; here, the authors focus on S. pombe. E-MAPs are generated through generating pairwise knock-outs and assaying their henotypes (usually growth in complex media), comparing them to the single-gene mutants. This E-map includes ~118,000 double mutants in 550 genes invloved in different aspects of cellular processes. In this set, similar to previous E-MAPs, the correlation between protein-protein interactions (PPI) and epistasis scores is apparent (see the figure below from the original paper).

The authors have also focused a great deal on dissecting the RNAi machinery in S. pombe. This study resulted in the identification of a novel component in this machinary (rsh1).

Thursday, November 6, 2008

Out with the old, in with the new : Allele Replacement

Source: Gray, M., Piccirillo, S., and Honigberg, S.M. Two-step method for constructing unmarked insertions, deletions and allele substitutions in the yeast genome. 2005. FEMS Microbiology Letters, 248: 31-6

In the spirit of the recent election, I decided to focus on a paper about change.

The authors developed a two-step process for removing, inserting or replacing regions in the yeast genome. They first remove the gene in question, replacing it with URA3. This is possible by flanking the URA3 with sequences homologous to the sequences flanking the gene in question. Due to recombination, URA3 is inserted in place of the gene. Selection for this occurs by growing cells on media lacking uracil ("Yes We Can...complete pyrimidine synthesis!"). The next step involves taking the replacement gene and flanking it with the same homologous sequences. Again, recombination replaces URA3 with this new insertion through recombination. Selection for this occurs by growing cells on media containing uracil and 5-FOA, a chemical that mimics uracil but is toxic. If a cell still has URA3 it will attempt to metabolize 5-FOA and kill itself. This method can be used to insert a new sequence (flankers originally touch), delete a sequence (URA3 replacement is just touching flankers), or replace a sequence. To insure replacement worked, primers from the replacing strand may contain a point mutation so that PCR would not amplify the inserted sequence and a gel-run would fail to show any DNA. Sequencing is always necessary to confirm due to the likelihood that any URA gene may mutate between functionality and pointless at any stage.

Friday, October 31, 2008

Imperfect Phenotypes: Transcripts and Peptides

Source: Ghaemmaghami, S., Huh, W., Bower, K., Howson, R.W., Belle, A., Dephoure, N., O’Shea, E.K., and Welssman, J.S. Global analysis of protein expression in yeast. 2003. Nature, 425:737-41

The authors created a library of strains with TAP-tagged to an ORF. They then could use a single antibody on each strain in mid-log phase to quantify the protein abundance. Comparing these values to those for complementary transcript abundance (via microarrays) and to codon bias scores, they found that there is a significant level of correlation between the measurements. Average transcript-to-protein ratio is fairly consistent near 4,000 proteins per transcript across the range of transcripts, but the protein variability for genes with the same transcript abundance was quite high. Codon bias had a similar good ratio, with high levels of variability at the same codon score. This variation was determined not to be measurement error, as a case of 206 essential proteins were retested on the TAP-tagged library in triplicate. The TAP library was able to detect a greater number of proteins than typical LC/MS methods, due to LC/MS high abundance bias.

So once again, we find that transcript differences aren't the best proxy for understanding cellular changes in response to perturbations. They may point in the right direction but a 10-fold change in one transcript and a 5-fold in another may result in the same peptide number difference. But are proteins any better, what with post-translational modifications and activations? While it is not as common place to measure proteome changes, I'm already thinking we may need to start looking into phosphorylome (?) changes to tie it all together and get a genuine comprehensive view.

Other notes:
-80% of proteome is expressed during normal growth conditions
-tandem affinity purification (TAP) tag is colmodulin binding peptide, TEV cleavage site, two IgG binding domains
-successful integrants for 98% of all ORFs in S. cerevisiae
-tagging does not hinder, TAP can be degraded
-detected 79% of essential, 83% of products corresponding to assigned gene names
=73% of all annotated ORFs
-very abundant mRNAs generally encode for abundant proteins
=some variation due to TAP tag, as subset that was retested had higher correlation than initially found
-similar, but lower, correlation between protein abundance and codon usage as measured by codon adaptation index (CAI)
=just noise at CAI < 0.2

Tuesday, October 28, 2008

Transcripts and Peptides: On-again, Off-again

Source: Foss, E.J., Radulovic, D., Shaffer, S.A., Ruderfer, D.M., Bedalov, A., Goodlett, D.R., and Kruglyak, L. Genetic basis of proteome variation in yeast. 2007. Nature Genetics, 39(11):1369-75

The authors used a mass spectrometry approach coupled with retention time shift software to measure the peptide abundances between the BY and RM strains and their segregants. Replicating a sample via quantitative western blot they found that they could reliably measure peptides as a quantitative trait, and that the levels showed inheritance patterns similar to transcript data from a previous study. A number of peptide levels differ significantly between the two parents and a linkage analysis found four major hotspots responsible for many of these linkages. Two of these hotspots were enriched for peptide synthesis, one being on LEU2, the other two were not enriched for any particular term. Almost all the peptides showed trans linkage. Only three of the four hotspots overlap to hotspots generated by transcript data of the 278 measurable peptides, and only rarely do these hotspots link to both a peptide and its corresponding transcripts. Interestingly, the 278 transcripts find the same hotspots as the total transcript database, suggesting a limited number of polymorphism control the whole proteome, thus the small sample size of proteins may still find the proteome hotspots.

•Other notes:
-weak correlation between transcript and peptide levels, suggesting post-transcriptional regulation acts as a buffer
-BY and RM differ at ~0.6% of genome
-if a representative sample this means 1/3 of all proteins differ between BY and RM
-221 peptides for 278 proteins
-heritability average was 62%
-expanded to 0.17 to get 85 proteins, 109 linkages
=only 7% of peptides map to a marker within 20kb of their ORF
-average correlation between protein and transcript abundance is 0.186
-3 of 4 hotspots overlap
=shared hotspots do not always overlap the same gene’s transcript and protein
=same polymorphism causes changes at different stages for different genes

Monday, October 20, 2008

The Age of Sequencing I: The Sanger Alliance

Source: Shendure and Ji (2008). Next-generation DNA sequencing. Nature Biotech 26:1135-1145.

There is no doubt that Sanger biochemistry (i.e. 'cycle sequencing') has set the fundamentals of sequencing; both in high-throughput clone based methods (e.g. shotgun) and single PCR product targeted sequenicng. In each cycle, labeled ddNTP molecules mark a single nucleotide at its end and through high-resolution electrophoretic separation, the final sequence is read. Each reaction can read about ~1000 bp, with the max accuracy of 99.999% and the cost of $0.5 per kb in shotgun.

This platform has been used for the emergence of the second generation sequencing pipelines: 454, Solexa, SOLiD, Polonator and HeliScope. These methods, despite many differences in their methodologies, follow the same 'cycling' logic followed by an optical signal of some sort. Library preparation, followed by adapter ligation and amplification is the first step of all these methods. Amplification is done either through in situ polonies, emulsion PCR or bridge PCR. The amplification step, however, should ensure the spacial clustering of each clone.

The bottom line is, these methods bypass the steps that are required in classic Sanger method and also use array-based approaches to enforce parallelism.

Saturday, October 11, 2008

Discovering Species-level Functions in Complex Microbial Communities

Source: Kalyuzhnaya et al. (2008). High-resolution metagenomics targets specific functional types in complex microbial communities. Nature Biotrech 26 (9): 1029-1034.

Environmental genomics (metagenomics) has become a hot topic in molecular biology; however, it is costly and highly dependent on the number of reference genomes available. Thus, studying highly complex communities like those of soil or lake sediments are not currently feasible. In this study, the authors use a labeling method to target the species that are specific for a given function. In this case, they have studied methylotrophy and through the usage of labled methane, methanol, methylamine, formaldehyde and formate, they have focused on the species directly assimilating these substrates through extracting the labeled fraction of the genomic DNA extracted from the community (using isopycnic centrifugation).

The 16S rRNA analysis shows that the samples enriched for methylotrophy are way less complex than the initial sample largely including the bona fide methylotrophs: Methylobacter tundripaludum, Methylomonas sp., Methylotenera mobilis, Methyloversatilis universalis, Ralstonia eutropha. It should be noted that some of the enriched species may not be methylotrophs but rather secondary links in the food chain (e.g. using 13C-CO2 produced by methylotrophs).

The authors demonstrate the utility of their approach through identifying a novel methylotroph and reconstructiong its genome and metabolic network.