Showing posts with label smallRNA. Show all posts
Showing posts with label smallRNA. Show all posts

Thursday, May 07, 2015

A clarity on the Illumina TruSeq Small RNA prep kit manual

In the TruSeq® Small RNALibrary Prep Guide, below the Figure 1, there is a sentence: "The RNA 3' adapter is modified to target microRNAs and other small RNAs that have a 3' hydroxyl group resulting from enzymatic cleavage by Dicer or other RNA processing enzymes." It's right, but could be very misleading if you are not clear of the diverse picture of transcriptome (scroll down for more detail). I want to emphasize that the 3' hydroxyl group (and the 5'-phosphate group) is NOT specific to microRNAs or any small RNAs. And it doesn't necessarily result from enzymatic cleavage by Dicer. Sonic fragmentation can also break the full length mRNA (with 5'-cap and 3'-polyA) into truncated RNA pieces with 5'-phosphate and 3' hydroxyl free ends. I just called Illumina to confirm that the 3' and 5' ligation steps don't guarantee the selection of miRNAs (but rather any RNAs with 5'-phosphate and 3' hydroxyl ends, if more accurately). The last step of gel purification is the key to select (or enrich, if more accurately) miRNAs.


OK. Here is what I learned from my colleagues about the different RNA species in the trancriptome:

There are 4 species in the transcriptome, where the later 3 are intermediates of transcription (or half product of degradation).
  • me7Gppp-------------------------3' (1) 
  •                 p------------- 3' (2) 
  •                 OH-------------3' (3) 
  •        ppp---------------------3' (4) 
Only group(2) will ligate to 5’adaptor. The 3' end can also have different format, at least two:
  • 5' ----------------- AAAAAA (1) 
  • 5' ---------- OH (2) 
Also note there are two enzymes used to repair the 5' ends: CIP and TAP. CIP (Calf Intestinal alkaline phosphatase) can remove the 5’ phosphate group of DNA strand. The TAP (Tobacco acid pyrophosphatase) is to remove the 5' cap structure (or 5'-5' triphosphate linkage) and leave a mono-phosphate at the 5' end. So, applying first CIP and then TAP will convert the above (2) and (4) to (3), then convert (1) to (2). That's one way to do capture the 5' cap structure, same purpose as CAGE (but CAGE use the type IIs restriction enzyme MmeI and type III restriction enzyme EcoP15I)

Friday, July 26, 2013

smallRNA in neurodegenerative diseases

Basic question: how does the small RNA function in neurodegenerative disease brain?

It's been shown that small RNAs (miRNA, piRNA, circular RNA) present in the brain. Recent researches even show that piRNA and circular RNA also expressed highly in brain tissues.

Rajasethupathy P, Antonov I, Sheridan R, Frey S, Sander C, Tuschl T, Kandel ER.
Cell. 2012 Apr 27;149(3):693-707. doi: 10.1016/j.cell.2012.02.057.



The authors found several piRNA clusters highly expressed in Aplysia CNS, with the biogenesis pattern of piRNA in germline, including 5' U, piwi protein loaded, clustered, nuclear localized, sensitive to serotonin. They found that "the Piwi/piRNA complex facilitates serotonin-dependent methylation of a conserved CpG island in the promoter of CREB2, the major inhibitory constraint of memory in Aplysia, leading to enhanced long-term synaptic facilitation". They speculated that the piwi/piRNA complex methylates the promoter of target gene (CREB2 in this case) by first accessing its nascent transcript. Not sure if this also happened to other neuron-functioning genes, such as HES4 and HTT for HD.

Hansen TB, Jensen TI, Clausen BH, Bramsen JB, Finsen B, Damgaard CK, Kjems J.
Nature. 2013 Mar 21;495(7441):384-8. doi: 10.1038/nature11993. Epub 2013 Feb 27.


Memczak S, Jens M, Elefsinioti A, Torti F, Krueger J, Rybak A, Maier L, Mackowiak SD, Gregersen LH, Munschauer M, Loewer A, Ziebold U, Landthaler M, Kocks C, le Noble F, Rajewsky N.
Nature. 2013 Mar 21;495(7441):333-8. doi: 10.1038/nature11928. Epub 2013 Feb 27.


In these two new studies, the primary circular RNA of interest is an antisense transcript to CDR1 termed CDR1as (CDR1 antisense, used herein) or ciRS-7 (circular RNA sponge for miR-7). This RNA was validated by rigorous bioinformatic and molecular studies. Seventy-four copies of a miR-7 binding site are present in this circular RNA owing to the presence of a tandem repeat in the transcribed sequence of CDR1as. These binding sites are conserved across eutherian (placental) mammals, whereas the intervening sequences generally are not.
(Cite: http://www.nature.com/mt/journal/v21/n6/full/mt2013101a.html)

It's interesting that the circular RNA discovered this year also anti-sense to the target/host gene. Would that be a general mechanism that neurodegenerative disease are controlled by the anti-sense RNA?


Another question: why the repeated peptide/trinucleotide insertion or expansion is a common feature for most neurodegenerative disease?

Here is the table for the at least 20 inherit neuro-diseases and the associated repeat insertion:

DiseaseMain clinical featuresCausal repeat (gene)Repeat locationMechanism or categoryComments
DM1Muscle weakness, myotonia, cardiac-endocrine-GI disease, MRCTG (DM1, also known as DMPK)3′ UTRRNA GOFA very common form of muscular dystrophy
DM2Muscle weakness, myotonia, cardiac-endocrine-GI diseaseCTG (ZNF9, also known as CNBP)IntronRNA GOFA striking phenocopy of DM1
DRPLASeizures, choreoathetosis, ataxia, cognitive declineCAG (ATN1)Coding regionPolyglutamine GOFVery rare, most patients are in Japan
FMR1MR, facial dysmorphism, autismCGG (FMR1)5′ UTRHypermethylation of promoter, LOFMost common inherited MR
FMR2MR, hyperactivityGCC (FMR2)5′ UTRLOFNeeds to be ruled out in X-linked MR
FRDAAtaxia, sensory loss, weakness, diabetes mellitus, cardiomyopathyGAA (FXN)IntronLOF, phenocopy of mitochondrial diseaseMost common inherited ataxia in Caucasian ethnicity
FXTASAtaxia, intention tremor, parkinsonismCGG (FMR1)5′ UTRRNA GOFPremutation carriers only
HDChorea, dystonia, cognitive decline, psychiatric diseaseCAG (HTT)Coding regionPolyglutamine GOFOne of the most common inherited diseases in humans
HDL2Chorea, dystonia, cognitive declineCTG (JPH3)3′ UTR, coding regionRNA GOF, poly-amino acid GOF and/or LOF?A striking phenocopy of HD
Myoclonic epilepsy of Unverricht and LundborgPhotosensitive myoclonus, tonic–clonic seizures, cerebellar degenerationCCCCGCCCCGCG (CSTB)PromoterLOFRare autosomal recessive disorder found in Finland and N. Africa
OPMDEyelid weakness, dysphagia, proximal limb weaknessGCG (PABPN1)Coding regionPolyalanine GOFModest expansion causes disease
SBMAProximal limb weakness, lower motor neuron diseaseCAG (AR)Coding regionPolyglutamine GOFPhenotype includes LOF androgen insensitivity
SCA1Ataxia, dysarthria, spasticity, ophthalmoplegiaCAG (ATXN1)Coding regionPolyglutamine GOFAccounts for 6% of all dominant ataxia
SCA2Ataxia, slow eye movement, hyporeflexia, motor disease, occasional parkinsonismCAG (ATXN2)Coding regionPolyglutamine GOFATXN2 protein may not reside in the nucleus
SCA3Ataxia, dystonia, lower motor neuron diseaseCAG (ATXN3)Coding regionPolyglutamine GOFMost common dominant ataxia
SCA6Ataxia, dysarthria, sensory loss, occasionally episodicCAG (CACNA1A)Coding regionPolyglutamine GOFCausal gene encodes a subunit of a P/Q-type Ca2+ channel
SCA7Ataxia, dysarthria, cone-rod dystrophy retinal diseaseCAG (ATXN7)Coding regionPolyglutamine GOFClinically distinct as patients have retinal disease
SCA8Ataxia, dysarthria, nystagmus, spasticityCTG/CAG (ATXN8)Untranslated RNA, coding regionRNA GOF and polyglutamine GOFMany cases of reduced penetrance
SCA10Ataxia, dysarthria, seizures, dysphagiaATTCT (ATXN10)IntronRNA GOF?Huge repeats; only Mexican ancestry?
SCA12Tremor, ataxia, spasticity, dementiaCAG (PPP2R2B)Promoter, 5′ UTR?UnknownCausal gene encodes a phosphatase
SCA17Ataxia, dementia, chorea, seizures, dystoniaCAG (TBP)Coding regionPolyglutamine GOFCausal gene encodes a common transcription factor (TBP)
Syndromic/non-syndromic X-linked mental retardationMR alone, with seizures or with dysarthria and dystoniaGCG (ARX)Coding regionProbably LOFAssociated with West syndrome or Partington syndrome
AR, androgen receptor; ARX, aristaless-related homeobox; ATN1, atrophin 1; ATXN, ataxin; CACNA1A, voltage-dependent P/Q-type calcium channel subunit α-1A; CSTB, cystatin B; DM, myotonic dystrophy; DMPK, DRPLA, dentatorubral-pallidoluysian atrophy; FMR1, fragile X mental retardation syndrome; FMR2, fragile X E mental retardation; FRDA, Friedreich's ataxia; FXN, frataxin; FXTAS, fragile X tremor ataxia syndrome; GI, gastrointestinal; GOF, gain of function; HD, Huntington's disease; HDL2, Huntington's disease-like 2; HTT, huntingtin; JPH3, junctophilin 3; LOF, loss of function; MR, mental retardation; OPMD, oculopharyngeal muscular dystrophy; PABPN1, poly(A)-binding protein, nuclear 1; PPP2R2B, protein phosphatase 2 regulatory subunit B, β isoform; SBMA, spinal and bulbar muscular atrophy; SCA, spinocerebellar ataxia; TBP, TATA box-binding protein; ZNF9, zinc finger 9.
Ref: http://www.nature.com/nrg/journal/v11/n4/fig_tab/nrg2748_T1.html

Not only genes in the list, other genes like the CDR1 (Cerebellar Degeneration-Related) gene mentioned above also contains several hexapeptide repeats.

 CDS             387..1175
                     /gene="CDR1"
                     /gene_synonym="CDR; CDR34; CDR62A"
                     /note="Derived by automated computational analysis using
                     gene prediction method: BestRefseq."
                     /codon_start=1
                     /product="cerebellar degeneration-related antigen 1"
                     /protein_id="NP_004056.2"
                     /db_xref="GI:44889483"
                     /db_xref="CCDS:CCDS14670.1"
                     /db_xref="GeneID:1038"
                     /db_xref="HGNC:1798"
                     /db_xref="HPRD:02361"
                     /db_xref="MIM:302650"
                     /translation="MAWLEDVDFLEDVPLLEDIPLLEDVPLLEDVPLLEDTSRLEDIN
                     LMEDMALLEDVDLLEDTDFLEDLDFSEAMDLREDKDFLEDMDSLEDMALLEDVDLLED
                     TDFLEDPDFLEAIDLREDKDFLEDMDSLEDLEAIGRCGFSGRHGFFGRRRFSGRPKLS
                     GRLGLLGRRGFSGRLGGYWKTWIFWKTWIFWKTWIFRKTYIGWKTWIFSGRCGLTGRP
                     GFGGRRRFFWKTLTDWKTWISFWKTLIDWKTWISFWKTLIDWKI"

Monday, March 18, 2013

Question: why RNA repeats defined in repeatMasker are not found in annotated smallRNA?

I was fetching small RNA like tRNA/rRNA/snRNA/snoRNA from the latest Ensembl, and then I noticed that there are a class of repeat in RepBase called 'RNA repeat', which contains various RNA, tRNA, rRNA, snRNA, scRNA, and srpRNA. I expected they are annotated, but actually when I intersect between them and I found a large number of RNA repeats are not annotated. Here is it:


$ grep rRNA $repeats_from_RepBase | intersectBed -a stdin -b $snRNAsnoRNAtRNArRNA_from_Ensembl -s -v | sort -u | wc -l
1557

Here is an rRNA example:


and a tRNA example:
Both are not annotated as a rRNA/tRNA gene, but as a RNA repeat. 

So, what's the definition of RNA repeat for RepBase? pseudogene?

Thursday, November 29, 2012

measuring expression of smallRNA genes

As the early note mentioned, it's not a good idea to use bias-correction based method like cufflinks to calculate the RPKM/FPKM values for small RNA (simply because it's not designed for that purpose).

Since the smallRNA is generally shorter, the usual RNA-seq protocol does not capture them properly (only the 'tail' of the distribution of fragment length). A good solution is to use smallRNA RNA-seq protocol. I meant using short reads length and without size selection.

Based on that, we may be able to use tools like cufflinks/DESeq to calculate the RPKM/FPKM. But again, I am not sure cufflinks will still apply the bias-correction to the small RNA genes when the reads length is short enough. As this post discussed, "The Best way to normalize miRNA HTS data", methods can be:

- RPM ( Read Per Million )
- RPKM (Reads per kilobase per million mapped) : Mortazi et al, Nat. Methods, 2008 (http://www.ncbi.nlm.nih.gov/pubmed/18516045)
- Trimmed mean of M-values : Robinson, Oshlack, Genome Biology 2010 (http://genomebiology.com/2010/11/3/R25)
- upper-quantile : http://www.ncbi.nlm.nih.gov/pubmed/20167110

Simon (one of the authors of DESeq) suggested to use TMM or DESeq (which are reads count-based methods) for differential expression analysis, at least for small RNA genes.

As the DEseq document described, to get reads count you need alignment (SAM/BAM) and the annotation file as input for some tools, like
  • the function summarizeOverlaps in the BioConductor GenomicRanges package,
  • the geneCounts function in the Bioconductor easyRNASeq package,
  • the htseq-count script distributed with the HTSeq Python package (Note:  htseq-count produces one count file for each sample. DESeq offers the function newCountDataSetFromHTSeqCount to get an analysis started from these files)  
  • or the bedtools multicov -count -bams sample1.bam sample2.bam sample3.bam -bed genes.bed
Note the naming issue in Refseq or UCSC that same locus has been assigned with multiple gene names (e.g. SMN1/2), and htseq-count requires unique gene_id in the annotation. You might want to make sure of this before using htseq-count. One solution to this is to use Ensembl GTF.

Also, be aware of the difference between bedtools multicov and htseq-count, esp. when you use multi-mappers. Read details here: http://seqanswers.com/forums/showthread.php?t=24903