The transcriptome is the set of all RNA molecules (transcripts) in a cell or a population of cells. It includes all of the functional RNA molecules and all other transcripts that may arise by spurious transcription or transcription of non-functional regions such as pseudogenes or virus fragments. A major goal of modern molecular biology is to determine which transcripts are functional and which ones are junk RNA. The term transcriptome is a portmanteau of the words transcript and genome; it is associated with the process of transcript production during the biological process of transcription. The functional part of the transcriptome is dynamic — it changes with cell type, developmental stage, environment, and stimuli — and therefore represents the active gene expression state rather than the static DNA sequence (genome). Eukaryotic transcriptomes tend to be more complex than bacterial transcriptomes and the transcriptomes of multicellular eukaryotes are even more complex than those of unicellular eukaryotes.
Etymology and history The word transcriptome is a portmanteau of the words transcript and genome. It appeared along with other neologisms formed using the suffixes -ome and -omics to denote all studies conducted on a genome-wide scale in the fields of life sciences and technology. As such, transcriptome and transcriptomics were one of the first words to emerge along with genome and proteome. The first study to present a case of a collection of a cDNA library for silk moth mRNA was published in 1979. The first seminal study to mention and investigate the transcriptome of an organism was published in 1997 and it described 60,633 transcripts expressed in S. cerevisiae using serial analysis of gene expression (SAGE). With the rise of high-throughput technologies and bioinformatics and the subsequent increased computational power, it became increasingly efficient and easy to characterize and analyze enormous amount of data. Attempts to characterize the transcriptome became more prominent with the advent of automated DNA sequencing during the 1980s. During the 1990s, expressed sequence tag sequencing was used to identify genes and their fragments. This was followed by techniques such as serial analysis of gene expression (SAGE), cap analysis of gene expression (CAGE), and massively parallel signature sequencing (MPSS).
Transcription
The transcriptome encompasses all the ribonucleic acid (RNA) transcripts present in a given organism or experimental sample. The functional component of the transcriptome includes RNAs that carry genetic information that is responsible for the process of converting DNA into an organism's phenotype. A gene gives rise to a single-stranded RNA molecule through a molecular process known as transcription; this RNA is complementary to the strand of DNA it originated from. The enzyme RNA polymerase attaches to the template DNA strand and catalyzes the addition of ribonucleotides to the 3' end of the growing sequence of the RNA transcript. In order to initiate its function, RNA polymerase needs to recognize a promoter sequence, located near the transcription start site that defines the beginning of the gene. This process is usually mediated and regulated by transcription factors. Transcription ends at a terminator site that defines the other end of the gene. The terminator site is often identified by termination sequences.
Types of RNA transcripts Almost all functional transcripts are derived from known genes. The only exceptions are a small number of transcripts that might play a direct role in regulating gene expression near the prompters of known genes. (See Enhancer RNA.) Genes occupy most of prokaryotic genomes so most of their genomes are transcribed. Many eukaryotic genomes are very large and known genes may take up only a fraction of the genome. In mammals, for example, known genes only account for 40-50% of the genome. Nevertheless, identified transcripts often map to a much larger fraction of the genome suggesting that the transcriptome contains spurious transcripts that do not come from genes. Some of these transcripts are known to be non-functional because they map to transcribed pseudogenes or degenerative transposons and viruses. Others map to unidentified regions of the genome that may be junk DNA. Spurious transcription is very common in eukaryotes, especially those with large genomes that might contain a lot of junk DNA. Some scientists claim that if a transcript has not been assigned to a known gene then the default assumption must be that it is junk RNA until it has been shown to be functional. This would mean that much of the transcriptome in species with large genomes is probably junk RNA. (See Non-coding RNA) The transcriptome includes the transcripts of protein-coding genes (mRNA plus introns) as well as the transcripts of non-coding genes (functional RNAs plus introns).
Ribosomal RNA/rRNA: Usually the most abundant RNA in the transcriptome. Long non-coding RNA/lncRNA: Non-coding RNA transcripts that are more than 200 nucleotides long. Members of this group comprise the largest fraction of the non-coding transcriptome other than introns. It is not known how many of these transcripts are functional and how many are junk RNA. transfer RNA/tRNA micro RNA/miRNA: 19-24 nucleotides (nt) long. Micro RNAs up- or downregulate expression levels of mRNAs by the process of RNA interference at the post-transcriptional level. small interfering RNA/siRNA: 20-24 nt small nucleolar RNA/snoRNA Piwi-interacting RNA/piRNA: 24-31 nt. They interact with Piwi proteins of the Argonaute family and have a function in targeting and cleaving transposons. enhancer RNA/eRNA:
… excerpt ends here. Continue reading the full article.


