COLLEGE BIOLOGY • GENE EXPRESSION & REGULATION

Transcription and RNA Processing

How cells decode DNA into functional RNA molecules through a precisely orchestrated molecular machinery.

Historical Context & Motivation

The question of how genetic information stored in DNA is converted into functional gene products has driven molecular biology since the discovery of the double helix. By the mid-twentieth century, researchers understood that DNA carried hereditary information, but the mechanism by which that information was read and expressed remained a central mystery. The concept of an intermediate molecule — messenger RNA (mRNA) — that carries genetic instructions from the nucleus to the ribosome became one of the most transformative ideas in modern biology. Understanding transcription and RNA processing is foundational to comprehending gene expression, cellular differentiation, and the molecular basis of disease.

1956
The Central Dogma
Francis Crick articulated the central dogma of molecular biology, proposing that information flows from DNA → RNA → protein, establishing the conceptual framework for transcription.
1961
Discovery of mRNA
François Jacob and Jacques Monod proposed the existence of a short-lived RNA intermediate — messenger RNA — that carries genetic information from DNA to ribosomes. Sydney Brenner and colleagues provided experimental evidence supporting this model.
1969
RNA Polymerase Isolated
Multiple laboratories isolated RNA polymerase enzymes from bacteria, revealing the core enzymatic machinery responsible for synthesizing RNA from a DNA template.
1977
Discovery of Introns and RNA Splicing
Phillip Sharp and Richard Roberts independently discovered that eukaryotic genes contain introns — non-coding sequences that must be removed from the primary transcript before translation. This finding earned them the 1993 Nobel Prize.
2006
Structural Basis of Transcription
Roger Kornberg received the Nobel Prize in Chemistry for elucidating the atomic-resolution crystal structure of eukaryotic RNA Polymerase II, revealing the precise molecular architecture of the transcription machinery.

These discoveries collectively revealed that gene expression is not a simple one-step readout of DNA. Instead, cells employ a sophisticated multi-stage process in which a primary RNA transcript (pre-mRNA) undergoes extensive processing — including capping, polyadenylation, and splicing — before it is competent for translation. The central question this lesson addresses is: how does the cell accurately copy genetic information from DNA into RNA, and what processing steps convert that raw transcript into a mature, functional molecule?

Core Principles of Transcription

Transcription is the enzymatic synthesis of an RNA molecule from a DNA template. Although the fundamental chemistry of phosphodiester bond formation is conserved between prokaryotes and eukaryotes, the regulatory complexity and post-transcriptional processing differ substantially. Several core principles govern the transcription process across all domains of life.

1

Template-Directed Synthesis

RNA polymerase reads the template strand (also called the antisense strand) of DNA in the 3′ → 5′ direction, synthesizing the RNA transcript in the 5′ → 3′ direction. The resulting RNA is complementary and antiparallel to the template strand, and identical in sequence to the coding (sense) strand, except with uracil (U) replacing thymine (T).
2

Promoter Recognition

Transcription initiates at specific DNA sequences called promoters. In prokaryotes, the sigma (σ) factor of RNA polymerase holoenzyme recognizes the −10 and −35 elements. In eukaryotes, general transcription factors (GTFs) assemble at the TATA box and other core promoter elements to recruit RNA Polymerase II.
3

NTP Substrates

RNA polymerase uses nucleoside triphosphates (NTPs) — ATP, UTP, GTP, and CTP — as substrates. The enzyme catalyzes the formation of a phosphodiester bond between the 3′-OH of the growing chain and the α-phosphate of the incoming NTP, releasing pyrophosphate (PPi). Subsequent hydrolysis of PPi by pyrophosphatase drives the reaction forward.
4

Three Phases of Transcription

Transcription proceeds through three distinct phases: initiation (promoter binding and open complex formation), elongation (processive RNA synthesis), and termination (release of the transcript and dissociation of the polymerase).
5

Co-transcriptional Processing

In eukaryotes, RNA processing events — including 5′ capping, splicing, and 3′ polyadenylation — occur while the transcript is still being synthesized, coordinated through the C-terminal domain (CTD) of RNA Pol II.
KEY TAKEAWAY
Think of transcription as a molecular photocopier: the DNA double helix is the master document stored in a secure vault (the nucleus), and the mRNA is a working photocopy that can be carried out to the factory floor (the ribosome) for protein assembly. Just as a photocopier reads one page at a time and produces a single-sided copy, RNA polymerase reads one strand of DNA and produces a single-stranded RNA. In eukaryotes, the photocopy undergoes editing (splicing), a cover page (5′ cap), and a binding edge (poly-A tail) before it is considered ready for use.

Visualizing the Transcription Process

The following diagram illustrates the three major phases of transcription — initiation, elongation, and termination — as they occur at a eukaryotic gene. Note how the transcription bubble moves along the DNA, unwinding the double helix ahead of the polymerase and rewinding it behind. The RNA transcript emerges from the polymerase as a single-stranded molecule that will subsequently undergo processing.

The three phases of eukaryotic transcription. During initiation (left), RNA Pol II and general transcription factors assemble at the promoter (TATA box). During elongation (center), the polymerase moves along the template, unwinding DNA at the transcription bubble and synthesizing nascent RNA (pink). During termination (right), the polymerase encounters the poly(A) signal, the transcript is cleaved and released, and the polymerase dissociates from the DNA.

Several important details emerge from this diagram. First, notice that the transcription bubble represents a locally unwound region of approximately 12–14 base pairs where the template strand is exposed to the polymerase's active site. The RNA–DNA hybrid within the bubble spans about 8–9 base pairs before the nascent RNA peels away as a single-stranded molecule. Second, the directionality of synthesis is invariant: RNA polymerase always adds nucleotides to the 3′-OH end of the growing chain, reading the template in the 3′ → 5′ direction. Third, in eukaryotes, termination is coupled to the cleavage and polyadenylation machinery, which recognizes the consensus sequence AAUAAA in the nascent transcript.

Mechanistic Details of Transcription

Initiation: Assembling the Pre-Initiation Complex

In eukaryotes, transcription initiation at protein-coding genes requires the ordered assembly of a pre-initiation complex (PIC) at the core promoter. The process begins when TFIID — specifically its TBP (TATA-binding protein) subunit — recognizes and binds the TATA box, typically located approximately 25–30 base pairs upstream of the transcription start site (+1). TBP binding induces a dramatic bend in the DNA, facilitating subsequent recruitment of TFIIA, TFIIB, TFIIF (which escorts RNA Pol II to the promoter), and finally TFIIE and TFIIH. The helicase activity of TFIIH unwinds approximately 11–15 bp of DNA around the start site, converting the closed complex to an open complex. The kinase activity of TFIIH then phosphorylates Serine 5 of the CTD heptapeptide repeats (Tyr-Ser-Pro-Thr-Ser-Pro-Ser) on RNA Pol II's largest subunit, triggering promoter clearance.

Elongation: Processive RNA Synthesis

Once the polymerase clears the promoter, it enters the elongation phase. The enzyme moves along the template strand at a rate of approximately 20–50 nucleotides per second in eukaryotes (compared to ~40–80 nt/s in prokaryotes). During elongation, the polymerase maintains the transcription bubble, catalyzing the nucleophilic attack of the 3′-OH of the nascent RNA chain on the α-phosphate of the incoming NTP.

PHOSPHODIESTER BOND FORMATION
(RNA)ₙ + NTP → (RNA)ₙ₊₁ + PPᵢ
Where (RNA)n is the growing RNA chain of length n, NTP is the incoming nucleoside triphosphate, and PPi is inorganic pyrophosphate. The reaction is thermodynamically driven forward by the subsequent hydrolysis of PPi → 2 Pi (ΔG° ≈ −33 kJ/mol).

Elongation factors such as P-TEFb (positive transcription elongation factor b) phosphorylate Serine 2 of the CTD, promoting productive elongation and the recruitment of RNA processing factors. The polymerase also possesses intrinsic proofreading capability: it can reverse-track (backtrack) and cleave misincorporated nucleotides using its endonuclease activity, enhanced by the elongation factor TFIIS.

Termination: Releasing the Transcript

Eukaryotic transcription termination for RNA Pol II genes is mechanistically linked to 3′ end processing. Two models account for termination: the allosteric (anti-terminator) model proposes that passage through the poly(A) signal sequence (AAUAAA) triggers conformational changes in the elongation complex, destabilizing it. The torpedo model proposes that after cleavage at the poly(A) site, a 5′→3′ exonuclease (Rat1 in yeast, Xrn2 in mammals) degrades the downstream RNA still associated with the polymerase, eventually catching up and displacing it. Current evidence suggests elements of both models operate in vivo.

🔬 Prokaryotic vs. Eukaryotic Termination
In prokaryotes, termination occurs via two mechanisms: intrinsic (rho-independent) termination relies on a GC-rich hairpin followed by a poly-U tract in the nascent RNA. Rho-dependent termination involves the Rho helicase, which translocates along the transcript and unwinds the RNA–DNA hybrid at pause sites.

Eukaryotic RNA Processing

Eukaryotic pre-mRNA undergoes three major processing events before it can be exported from the nucleus as mature mRNA: 5′ capping, 3′ polyadenylation, and intron splicing. These modifications protect the transcript from degradation, facilitate nuclear export, and are critical for efficient translation. All three processes are coordinated through interactions with the phosphorylated CTD of RNA Pol II.

The four steps of eukaryotic mRNA processing. The primary transcript (pre-mRNA) contains exons (teal) interspersed with introns (purple, dashed). Processing adds a 5′ m⁷G cap (gold circle), a poly(A) tail (green), and removes introns via splicing, yielding a contiguous mature mRNA ready for export.

5′ Capping

The 5′ cap is added co-transcriptionally after the first 20–30 nucleotides have been synthesized. A 7-methylguanosine (m⁷G) is linked to the first nucleotide of the transcript via an unusual 5′–5′ triphosphate bridge. This cap structure is critical: it protects the mRNA from 5′ exonuclease degradation, is recognized by the nuclear cap-binding complex (CBC) for export, and is later bound by eIF4E to initiate cap-dependent translation.

3′ Polyadenylation

The 3′ end of most eukaryotic mRNAs is generated not by transcription termination itself, but by endonucleolytic cleavage approximately 10–30 nucleotides downstream of the polyadenylation signal (AAUAAA). The cleavage is carried out by CPSF (cleavage and polyadenylation specificity factor) and CstF (cleavage stimulation factor). Following cleavage, poly(A) polymerase (PAP) adds approximately 200 adenine residues to the 3′ end in a template-independent reaction. The poly(A) tail, bound by poly(A)-binding protein (PABP), enhances mRNA stability, promotes translation, and facilitates nuclear export.

RNA Splicing

Most eukaryotic protein-coding genes are interrupted by introns — non-coding sequences that must be precisely excised from the pre-mRNA, with the flanking exons ligated together. Splicing is catalyzed by the spliceosome, a massive ribonucleoprotein complex composed of five small nuclear RNAs (snRNAs: U1, U2, U4, U5, U6) and over 100 associated proteins. The spliceosome recognizes three conserved sequence elements within each intron: the 5′ splice site (GU), the branch point adenosine (typically 18–40 nt upstream of the 3′ splice site), and the 3′ splice site (AG). Splicing proceeds through two sequential transesterification reactions. In the first step, the 2′-OH of the branch point adenosine attacks the phosphodiester bond at the 5′ splice site, generating a lariat intermediate. In the second step, the 3′-OH of the freed upstream exon attacks the 3′ splice site, joining the two exons and releasing the intron as a lariat.

🧬 Alternative Splicing
A single gene can produce multiple mRNA variants — and therefore multiple protein isoforms — through alternative splicing. By selectively including or excluding certain exons, or by using alternative 5′ or 3′ splice sites, the cell dramatically expands its proteomic diversity. It is estimated that over 95% of human multi-exon genes undergo alternative splicing. The Drosophila Dscam gene can potentially produce over 38,000 distinct mRNA isoforms through combinatorial exon selection.

Worked Example: From Gene to Mature mRNA

Consider a eukaryotic gene with the following structure: a promoter with a TATA box at position −30, a transcription start site at +1, three exons (Exon 1: 150 nt, Exon 2: 200 nt, Exon 3: 300 nt), and two introns (Intron 1: 1,200 nt, Intron 2: 800 nt). We will trace the path from gene to mature mRNA, calculating the sizes of the primary transcript and mature mRNA.

Determining the Structure and Size of a Mature mRNA
1
Step 1 — Determine the primary transcript lengthThe primary transcript (pre-mRNA) includes all exons and introns from the transcription start site through the poly(A) signal. Assuming the poly(A) cleavage site is at the end of Exon 3's sequence, the total transcribed length is: Exon 1 + Intron 1 + Exon 2 + Intron 2 + Exon 3 = 150 + 1,200 + 200 + 800 + 300 = 2,650 nt.
Pre-mRNA length ≈ 2,650 nucleotides
2
Step 2 — Apply 5′ cappingA 7-methylguanosine (m⁷G) cap is added to the 5′ end of the pre-mRNA via a 5′–5′ triphosphate linkage. This occurs co-transcriptionally after synthesis of the first ~25 nt. The cap does not significantly change the nucleotide count but adds the modified guanosine at the 5′ terminus.
5′ end: m⁷GpppN...
3
Step 3 — Apply 3′ polyadenylationAfter the pre-mRNA is cleaved downstream of the AAUAAA signal, poly(A) polymerase (PAP) adds approximately 200 adenine residues to the 3′ end. This adds ~200 nt to the total length.
3′ end gains ~200 A residues
4
Step 4 — Perform intron splicingThe spliceosome removes both introns (Intron 1 = 1,200 nt, Intron 2 = 800 nt), excising a total of 2,000 nt and ligating the exons together: Exon 1–Exon 2–Exon 3.
Total intronic sequence removed: 2,000 nt
5
Step 5 — Calculate mature mRNA lengthThe mature mRNA consists of the joined exons plus the poly(A) tail (and the 5′ cap, which is a single modified nucleotide). Total coding/UTR region = 150 + 200 + 300 = 650 nt from exons + ~200 nt poly(A) tail = approximately 850 nt (plus the cap structure).
Mature mRNA ≈ ~850 nucleotides (m⁷G cap + 650 nt exonic sequence + ~200 nt poly(A) tail)
6
Step 6 — Compare pre-mRNA to mature mRNAThe pre-mRNA was 2,650 nt. The mature mRNA (exons + poly-A) is ~850 nt. This means that approximately 75.5% of the primary transcript consisted of introns that were removed during splicing. This ratio is typical for mammalian genes, where introns often vastly exceed exons in length.
Introns represented ~75.5% of the pre-mRNA (2,000/2,650)

Comparing Prokaryotic and Eukaryotic Transcription

While the fundamental chemistry of transcription is conserved across all life, significant differences exist between prokaryotic and eukaryotic systems. These differences reflect the distinct cellular architectures and regulatory requirements of the two domains. Understanding these contrasts is essential for appreciating the added complexity of eukaryotic gene expression.

Key differences between prokaryotic and eukaryotic transcription
FeatureProkaryotesEukaryotes
RNA PolymeraseSingle RNA polymerase (core enzyme: α₂ββ′ω); σ factor for promoter recognitionThree RNA polymerases: Pol I (rRNA), Pol II (mRNA, snRNA), Pol III (tRNA, 5S rRNA)
Promoter Elements−10 (Pribnow box: TATAAT) and −35 (TTGACA) consensus sequencesTATA box (~−30), Inr, BRE, DPE, MTE; recognized by general transcription factors (GTFs)
mRNA ProcessingNo capping, no polyadenylation (usually), no splicing; mRNA is translated co-transcriptionally5′ capping (m⁷G), 3′ polyadenylation (~200 A), intron splicing by spliceosome
Transcription–Translation CouplingCoupled: ribosomes attach to mRNA while it is still being transcribedUncoupled: transcription in nucleus, translation in cytoplasm (separated by nuclear envelope)
Gene OrganizationPolycistronic mRNAs (operons); genes are typically intronlessMonocistronic mRNAs; genes contain introns (often >90% of gene length)
TerminationRho-independent (intrinsic hairpin + poly-U) or Rho-dependent (Rho helicase)Coupled to polyadenylation; torpedo model (Xrn2/Rat1) and allosteric mechanisms
Elongation Rate~40–80 nucleotides/second~20–50 nucleotides/second (through chromatin)
KEY TAKEAWAY
The fundamental difference between prokaryotic and eukaryotic transcription can be compared to two different publishing models. In the prokaryotic model, a manuscript is read aloud (transcribed) and immediately translated into action by a live audience (ribosomes) — a streamlined, rapid process akin to live television broadcast. In the eukaryotic model, the manuscript must first go through a multi-step editing and production process — proofreading (capping), adding a conclusion (poly-A), cutting unnecessary chapters (splicing) — before being shipped to a separate facility (cytoplasm) for final production. This added complexity allows for much greater regulation and diversity of output.

Connections to Advanced Topics

The basic framework of transcription and RNA processing presented in this lesson connects to several advanced areas of molecular biology and biomedicine. Understanding these connections provides context for how the core mechanisms are regulated, perturbed in disease, and exploited therapeutically.

Connections from transcription fundamentals to advanced molecular biology and medicine
Core ConceptAdvanced ExtensionClinical / Research Significance
Promoter recognition by GTFsEnhancers, silencers, and Mediator complexMutations in enhancer regions are linked to cancer and developmental disorders; super-enhancers drive oncogene expression
CTD phosphorylationEpigenetics and chromatin remodelingHistone modifications (H3K4me3 at promoters, H3K36me3 in gene bodies) are deposited co-transcriptionally and influence future gene expression
Alternative splicingSplicing regulatory networksMis-splicing causes ~15% of human genetic diseases (e.g., spinal muscular atrophy, certain β-thalassemias); antisense oligonucleotides (ASOs) like nusinersen correct splicing defects
mRNA stability (cap + poly-A)mRNA decay pathways and RNA therapeuticsmRNA vaccines (e.g., COVID-19) use modified nucleosides and optimized 5′ UTR/cap structures to enhance translation and evade innate immunity
Transcription elongationPromoter-proximal pausingMany eukaryotic genes have Pol II paused ~30–60 nt downstream of the TSS; release by P-TEFb is a major regulatory checkpoint, and is targeted by BET inhibitor drugs in cancer therapy

These advanced topics illustrate that transcription and RNA processing are not merely mechanical copying operations — they represent highly regulated, dynamic processes that serve as critical control points for gene expression. The emerging fields of epitranscriptomics (the study of chemical modifications to RNA, such as m⁶A methylation) and RNA therapeutics (including mRNA vaccines, antisense oligonucleotides, and CRISPR guide RNAs) all build directly upon the foundational principles covered in this lesson. As you advance through molecular biology, you will see these themes recur at every level of gene regulation.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain why the sequence of the mRNA transcript is identical to the coding (sense) strand of DNA rather than the template (antisense) strand. In your answer, clarify the relationship between complementary base pairing and strand directionality, and explain why uracil appears in the mRNA where thymine would appear in the coding strand.
PROBLEM 2BASIC CALCULATION
A eukaryotic gene has 5 exons (80 nt, 120 nt, 200 nt, 150 nt, and 100 nt) and 4 introns (2,000 nt, 1,500 nt, 3,000 nt, and 500 nt). After transcription and processing (including addition of a ~250-nt poly(A) tail), what is the approximate length of the mature mRNA? What percentage of the primary transcript consisted of intronic sequence?
PROBLEM 3INTERMEDIATE
A researcher treats cells with α-amanitin, a potent inhibitor of RNA Polymerase II, and then analyzes the types of RNA being synthesized. Which classes of RNA would continue to be produced? Which would be abolished? Explain why, referencing the division of labor among eukaryotic RNA polymerases.
PROBLEM 4APPLIED
A patient presents with a genetic disorder caused by a point mutation at the 5′ splice site of intron 3 of a four-exon gene. The normal 5′ splice site consensus is GU, but the mutation changes it to AU. Predict the most likely consequences of this mutation on mRNA structure and the resulting protein. Consider at least two possible outcomes.
PROBLEM 5CRITICAL THINKING
The discovery that the spliceosome uses RNA catalysis (with snRNAs forming the catalytic core) has led to the hypothesis that splicing is an evolutionary relic of the 'RNA world.' Evaluate this hypothesis: what evidence supports the idea that self-splicing introns preceded the spliceosome, and what alternative hypothesis has been proposed? How might the existence of alternative splicing complicate a purely 'introns-early' interpretation?

Lesson Summary

Transcription is the first step of gene expression, in which RNA polymerase synthesizes an RNA copy of a gene's template strand in the 5′ → 3′ direction. The process proceeds through three phases — initiation (promoter recognition and open complex formation), elongation (processive NTP addition), and termination (transcript release). In eukaryotes, initiation requires assembly of the pre-initiation complex (PIC) including general transcription factors (TFIIA–H) and RNA Pol II at the TATA box of the core promoter.

Eukaryotic pre-mRNA undergoes three essential processing steps: addition of a 5′ m⁷G cap (protects from degradation, promotes translation initiation), addition of a 3′ poly(A) tail (~200 adenines; enhances stability and export), and removal of introns by the spliceosome through two transesterification reactions that join exons into a contiguous coding sequence. Alternative splicing vastly expands proteomic diversity from a limited gene set. These processes distinguish eukaryotic from prokaryotic gene expression, where transcription and translation are coupled and mRNAs typically require no processing.

Varsity Tutors • College Biology • Transcription and RNA Processing