AP BIOLOGY • GENE EXPRESSION AND REGULATION

Transcription and RNA Processing

How cells convert DNA instructions into functional RNA molecules through precise enzymatic mechanisms and post-transcriptional modifications.

Historical Context & Motivation

The discovery that genetic information flows from DNA to RNA to protein stands as one of the most consequential insights in molecular biology. After Watson and Crick elucidated the double-helical structure of DNA in 1953, the scientific community faced an urgent question: how does the information encoded in a stable, nuclear molecule direct the synthesis of proteins in the cytoplasm? The answer required identifying an intermediary molecule — messenger RNA (mRNA) — and the enzyme responsible for copying DNA into RNA, known as RNA polymerase. These discoveries launched decades of research into the mechanisms of transcription and the elaborate processing steps that convert a primary transcript into a mature, functional RNA.

1956
Discovery of RNA Polymerase Activity
Several labs, including those of Severo Ochoa and Arthur Kornberg, identified enzymatic activities capable of synthesizing RNA from nucleotide precursors, establishing that RNA synthesis is an enzyme-catalyzed process.
1961
The Messenger RNA Hypothesis
François Jacob and Jacques Monod proposed that a short-lived RNA intermediate — mRNA — carries genetic information from DNA to ribosomes. Sydney Brenner, Jacob, and Matthew Meselson provided experimental evidence the same year.
1970
Sigma Factor and Promoter Recognition
Richard Burgess and Andrew Travers isolated the sigma (σ) subunit of bacterial RNA polymerase, demonstrating that specific protein factors direct the enzyme to promoter sequences on DNA.
1977
Discovery of Introns and RNA Splicing
Phillip Sharp and Richard Roberts independently discovered that eukaryotic genes contain non-coding intervening sequences (introns) that are removed from the primary transcript by RNA splicing, fundamentally revising views of gene structure.
1993
Spliceosome Structure Emerges
Biochemical and genetic studies revealed that intron removal is catalyzed by the spliceosome, a large ribonucleoprotein complex composed of five snRNPs. Later cryo-EM work provided atomic-level structural details.

The central question that unifies this lesson is: how does a cell accurately copy a specific segment of its genome into RNA, and what additional modifications must eukaryotic cells perform before that RNA can function? Understanding transcription and RNA processing is essential not only for the AP Biology exam but for appreciating how gene expression is regulated at multiple levels — from promoter access to alternative splicing — with profound implications for development, disease, and biotechnology.

Core Principles of Transcription

Transcription is the process by which the information stored in a DNA template is enzymatically copied into a complementary RNA molecule. Although the fundamental logic of transcription — base pairing between DNA and incoming ribonucleotides — is conserved across all domains of life, eukaryotic transcription involves greater complexity in terms of enzyme diversity, regulatory architecture, and post-transcriptional processing. Before exploring the mechanism in detail, it is important to establish the foundational principles that govern how transcription operates.

1

Template-Directed Synthesis

RNA polymerase reads the template strand (also called the antisense strand) of DNA in the 3′→5′ direction, synthesizing a complementary RNA molecule in the 5′→3′ direction. The resulting RNA has the same base sequence as the coding strand (sense strand), except with uracil (U) replacing thymine (T).
2

Promoter-Driven Initiation

Transcription begins at specific DNA sequences called promoters. In bacteria, the sigma factor directs RNA polymerase to consensus sequences (e.g., −10 and −35 elements). In eukaryotes, general transcription factors and RNA Polymerase II recognize the TATA box and other core promoter elements.
3

No Primer Required

Unlike DNA polymerase, RNA polymerase can initiate synthesis de novo — it does not require an RNA primer. This is a critical distinction from DNA replication and reflects the enzyme's ability to catalyze the first phosphodiester bond without a pre-existing 3′-OH group.
4

Selective Gene Activation

Not all genes are transcribed at the same time or in every cell type. Differential gene expression is achieved through combinatorial control by transcription factors, enhancers, silencers, and chromatin remodeling — allowing the same genome to produce vastly different cell phenotypes.
5

Eukaryotic RNA Processing

In eukaryotes, the primary transcript (pre-mRNA) undergoes three major processing events: 5′ capping, 3′ polyadenylation, and RNA splicing. These modifications are required for mRNA stability, nuclear export, and efficient translation.
KEY TAKEAWAY
Think of transcription like a recording studio: the DNA template strand is the master recording that never leaves the vault, RNA polymerase is the recording equipment that produces a working copy, and the mRNA is the distributed copy that gets sent out for use (translation). In eukaryotes, the raw recording (pre-mRNA) goes through an editing suite — capping, polyadenylation, and splicing — before it is released as a polished product.

Visual Overview of Transcription

The following diagram illustrates the three major stages of transcription — initiation, elongation, and termination — showing the spatial relationship between DNA, RNA polymerase, and the growing RNA transcript. Pay particular attention to strand polarity and the direction of polymerase movement.

The three stages of transcription are shown from left to right. During initiation, RNA polymerase (RNAP) binds the promoter and locally unwinds DNA. During elongation, RNAP moves along the template strand (3′→5′), synthesizing mRNA in the 5′→3′ direction. During termination, a terminator signal causes RNAP and the completed mRNA to dissociate from the DNA. Note the base-pairing rules in the lower panel: uracil replaces thymine in RNA.

Note that in the diagram, the coding strand runs 5′→3′ from left to right (pink), while the template strand runs antiparallel 3′→5′ (blue). RNA polymerase reads the template strand and synthesizes the mRNA in the 5′→3′ direction, which means the mRNA sequence is complementary to the template but essentially identical to the coding strand (with U instead of T). On the AP exam, you may be asked to determine the mRNA sequence given either the template or the coding strand — always keep the base-pairing rules and strand polarity in mind.

Detailed Mechanism of Transcription

Initiation: Assembling the Transcription Machinery

In eukaryotes, transcription of protein-coding genes is carried out by RNA Polymerase II (Pol II), a multi-subunit enzyme that cannot recognize promoters on its own. Instead, a series of general transcription factors (GTFs) — TFIIA, TFIIB, TFIID, TFIIE, TFIIF, and TFIIH — assemble in a stepwise manner at the core promoter to form the pre-initiation complex (PIC). The process begins when the TATA-binding protein (TBP), a subunit of TFIID, binds the TATA box approximately 25–30 base pairs upstream of the transcription start site. TFIIH then uses its helicase activity to unwind about 11–15 base pairs of DNA, creating the transcription bubble, and its kinase activity phosphorylates the C-terminal domain (CTD) of Pol II, triggering promoter clearance and the transition to elongation.

Elongation: Synthesizing the RNA Transcript

Once Pol II clears the promoter, it enters the elongation phase. The enzyme moves along the template strand in the 3′→5′ direction, adding complementary ribonucleoside triphosphates (NTPs) — ATP, UTP, GTP, and CTP — to the growing RNA chain in the 5′→3′ direction. Each incoming NTP is joined to the 3′-OH of the preceding nucleotide by a phosphodiester bond, releasing pyrophosphate (PPi), whose subsequent hydrolysis drives the reaction forward. The rate of elongation in eukaryotes is roughly 1,000–2,000 nucleotides per minute. Pol II maintains a short RNA:DNA hybrid of about 8 base pairs within the active site; behind the polymerase, the DNA double helix re-anneals and the nascent RNA strand peels away as a single-stranded molecule.

Termination: Releasing the Transcript

Termination mechanisms differ between prokaryotes and eukaryotes. In bacteria, two well-characterized pathways exist: rho-independent (intrinsic) termination, in which a GC-rich hairpin loop followed by a poly-U stretch in the nascent RNA destabilizes the RNA:DNA hybrid and causes dissociation, and rho-dependent termination, in which the rho (ρ) protein, an ATP-dependent helicase, catches up to a paused polymerase and unwinds the RNA:DNA hybrid. In eukaryotes, Pol II transcription of mRNA genes terminates via the cleavage-polyadenylation model: the pre-mRNA is cleaved at a specific site downstream of the AAUAAA polyadenylation signal, and the remaining RNA still being synthesized by Pol II is degraded by a 5′→3′ exonuclease, which eventually catches the polymerase and triggers termination.

🔬 Prokaryotes vs. Eukaryotes
In prokaryotes, transcription and translation are coupled — ribosomes begin translating the mRNA while it is still being transcribed, because there is no nuclear membrane separating the two processes. In eukaryotes, transcription occurs in the nucleus and the mature mRNA must be exported to the cytoplasm before translation can begin, creating opportunities for additional layers of regulation.

Eukaryotic RNA Processing

In eukaryotic cells, the primary transcript produced by RNA Polymerase II — called pre-mRNA — undergoes three critical processing events before it is exported from the nucleus as mature mRNA. These modifications occur co-transcriptionally (while transcription is still underway) and are essential for mRNA stability, proper nuclear export, and efficient translation. The three events are 5′ capping, 3′ polyadenylation, and RNA splicing.

This diagram traces the transformation of pre-mRNA into mature mRNA through four steps. The 5′ cap (green, 7-methylguanosine) is added first, followed by cleavage and polyadenylation at the 3′ end (red). Splicing removes introns (purple/pink) and joins exons (gold then cyan) to produce the continuous open reading frame of the mature mRNA.

5′ Capping

As soon as the first 20–30 nucleotides of the pre-mRNA emerge from RNA Polymerase II, capping enzymes add a 7-methylguanosine (7mG) residue to the 5′ end through an unusual 5′→5′ triphosphate linkage. This cap serves multiple functions: it protects the mRNA from degradation by 5′ exonucleases, facilitates recognition by the nuclear cap-binding complex (CBC) for export through nuclear pores, and is later recognized by the eukaryotic initiation factor eIF4E during translation initiation. Without the cap, the mRNA would be rapidly degraded and would fail to engage the ribosomal machinery.

3′ Polyadenylation

The 3′ end of the pre-mRNA is processed by endonucleolytic cleavage at a site 10–30 nucleotides downstream of the highly conserved AAUAAA polyadenylation signal. After cleavage, the enzyme poly-A polymerase (PAP) adds approximately 100–250 adenine nucleotides to form the poly-A tail. This tail enhances mRNA stability by slowing 3′ exonuclease degradation, assists in nuclear export, and stimulates translation efficiency. The length of the poly-A tail can influence mRNA lifespan — a shorter tail is associated with more rapid degradation.

RNA Splicing

Most eukaryotic protein-coding genes contain introns (non-coding intervening sequences) interspersed among exons (expressed sequences). In some human genes, introns may comprise more than 90% of the primary transcript. The removal of introns and the joining of exons is catalyzed by the spliceosome, a large ribonucleoprotein (RNP) complex composed of five small nuclear ribonucleoproteins (snRNPs: U1, U2, U4, U5, U6) and many associated proteins. Splicing proceeds through two sequential transesterification reactions. First, the 2′-OH of a conserved branch-point adenine within the intron attacks the 5′ splice site, forming a lariat intermediate. Second, the free 3′-OH of the upstream exon attacks the 3′ splice site, joining the two exons and releasing the intron lariat for degradation.

🧬 Alternative Splicing — Expanding Protein Diversity
A single gene can produce multiple mRNA variants through alternative splicing, in which different combinations of exons are included or excluded. This mechanism explains how the ~20,000 human protein-coding genes can give rise to an estimated 100,000+ distinct proteins. For example, the Drosophila DSCAM gene can theoretically produce over 38,000 mRNA isoforms through combinatorial exon selection. Alternative splicing is a critical concept for AP Biology because it illustrates how post-transcriptional regulation expands the informational capacity of the genome.

Worked Example: From DNA to Mature mRNA

Consider the following problem, which integrates transcription and RNA processing. You are given a segment of a eukaryotic gene and asked to determine the sequence of the mature mRNA.

Determining Mature mRNA Sequence
1
Step 1 — Identify the Given InformationA portion of a eukaryotic gene has the following coding (sense) strand sequence (written 5′→3′): 5′-ATG·GCA·|·GTA·AGG·TCA·|·TTT·GCC·TAA-3′. The vertical bars (|) denote exon-intron boundaries. Exon 1 = ATG GCA, Intron = GTA AGG TCA, Exon 2 = TTT GCC TAA. The template strand is the complement read 3′→5′.
2
Step 2 — Write the Pre-mRNA SequenceThe pre-mRNA has the same sequence as the coding strand but with U replacing T: 5′-AUG·GCA·GUA·AGG·UCA·UUU·GCC·UAA-3′. This is the primary transcript before any processing.
Pre-mRNA: 5′-AUG GCA GUA AGG UCA UUU GCC UAA-3′
3
Step 3 — Apply 5′ Capping and 3′ PolyadenylationA 7-methylguanosine cap is added to the 5′ end (shown as 7mG–), and a poly-A tail (~100–250 As) is added to the 3′ end. The coding sequence itself is not altered by these modifications, but they are crucial structural features of the mature mRNA.
7mG–5′-AUG GCA [GUA AGG UCA] UUU GCC UAA-3′-AAAAAA...A
4
Step 4 — Remove the Intron via SplicingThe spliceosome recognizes the intron boundaries and removes the intron sequence GUA AGG UCA, joining Exon 1 directly to Exon 2. The exons are ligated by a phosphodiester bond.
Mature mRNA: 7mG–5′-AUG GCA UUU GCC UAA-3′-AAAAAA...A
5
Step 5 — Verify the Reading FrameReading the mature mRNA in triplet codons from the start codon AUG: AUG (Met) – GCA (Ala) – UUU (Phe) – GCC (Ala) – UAA (stop). The mRNA encodes a 4-amino-acid polypeptide: Met-Ala-Phe-Ala. Notice that the intron interrupted the coding sequence but, upon removal, the reading frame was preserved — this is a fundamental requirement of proper splicing.
Polypeptide: Met-Ala-Phe-Ala

Prokaryotic vs. Eukaryotic Transcription

Although the fundamental chemistry of RNA synthesis is conserved, prokaryotic and eukaryotic transcription differ significantly in terms of enzyme complexity, regulatory mechanisms, and post-transcriptional processing. The following table summarizes the key distinctions that are frequently tested on the AP Biology exam.

Key differences between prokaryotic and eukaryotic transcription
FeatureProkaryotesEukaryotes
RNA PolymeraseSingle type (core enzyme + σ factor)Three types: Pol I (rRNA), Pol II (mRNA), Pol III (tRNA, 5S rRNA)
Promoter Elements−10 (Pribnow box) and −35 consensus sequencesTATA box, Inr, DPE; enhancers may be thousands of bp away
Transcription Factorsσ factor for initiation; few additional regulatorsMultiple GTFs (TFIIA–H); Mediator complex; numerous activators/repressors
mRNA ProcessingNone (mRNA is ready for translation immediately)5′ cap, poly-A tail, intron splicing
Coupling with TranslationYes — co-transcriptional translationNo — transcription in nucleus, translation in cytoplasm
TerminationRho-dependent or rho-independent (intrinsic)Cleavage-polyadenylation coupled with torpedo model
Gene StructureGenes typically lack introns; operons group related genesMost genes contain introns; monocistronic mRNAs
KEY TAKEAWAY
Think of prokaryotic gene expression as a streamlined assembly line with minimal quality control — the product (mRNA) goes directly to use. Eukaryotic gene expression is more like a pharmaceutical manufacturing pipeline, with multiple checkpoints (capping, splicing, polyadenylation, export) that ensure only properly processed transcripts reach the translation machinery. This added complexity allows for far more sophisticated regulation but also creates more potential points of failure, which is relevant to understanding how mutations in splicing signals can cause genetic disease.

Connections to Regulation and Disease

Transcription and RNA processing are not merely mechanical steps in gene expression — they are heavily regulated and serve as critical control points that determine which proteins a cell produces, when, and in what quantities. Understanding these connections is essential for the AP Biology exam and for appreciating the broader significance of gene regulation in development, cellular differentiation, and disease.

How transcription and RNA processing connect to advanced topics
Concept in This LessonAdvanced Connection
Promoter recognition by Pol IIEpigenetic modifications (DNA methylation, histone acetylation) regulate promoter accessibility. CpG island hypermethylation can silence tumor suppressor genes, contributing to cancer.
Enhancers and transcription factorsCombinatorial control by multiple transcription factors enables cell-type-specific gene expression. Mutations in enhancers can cause developmental disorders (e.g., limb malformations from Shh enhancer mutations).
Alternative splicingAberrant splicing accounts for ~15% of human genetic diseases. For example, certain mutations in the SMN1 gene cause spinal muscular atrophy by disrupting normal splicing patterns.
Poly-A tail lengthmRNA stability regulation through poly-A tail shortening (deadenylation) is a post-transcriptional control mechanism. MicroRNAs (miRNAs) can accelerate deadenylation and mRNA degradation.
5′ cap recognitionmRNA vaccines (e.g., COVID-19) use synthetic mRNAs with modified 5′ caps and optimized UTRs to enhance stability and translation in host cells — a direct biomedical application of RNA processing principles.

As you advance beyond the foundational material covered in this lesson, keep in mind that transcription is the first — and often the most important — regulatory step in gene expression. The concepts of transcriptional regulation (operons in prokaryotes, enhancers and chromatin remodeling in eukaryotes) and post-transcriptional regulation (alternative splicing, mRNA stability, RNA interference) are deeply tested on the AP Biology exam. Future lessons will explore these regulatory layers in detail, building upon the mechanistic foundations established here.

Practice Problems

1
Which of the following best explains why eukaryotic mRNA must undergo processing before translation, whereas prokaryotic mRNA does not?
2
A segment of the template (antisense) strand of a gene reads 3′-TACGGATTC-5′. What is the sequence of the mRNA transcribed from this template?
3
A researcher discovers a mutation in a eukaryotic gene that changes the conserved GU dinucleotide at the 5′ splice site of intron 3 to an AU dinucleotide. Which of the following is the most likely consequence of this mutation?
PROBLEM 4APPLIED
A research team hypothesizes that a newly identified protein, Factor X, is required for proper 5′ capping of mRNA in eukaryotic cells. Design an experiment to test this hypothesis. In your response: (a) Describe the experimental setup, including appropriate controls. (b) Identify the independent and dependent variables. (c) Describe the expected results if the hypothesis is supported. (d) Explain how improper 5′ capping would affect downstream gene expression.
PROBLEM 5CRITICAL THINKING
Researchers studying the human tropomyosin gene find that this single gene produces different mRNA isoforms in skeletal muscle cells versus fibroblasts due to alternative splicing. In skeletal muscle, exons 1, 2, 3, 5, 7, 8, and 9 are included, while in fibroblasts, exons 1, 2, 4, 6, 8, and 9 are included. The pre-mRNA contains 9 exons and 8 introns. (a) Explain the molecular mechanism that allows a single gene to produce tissue-specific mRNA variants. (b) Predict how the proteins produced from these two mRNA isoforms would differ structurally and functionally. (c) A mutation eliminates the branch-point sequence in intron 4 (between exons 4 and 5). Predict the effect of this mutation on mRNA production in each tissue type. (d) Explain why alternative splicing is considered a mechanism that increases proteomic diversity beyond what would be predicted by gene number alone.

Lesson Summary

Transcription is the process by which RNA polymerase reads the template strand of DNA (3′→5′) and synthesizes a complementary RNA molecule (5′→3′). The process proceeds through three stages: initiation at a promoter (involving general transcription factors and the pre-initiation complex in eukaryotes), elongation (NTP addition via phosphodiester bond formation), and termination (rho-dependent/independent in prokaryotes; cleavage-polyadenylation in eukaryotes).

In eukaryotes, the primary transcript (pre-mRNA) undergoes three essential processing modifications: addition of a 5′ 7-methylguanosine cap for stability and translation initiation, addition of a 3′ poly-A tail (~100–250 adenines) for stability and export, and removal of introns by the spliceosome (composed of snRNPs), which joins exons into a continuous coding sequence. Alternative splicing allows a single gene to produce multiple mRNA isoforms, vastly expanding the proteomic diversity of eukaryotic organisms. These processing steps are absent in prokaryotes, where transcription and translation are coupled in the cytoplasm.

Varsity Tutors • AP Biology • Transcription and RNA Processing