Historical Context & Motivation
For centuries, people knew that offspring resemble their parents, but no one understood the molecular machinery behind inheritance. The discovery of DNA's structure in 1953 was a breakthrough, yet it raised an equally important question: how does a sequence of nucleotides actually become a functional protein? The story of gene expression — the process by which genetic information flows from DNA to RNA to protein — is one of the most important narratives in modern biology.
These discoveries collectively answered a pivotal question: how does the information stored in a linear sequence of DNA bases get expressed as a three-dimensional, functional protein? Understanding gene expression is essential because it explains how a single fertilized egg can produce hundreds of different cell types, how organisms respond to their environment, and how mutations can lead to disease.
Core Principles of Gene Expression
Gene expression has two major stages: transcription and translation. Between these steps in eukaryotic cells, the initial RNA transcript undergoes RNA processing. Together, these processes convert the language of nucleotides into the language of amino acids, ultimately producing the proteins that carry out nearly every function in a living organism.
Transcription
RNA Processing (Eukaryotes)
Translation
The Genetic Code
Gene Regulation
Visualizing Gene Expression
Notice that the diagram illustrates two key compartments in a eukaryotic cell. Transcription and RNA processing take place inside the nucleus, where DNA is safely housed. Only the mature, processed mRNA is allowed to leave the nucleus through nuclear pores. Translation then occurs in the cytoplasm at ribosomes. This physical separation allows eukaryotic cells an extra level of control over gene expression that prokaryotic cells (which lack a nucleus) do not have — in prokaryotes, transcription and translation can occur simultaneously.
The Mechanism of Gene Expression in Detail
Transcription: Reading the DNA Template
Transcription begins when RNA polymerase binds to a region of DNA called the promoter. In eukaryotes, transcription factors must first assemble at the promoter before RNA polymerase can attach. Once bound, the enzyme unwinds the double helix and reads the template strand in the 3′ → 5′ direction, synthesizing a complementary mRNA strand in the 5′ → 3′ direction. The enzyme follows base-pairing rules: adenine (A) on DNA pairs with uracil (U) on mRNA, thymine (T) pairs with adenine (A), guanine (G) pairs with cytosine (C), and cytosine pairs with guanine. Transcription ends when RNA polymerase reaches a terminator sequence, and the newly made pre-mRNA is released.
RNA Processing: Preparing the Message
Before the pre-mRNA can leave the nucleus, it undergoes three critical modifications. First, a modified guanine nucleotide called the 5′ cap is added to the beginning of the transcript, which helps ribosomes recognize the mRNA and protects it from degradation. Second, an enzyme adds a chain of 100–250 adenine nucleotides to the 3′ end, forming the poly-A tail, which also protects the mRNA and aids in export from the nucleus. Third, spliceosomes — large complexes of RNA and protein — remove the introns (non-coding sequences) and join the exons (coding sequences). Through alternative splicing, a single gene can produce multiple different proteins by including different combinations of exons, greatly increasing protein diversity.
Translation: Building the Protein
Translation takes place at the ribosome and involves three main phases: initiation, elongation, and termination. During initiation, the small ribosomal subunit binds to the mRNA at the 5′ cap and scans along until it finds the start codon AUG, which codes for the amino acid methionine. A tRNA molecule carrying methionine binds to this codon through complementary base pairing of its anticodon (UAC). The large ribosomal subunit then joins to complete the ribosome.
During elongation, the ribosome moves along the mRNA one codon at a time. At each codon, a tRNA with a matching anticodon delivers the appropriate amino acid. The ribosome catalyzes the formation of a peptide bond between adjacent amino acids, and the growing polypeptide chain extends. Translation terminates when the ribosome encounters a stop codon (UAA, UAG, or UGA). No tRNA recognizes stop codons; instead, release factors bind and cause the ribosome to disassemble, releasing the completed polypeptide.
The Genetic Code & Codon Table
The genetic code is the set of rules by which information encoded in mRNA sequences is translated into amino acid sequences. With four possible bases at each position and codons three nucleotides long, there are 4³ = 64 possible codons. These 64 codons specify only 20 amino acids plus the stop signal, which means the code is degenerate — most amino acids are encoded by more than one codon. This redundancy provides a buffer against the effects of some point mutations.
To use the codon table, you read an mRNA sequence three bases at a time starting from the start codon (AUG). For each codon, find the first base in the left column, the second base along the top row, and the third base narrows it to one specific amino acid. For instance, the codon GCA codes for alanine (Ala). Practicing this skill is essential for IB Biology because exam questions frequently ask you to determine an amino acid sequence from a given DNA or mRNA template.
Worked Example: From DNA to Polypeptide
Let's walk through a complete gene expression problem. You are given the following DNA template strand sequence and asked to determine the amino acid sequence of the polypeptide produced.
Comparing Gene Expression in Prokaryotes & Eukaryotes
Although the basic principles of transcription and translation are conserved across life, there are important differences between how prokaryotic and eukaryotic cells carry out gene expression. Understanding these differences is a common IB exam topic.
| Feature | Prokaryotes | Eukaryotes |
|---|---|---|
| Location of transcription | Cytoplasm (no nucleus) | Nucleus |
| Location of translation | Cytoplasm | Cytoplasm (rough ER or free ribosomes) |
| Simultaneous transcription & translation | Yes — translation begins before transcription finishes | No — mRNA must be processed and exported first |
| RNA processing | Minimal — no introns, no splicing, no 5′ cap or poly-A tail | Extensive — intron removal, 5′ capping, poly-A tail addition |
| Introns present | Rarely | Yes — often many introns per gene |
| RNA polymerase types | One type | Three types (I, II, III); RNA Pol II transcribes mRNA |
| Ribosomes | 70S (50S + 30S subunits) | 80S (60S + 40S subunits) |
Gene Regulation & the Impact of Mutations
Gene expression does not operate in an unregulated manner. Cells carefully control which genes are turned on or off, and changes to the DNA sequence — mutations — can alter gene expression with significant consequences. Understanding regulation and mutations connects gene expression to broader IB topics such as evolution, biotechnology, and disease.
| Mutation Type | Description | Effect on Protein |
|---|---|---|
| Silent mutation | A base substitution that results in a codon for the same amino acid (due to degeneracy). | No change — the protein sequence is identical. |
| Missense mutation | A base substitution that changes one amino acid to a different one. | May alter protein function (e.g., sickle cell anemia: Glu → Val in hemoglobin). |
| Nonsense mutation | A base substitution that creates a premature stop codon. | Truncated (shortened) protein — usually nonfunctional. |
| Frameshift mutation | An insertion or deletion of one or two bases shifts the reading frame. | All downstream amino acids change — usually devastating to protein function. |
Cells regulate gene expression at multiple levels. At the transcriptional level, transcription factors can activate or repress genes by binding to specific DNA sequences near the promoter. Epigenetic mechanisms such as DNA methylation and histone modification can make genes more or less accessible without changing the DNA sequence itself. At the post-transcriptional level, the stability of mRNA and the efficiency of translation can be regulated. These layers of control explain how a skin cell and a neuron can contain identical DNA yet perform radically different functions — they express different subsets of genes.
Practice Problems
Summary of Gene Expression
Gene expression is the two-stage process by which the information in DNA is used to build functional proteins. During transcription, RNA polymerase reads the template strand (3′ → 5′) and synthesizes a complementary mRNA molecule (5′ → 3′). In eukaryotes, the pre-mRNA is processed — receiving a 5′ cap, poly-A tail, and having introns removed by spliceosomes — before the mature mRNA exits the nucleus.
During translation, ribosomes read mRNA codons (sets of three nucleotides) and assemble amino acids delivered by tRNA molecules into a polypeptide chain. The genetic code is universal, degenerate, and non-overlapping. Mutations — including silent, missense, nonsense, and frameshift types — can alter the protein produced, linking molecular changes to phenotypic effects such as sickle cell anemia. Cells regulate gene expression at multiple levels, ensuring the right proteins are made at the right time in the right cells.