IB BIOLOGY • CONTINUITY AND CHANGE

Apply Gene Expression

Discover how cells read DNA instructions to build the proteins that drive every living process.

Historical Context & Motivation

For centuries, people knew that offspring resemble their parents, but no one understood the molecular machinery behind inheritance. The discovery of DNA's structure in 1953 was a breakthrough, yet it raised an equally important question: how does a sequence of nucleotides actually become a functional protein? The story of gene expression — the process by which genetic information flows from DNA to RNA to protein — is one of the most important narratives in modern biology.

1941
One Gene–One Enzyme Hypothesis
George Beadle and Edward Tatum demonstrated that individual genes direct the synthesis of specific enzymes, establishing the first clear link between genotype and phenotype.
1953
Structure of DNA Revealed
James Watson and Francis Crick, building on X-ray data from Rosalind Franklin, published the double-helix model. This structure immediately suggested a mechanism for copying genetic information.
1958
The Central Dogma
Francis Crick proposed the central dogma of molecular biology: information flows from DNA → RNA → protein. This framework guides our understanding of gene expression to this day.
1961
Cracking the Genetic Code
Marshall Nirenberg and Heinrich Matthaei deciphered the first codon (UUU = phenylalanine), opening the door to translating the entire genetic code into amino acid sequences.
1977
Discovery of Introns and Splicing
Richard Roberts and Phillip Sharp independently discovered that eukaryotic genes contain non-coding sequences called introns, which must be removed from mRNA before translation.

These discoveries collectively answered a pivotal question: how does the information stored in a linear sequence of DNA bases get expressed as a three-dimensional, functional protein? Understanding gene expression is essential because it explains how a single fertilized egg can produce hundreds of different cell types, how organisms respond to their environment, and how mutations can lead to disease.

Core Principles of Gene Expression

Gene expression has two major stages: transcription and translation. Between these steps in eukaryotic cells, the initial RNA transcript undergoes RNA processing. Together, these processes convert the language of nucleotides into the language of amino acids, ultimately producing the proteins that carry out nearly every function in a living organism.

1

Transcription

RNA polymerase reads the template strand of DNA (3′ → 5′) and synthesizes a complementary mRNA strand in the 5′ → 3′ direction. This occurs in the nucleus of eukaryotic cells.
2

RNA Processing (Eukaryotes)

The pre-mRNA transcript receives a 5′ cap, a 3′ poly-A tail, and has its introns removed by spliceosomes. The remaining exons are joined together to form mature mRNA.
3

Translation

Ribosomes read mRNA in sets of three nucleotides called codons. Each codon is matched by a tRNA carrying a specific amino acid. Amino acids are linked by peptide bonds into a polypeptide chain.
4

The Genetic Code

The genetic code is universal (shared across nearly all organisms), degenerate (multiple codons can code for the same amino acid), and non-overlapping (codons are read sequentially without gaps).
5

Gene Regulation

Not all genes are expressed at all times. Cells regulate gene expression through mechanisms like transcription factors, epigenetic modifications, and mRNA degradation, ensuring the right proteins are made in the right cells at the right time.
KEY TAKEAWAY
Think of gene expression like a recipe book in a locked library. Transcription is like making a photocopy of one recipe (mRNA) from the master book (DNA) so you can take it to the kitchen. Translation is the actual cooking — the ribosome reads the recipe and assembles the ingredients (amino acids) into the finished dish (protein). The library never leaves the building (nucleus), but the recipe copy travels to where the work is done (cytoplasm).

Visualizing Gene Expression

This diagram shows the complete flow of gene expression. On the left, within the nucleus, DNA is transcribed into pre-mRNA, which is then processed (capped, spliced, and tailed) to form mature mRNA. The mature mRNA exits through a nuclear pore into the cytoplasm, where ribosomes translate it into a polypeptide chain that folds into a functional protein.

Notice that the diagram illustrates two key compartments in a eukaryotic cell. Transcription and RNA processing take place inside the nucleus, where DNA is safely housed. Only the mature, processed mRNA is allowed to leave the nucleus through nuclear pores. Translation then occurs in the cytoplasm at ribosomes. This physical separation allows eukaryotic cells an extra level of control over gene expression that prokaryotic cells (which lack a nucleus) do not have — in prokaryotes, transcription and translation can occur simultaneously.

The Mechanism of Gene Expression in Detail

Transcription: Reading the DNA Template

Transcription begins when RNA polymerase binds to a region of DNA called the promoter. In eukaryotes, transcription factors must first assemble at the promoter before RNA polymerase can attach. Once bound, the enzyme unwinds the double helix and reads the template strand in the 3′ → 5′ direction, synthesizing a complementary mRNA strand in the 5′ → 3′ direction. The enzyme follows base-pairing rules: adenine (A) on DNA pairs with uracil (U) on mRNA, thymine (T) pairs with adenine (A), guanine (G) pairs with cytosine (C), and cytosine pairs with guanine. Transcription ends when RNA polymerase reaches a terminator sequence, and the newly made pre-mRNA is released.

RNA Processing: Preparing the Message

Before the pre-mRNA can leave the nucleus, it undergoes three critical modifications. First, a modified guanine nucleotide called the 5′ cap is added to the beginning of the transcript, which helps ribosomes recognize the mRNA and protects it from degradation. Second, an enzyme adds a chain of 100–250 adenine nucleotides to the 3′ end, forming the poly-A tail, which also protects the mRNA and aids in export from the nucleus. Third, spliceosomes — large complexes of RNA and protein — remove the introns (non-coding sequences) and join the exons (coding sequences). Through alternative splicing, a single gene can produce multiple different proteins by including different combinations of exons, greatly increasing protein diversity.

Translation: Building the Protein

Translation takes place at the ribosome and involves three main phases: initiation, elongation, and termination. During initiation, the small ribosomal subunit binds to the mRNA at the 5′ cap and scans along until it finds the start codon AUG, which codes for the amino acid methionine. A tRNA molecule carrying methionine binds to this codon through complementary base pairing of its anticodon (UAC). The large ribosomal subunit then joins to complete the ribosome.

During elongation, the ribosome moves along the mRNA one codon at a time. At each codon, a tRNA with a matching anticodon delivers the appropriate amino acid. The ribosome catalyzes the formation of a peptide bond between adjacent amino acids, and the growing polypeptide chain extends. Translation terminates when the ribosome encounters a stop codon (UAA, UAG, or UGA). No tRNA recognizes stop codons; instead, release factors bind and cause the ribosome to disassemble, releasing the completed polypeptide.

🧬 Codon–Anticodon Pairing
Remember that codons are read on the mRNA (5′ → 3′), and anticodons are on the tRNA (3′ → 5′). They bind by complementary base pairing. For example, the mRNA codon 5′-AUG-3′ pairs with the tRNA anticodon 3′-UAC-5′.

The Genetic Code & Codon Table

The genetic code is the set of rules by which information encoded in mRNA sequences is translated into amino acid sequences. With four possible bases at each position and codons three nucleotides long, there are 4³ = 64 possible codons. These 64 codons specify only 20 amino acids plus the stop signal, which means the code is degenerate — most amino acids are encoded by more than one codon. This redundancy provides a buffer against the effects of some point mutations.

The codon table above lists all 64 mRNA codons and their corresponding amino acids. The start codon (AUG) is highlighted in green, while the three stop codons (UAA, UAG, UGA) are shown in red. Notice how multiple codons can code for the same amino acid — this is the degeneracy of the genetic code.

To use the codon table, you read an mRNA sequence three bases at a time starting from the start codon (AUG). For each codon, find the first base in the left column, the second base along the top row, and the third base narrows it to one specific amino acid. For instance, the codon GCA codes for alanine (Ala). Practicing this skill is essential for IB Biology because exam questions frequently ask you to determine an amino acid sequence from a given DNA or mRNA template.

Worked Example: From DNA to Polypeptide

Let's walk through a complete gene expression problem. You are given the following DNA template strand sequence and asked to determine the amino acid sequence of the polypeptide produced.

Determining an Amino Acid Sequence from a DNA Template Strand
1
Step 1 — Identify the DNA Template StrandYou are given the DNA template strand: 3′ − TAC GGA CTC AAA ATT − 5′. Remember that RNA polymerase reads the template strand in the 3′ → 5′ direction to build the mRNA in the 5′ → 3′ direction.
Template strand identified: 3′ TAC GGA CTC AAA ATT 5′
2
Step 2 — Transcribe DNA to mRNAApply the base-pairing rules. Each DNA base on the template strand pairs with its RNA complement: T → A, A → U, C → G, G → C. Reading the template strand from 3′ to 5′, we build the mRNA from 5′ to 3′.
mRNA: 5′ − AUG CCU GAG UUU UAA − 3′
3
Step 3 — Divide mRNA into CodonsStarting from the start codon AUG, separate the mRNA into triplets (codons): AUG | CCU | GAG | UUU | UAA. Notice that AUG is the start codon and UAA is a stop codon.
Codons: AUG, CCU, GAG, UUU, UAA
4
Step 4 — Translate Each Codon Using the Codon TableLook up each codon in the codon table. AUG = Methionine (Met), CCU = Proline (Pro), GAG = Glutamic acid (Glu), UUU = Phenylalanine (Phe), UAA = Stop. Translation begins at AUG and ends when the ribosome reaches UAA.
Amino acid sequence: Met − Pro − Glu − Phe
5
Step 5 — State the Final PolypeptideThe resulting polypeptide is four amino acids long. In many organisms, the initial methionine may be removed after translation during post-translational modification. However, for IB Biology purposes, include methionine in your answer unless the question specifies otherwise. This polypeptide would then fold into its functional three-dimensional shape.
Final polypeptide: Met–Pro–Glu–Phe (4 amino acids)

Comparing Gene Expression in Prokaryotes & Eukaryotes

Although the basic principles of transcription and translation are conserved across life, there are important differences between how prokaryotic and eukaryotic cells carry out gene expression. Understanding these differences is a common IB exam topic.

Key differences in gene expression between prokaryotes and eukaryotes
FeatureProkaryotesEukaryotes
Location of transcriptionCytoplasm (no nucleus)Nucleus
Location of translationCytoplasmCytoplasm (rough ER or free ribosomes)
Simultaneous transcription & translationYes — translation begins before transcription finishesNo — mRNA must be processed and exported first
RNA processingMinimal — no introns, no splicing, no 5′ cap or poly-A tailExtensive — intron removal, 5′ capping, poly-A tail addition
Introns presentRarelyYes — often many introns per gene
RNA polymerase typesOne typeThree types (I, II, III); RNA Pol II transcribes mRNA
Ribosomes70S (50S + 30S subunits)80S (60S + 40S subunits)
KEY TAKEAWAY
The biggest structural difference is that eukaryotic cells have a nucleus that physically separates transcription from translation, creating a checkpoint for RNA processing. Think of it like mailing a letter: in prokaryotes, you hand the letter directly to the recipient (simultaneous transcription and translation). In eukaryotes, the letter goes through a post office (RNA processing in the nucleus) before it reaches its destination (ribosome in the cytoplasm).

Gene Regulation & the Impact of Mutations

Gene expression does not operate in an unregulated manner. Cells carefully control which genes are turned on or off, and changes to the DNA sequence — mutations — can alter gene expression with significant consequences. Understanding regulation and mutations connects gene expression to broader IB topics such as evolution, biotechnology, and disease.

Types of gene mutations and their effects on protein structure
Mutation TypeDescriptionEffect on Protein
Silent mutationA base substitution that results in a codon for the same amino acid (due to degeneracy).No change — the protein sequence is identical.
Missense mutationA base substitution that changes one amino acid to a different one.May alter protein function (e.g., sickle cell anemia: Glu → Val in hemoglobin).
Nonsense mutationA base substitution that creates a premature stop codon.Truncated (shortened) protein — usually nonfunctional.
Frameshift mutationAn insertion or deletion of one or two bases shifts the reading frame.All downstream amino acids change — usually devastating to protein function.

Cells regulate gene expression at multiple levels. At the transcriptional level, transcription factors can activate or repress genes by binding to specific DNA sequences near the promoter. Epigenetic mechanisms such as DNA methylation and histone modification can make genes more or less accessible without changing the DNA sequence itself. At the post-transcriptional level, the stability of mRNA and the efficiency of translation can be regulated. These layers of control explain how a skin cell and a neuron can contain identical DNA yet perform radically different functions — they express different subsets of genes.

🔬 Connecting to IB: Sickle Cell Anemia
Sickle cell anemia is a classic IB example of how a single base substitution (A → T in the template strand) changes one codon from GAG to GUG, replacing glutamic acid with valine in the β-globin chain of hemoglobin. This one amino acid change causes the protein to polymerize under low-oxygen conditions, deforming red blood cells into a sickle shape. It beautifully illustrates how gene expression links genotype to phenotype.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain why the genetic code is described as 'degenerate' and discuss one advantage this provides to organisms.
PROBLEM 2BASIC CALCULATION
A DNA template strand has the sequence: 3′ − TAC AAG CGA ATC − 5′. Determine the mRNA sequence and identify the amino acid sequence produced.
PROBLEM 3INTERMEDIATE
A mutation changes the DNA template strand sequence from 3′ − TAC − 5′ to 3′ − TAA − 5′ at the start of a gene. Predict the effect of this mutation on gene expression and explain your reasoning.
PROBLEM 4APPLIED
A researcher isolates mRNA from a eukaryotic cell and finds that it is 900 nucleotides long (excluding the 5′ cap and poly-A tail). However, the gene from which it was transcribed is 2,700 nucleotides long. Explain this discrepancy and calculate the maximum number of amino acids the protein could contain.
PROBLEM 5CRITICAL THINKING
In prokaryotes, transcription and translation are coupled — ribosomes begin translating the mRNA while it is still being transcribed. In eukaryotes, this does not occur. Analyze how this difference relates to (a) the presence of introns in eukaryotic genes, and (b) the ability of eukaryotic cells to regulate gene expression more precisely than prokaryotes.

Summary of Gene Expression

Gene expression is the two-stage process by which the information in DNA is used to build functional proteins. During transcription, RNA polymerase reads the template strand (3′ → 5′) and synthesizes a complementary mRNA molecule (5′ → 3′). In eukaryotes, the pre-mRNA is processed — receiving a 5′ cap, poly-A tail, and having introns removed by spliceosomes — before the mature mRNA exits the nucleus.

During translation, ribosomes read mRNA codons (sets of three nucleotides) and assemble amino acids delivered by tRNA molecules into a polypeptide chain. The genetic code is universal, degenerate, and non-overlapping. Mutations — including silent, missense, nonsense, and frameshift types — can alter the protein produced, linking molecular changes to phenotypic effects such as sickle cell anemia. Cells regulate gene expression at multiple levels, ensuring the right proteins are made at the right time in the right cells.

Varsity Tutors • IB Biology • Apply Gene Expression