IB BIOLOGY • CONTINUITY AND CHANGE

Understand Gene Expression

How cells read DNA instructions and build the proteins that drive all life processes.

Historical Context & Motivation

Every cell in your body contains the same DNA, yet a neuron looks and behaves nothing like a red blood cell. The question of how identical genetic information produces wildly different cell types puzzled biologists for decades. Understanding gene expression — the process by which information encoded in a gene is used to build a functional product, usually a protein — became one of the most transformative goals in modern biology. The discoveries that revealed this process did not happen overnight; they unfolded across the twentieth century as researchers connected genetics, biochemistry, and molecular biology.

1941
One Gene–One Enzyme Hypothesis
Beadle and Tatum used the bread mold Neurospora to demonstrate that each gene directs the production of a specific enzyme, establishing the first clear link between genes and proteins.
1953
Structure of DNA Revealed
Watson and Crick described the double-helix structure of DNA, providing the physical basis for understanding how genetic information is stored and copied.
1958
The Central Dogma
Francis Crick proposed the central dogma of molecular biology: information flows from DNA → RNA → Protein, framing the conceptual pathway of gene expression.
1961
Cracking the Genetic Code
Nirenberg and Matthaei deciphered the first codon (UUU → phenylalanine), launching the effort to map all 64 codons in the genetic code.
1977
Discovery of Introns and Splicing
Sharp and Roberts independently discovered that eukaryotic genes contain non-coding sequences called introns, which are removed during RNA processing before a protein can be made.

These breakthroughs raised a central question that still drives biological research today: if all cells share the same genome, what determines which genes are turned on or off in any given cell? Answering this question requires understanding the full pathway of gene expression, from the unwinding of DNA in the nucleus to the folding of a finished protein in the cytoplasm.

Core Principles of Gene Expression

Gene expression can be broken down into a set of foundational principles that apply across all living organisms. While the details differ between prokaryotes and eukaryotes, the core logic remains the same: DNA stores information as a sequence of nucleotide bases, this information is copied into an RNA intermediate, and that RNA message directs the assembly of a polypeptide chain. The following concepts form the backbone of this process.

1

Transcription

The enzyme RNA polymerase reads one strand of DNA (the template strand) and synthesizes a complementary messenger RNA (mRNA) molecule in the 5′ → 3′ direction.
2

RNA Processing

In eukaryotes, the initial mRNA transcript (pre-mRNA) undergoes modifications: a 5′ cap and 3′ poly-A tail are added, and introns are removed by splicing.
3

Translation

Ribosomes read the mRNA in sets of three nucleotides called codons. Each codon specifies an amino acid, delivered by transfer RNA (tRNA) molecules carrying complementary anticodons.
4

The Genetic Code

The code is universal (shared by nearly all organisms), degenerate (multiple codons can code for the same amino acid), and non-overlapping (codons are read sequentially without sharing bases).
5

Regulation

Cells control gene expression at multiple levels — before, during, and after transcription — using transcription factors, epigenetic modifications, and RNA degradation pathways to ensure the right proteins are made at the right time.
KEY TAKEAWAY
Think of gene expression like a recipe system. Your DNA is the master cookbook stored safely in the kitchen (nucleus). When a cell needs a dish (protein), it doesn't drag the whole cookbook to the stove — instead, it copies just the recipe it needs onto a slip of paper (mRNA). That slip is carried to the kitchen counter (ribosome), where ingredients (amino acids) are assembled step by step according to the instructions. Regulation is like deciding which recipe to copy based on what the family is hungry for today.

Visualizing the Central Dogma

The diagram below illustrates the flow of genetic information from DNA through RNA to protein. Notice how the process moves from the nucleus (where transcription occurs) to the cytoplasm (where translation takes place at the ribosome). Each stage involves distinct molecular machinery, and the diagram highlights the key players at each step.

The central dogma of molecular biology. Transcription occurs in the nucleus, where DNA is copied into pre-mRNA. After processing (capping, splicing, polyadenylation), mature mRNA is exported to the cytoplasm, where translation at the ribosome assembles amino acids into a polypeptide chain.

In the diagram, notice the clear division between nucleus and cytoplasm. Transcription begins when RNA polymerase binds to a region of DNA called the promoter and unwinds the double helix. It reads the template strand in the 3′ → 5′ direction, building the mRNA strand in the 5′ → 3′ direction. The resulting pre-mRNA is then processed: introns are removed, a protective cap is added to the 5′ end, and a poly-A tail is added to the 3′ end. Only after these modifications can the mature mRNA exit the nucleus through nuclear pores and reach a ribosome for translation.

The Mechanism of Gene Expression in Detail

Transcription: From DNA to mRNA

Transcription occurs in three stages. During initiation, RNA polymerase (along with transcription factors in eukaryotes) binds to the promoter region upstream of the gene. The promoter is not transcribed but acts as a signal that tells the enzyme where to start. In the elongation phase, RNA polymerase moves along the template strand, adding complementary RNA nucleotides one at a time. Remember that in RNA, uracil (U) replaces thymine (T), so wherever the DNA template has an adenine (A), the mRNA will have a uracil. Finally, during termination, the polymerase encounters a terminator sequence and releases the newly formed mRNA transcript.

RNA Processing (Eukaryotes Only)

Before the mRNA can leave the nucleus, three key modifications occur. First, a modified guanine nucleotide called the 5′ cap is added to the beginning of the transcript, which protects the mRNA from degradation and helps ribosomes recognize it. Second, a string of 100–200 adenine nucleotides called the poly-A tail is attached to the 3′ end, further stabilizing the molecule. Third, and perhaps most importantly, RNA splicing removes non-coding sequences (introns) and joins coding sequences (exons) together. A molecular machine called the spliceosome carries out this precise cut-and-paste operation.

Translation: From mRNA to Protein

Translation also proceeds through initiation, elongation, and termination. During initiation, the small ribosomal subunit binds to the mRNA and scans for the start codon (AUG), which codes for the amino acid methionine. A tRNA carrying methionine binds to this codon, and the large ribosomal subunit joins to form the complete ribosome. During elongation, the ribosome moves along the mRNA three bases at a time. At each codon, a matching tRNA delivers its amino acid, and a peptide bond forms between adjacent amino acids. Translation continues until the ribosome reaches a stop codon (UAA, UAG, or UGA), at which point a release factor causes the ribosome to detach and the completed polypeptide to be released.

🧬 Base-Pairing Rules in Gene Expression
During transcription, DNA bases pair with RNA bases: A → U, T → A, G → C, C → G. During translation, mRNA codons pair with tRNA anticodons using standard RNA base pairing: A–U and G–C. Always remember that RNA uses uracil instead of thymine.

Regulation of Gene Expression

Not every gene is expressed all the time. In fact, most cells use only a fraction of their genome at any given moment. Gene expression is regulated at multiple levels, and understanding these control points is essential for grasping how a single genome can produce hundreds of distinct cell types. The diagram below illustrates the major levels at which regulation can occur.

Gene expression can be regulated at five key levels: epigenetic (chromatin structure), transcriptional (transcription factors and promoters), post-transcriptional (RNA splicing and stability), translational (ribosome activity), and post-translational (protein modification and degradation).

The most significant control point is usually at the transcriptional level, because it is energetically efficient to prevent unwanted mRNAs from being made in the first place. Transcription factors are proteins that bind to specific DNA sequences near a gene's promoter; activators increase transcription while repressors decrease it. Enhancers are DNA regions, sometimes thousands of bases away from the gene, that can boost transcription when bound by activator proteins. At the epigenetic level, chemical modifications such as DNA methylation (adding methyl groups to cytosine bases) and histone acetylation (adding acetyl groups to histone proteins) alter how tightly DNA is packaged, making genes more or less accessible to RNA polymerase.

Worked Example: From DNA to Amino Acid Sequence

Let's walk through the process of gene expression step by step, starting with a short DNA template strand and determining the amino acid sequence of the resulting polypeptide.

Determining the amino acid sequence from a DNA template strand
1
Step 1 — Identify the DNA Template StrandYou are given the following DNA template strand (read 3′ → 5′ by RNA polymerase): 3′ — TAC GGA CTC AAT ATT — 5′
Template strand: 3′ TAC GGA CTC AAT ATT 5′
2
Step 2 — Transcribe DNA into mRNAApply the base-pairing rules for transcription: A → U, T → A, G → C, C → G. RNA polymerase reads the template 3′ → 5′ and builds mRNA 5′ → 3′. T → A, A → U, C → G | G → C, G → C, A → U | C → G, T → A, C → G | A → U, A → U, T → A | A → U, T → A, T → A
mRNA: 5′ — AUG CCU GAG UUA UAA — 3′
3
Step 3 — Identify the CodonsBreak the mRNA into triplets (codons) starting from the start codon AUG: AUG | CCU | GAG | UUA | UAA
Five codons identified: AUG, CCU, GAG, UUA, UAA
4
Step 4 — Use the Codon Table to Find Amino AcidsLook up each codon in the genetic code table: AUG → Methionine (Met) — also the start codon CCU → Proline (Pro) GAG → Glutamic acid (Glu) UUA → Leucine (Leu) UAA → Stop codon (translation terminates)
Amino acid sequence: Met – Pro – Glu – Leu
5
Step 5 — Interpret the ResultThe polypeptide produced from this template strand is four amino acids long. Translation began at the start codon (AUG) and ended when the ribosome reached the stop codon (UAA). In a real cell, this short peptide would typically undergo further processing and folding to become functional.
Final polypeptide: Met – Pro – Glu – Leu (4 amino acids)

Prokaryotic vs. Eukaryotic Gene Expression

While the central dogma applies to all living organisms, there are important structural and functional differences between how prokaryotes and eukaryotes carry out gene expression. Understanding these differences is a common IB exam topic and will deepen your understanding of cellular organization.

Key differences in gene expression between prokaryotes and eukaryotes
FeatureProkaryotesEukaryotes
LocationCytoplasm (no nucleus)Transcription in nucleus; translation in cytoplasm
CouplingTranscription and translation occur simultaneouslySeparated by nuclear envelope; sequential
RNA ProcessingMinimal; no introns, no splicing5′ capping, poly-A tail, intron splicing
mRNA StructureOften polycistronic (multiple genes per mRNA)Monocistronic (one gene per mRNA)
Gene RegulationOperons (e.g., lac operon); mainly transcriptionalMultiple levels; epigenetics, enhancers, miRNA
Ribosomes70S (50S + 30S subunits)80S (60S + 40S subunits)
KEY TAKEAWAY
The most striking difference is coupling. In bacteria, ribosomes can begin translating an mRNA molecule even while RNA polymerase is still transcribing it — like a factory worker assembling parts as they roll off the production line. In eukaryotes, the nuclear envelope enforces a strict separation, more like a warehouse system where raw materials (pre-mRNA) must be inspected and packaged before being shipped to the assembly floor (cytoplasm).

Connections to Advanced Topics

Gene expression is not just a topic for introductory biology — it connects directly to cutting-edge research in medicine, biotechnology, and evolution. Understanding the basics you've learned here prepares you for more advanced concepts that appear in higher-level IB Biology and university courses.

How gene expression connects to advanced biology and medicine
Concept in This LessonAdvanced Extension
Transcription factors regulate gene expressionMutations in transcription factor genes can lead to cancer (oncogenes and tumor suppressors)
Alternative splicing produces different mRNAs from one geneThe human genome has ~20,000 genes but produces >100,000 different proteins through alternative splicing
Epigenetic modifications (methylation, acetylation)Epigenetic changes can be inherited across generations without altering DNA sequence (transgenerational epigenetics)
The genetic code is (nearly) universalCRISPR-Cas9 gene editing exploits the universality of the code to modify genes in any organism
mRNA carries genetic instructions to the ribosomemRNA vaccines (e.g., COVID-19) deliver synthetic mRNA so cells produce viral proteins to trigger immunity

These connections illustrate why gene expression is one of the most consequential topics in all of biology. The same mechanisms that allow your body to develop from a single fertilized egg also explain why cancers form, how vaccines work, and why genetic engineering is possible. As you continue studying biology, every new topic — from evolution to immunology to ecology — will circle back to the fundamental question of which genes are expressed, where, and when.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain why a muscle cell and a nerve cell in the same organism contain identical DNA but have different structures and functions. Use the term "gene expression" in your answer.
PROBLEM 2BASIC CALCULATION
A DNA template strand has the sequence: 3′ — TAC AAG GCA ACT — 5′. Write the corresponding mRNA sequence and identify the amino acids encoded (use a codon table).
PROBLEM 3INTERMEDIATE
A pre-mRNA molecule contains 5 exons and 4 introns. After RNA processing, if exon 3 is skipped through alternative splicing, how does this affect the final protein? Describe the biological significance of alternative splicing.
PROBLEM 4APPLIED
Certain antibiotics, such as chloramphenicol, work by binding to the 50S subunit of prokaryotic ribosomes and blocking translation. Explain why these antibiotics can kill bacteria without harming human cells, and predict what would happen to bacterial protein synthesis when the drug is present.
PROBLEM 5CRITICAL THINKING
A researcher observes that a specific gene is being actively transcribed in a liver cell (confirmed by high levels of pre-mRNA), yet no corresponding protein is detected. Propose three different levels of regulation that could explain this observation, and describe a mechanism at each level.

Summary: Understand Gene Expression

Gene expression is the process by which information in DNA is used to synthesize functional products, primarily proteins. The process follows the central dogma: DNA is transcribed into mRNA by RNA polymerase, the pre-mRNA is processed (5′ capping, splicing, poly-A tailing) in eukaryotes, and the mature mRNA is translated at ribosomes using tRNA molecules that match codons with specific amino acids. The genetic code is universal, degenerate, and non-overlapping.

Cells regulate gene expression at multiple levels — epigenetic, transcriptional, post-transcriptional, translational, and post-translational — to ensure the right proteins are produced in the right cells at the right time. Key differences between prokaryotic and eukaryotic gene expression include RNA processing, ribosome size, and the spatial separation of transcription and translation. These mechanisms underpin modern applications from mRNA vaccines to CRISPR gene editing.

Varsity Tutors • IB Biology • Understand Gene Expression