IB BIOLOGY • CONTINUITY AND CHANGE

Understand Protein Synthesis

How cells read DNA instructions to build the proteins that drive every life process.

Historical Context & Motivation

For centuries, scientists knew that living organisms were made of complex substances, but they had no idea how cells actually manufactured them. By the mid-twentieth century, researchers understood that DNA carried hereditary information, yet the mechanism by which this information was converted into functional proteins remained a mystery. Unravelling this process — known as protein synthesis — became one of the greatest achievements in molecular biology and reshaped our understanding of how genes control the traits of every living organism.

1953
DNA Structure Solved
James Watson and Francis Crick, building on Rosalind Franklin's X-ray crystallography data, proposed the double-helix model of DNA, revealing how genetic information could be stored and copied.
1958
The Central Dogma
Francis Crick proposed the central dogma of molecular biology: information flows from DNA → RNA → Protein. This framework guided decades of research.
1961
Messenger RNA Identified
François Jacob and Jacques Monod identified mRNA as the intermediate molecule that carries the DNA message to the ribosome, where proteins are assembled.
1966
The Genetic Code Cracked
Marshall Nirenberg and Har Gobind Khorana deciphered the complete genetic code, showing which three-nucleotide sequences (codons) correspond to each amino acid.

These discoveries answered a fundamental question: how does a cell convert a sequence of nucleotide bases in DNA into a precise sequence of amino acids in a protein? The answer involves two major stages — transcription and translation — each with its own molecular machinery. Understanding these stages is essential for the IB Biology curriculum and for grasping how mutations, gene expression, and biotechnology work.

Core Principles of Protein Synthesis

Protein synthesis is the process by which cells use the instructions encoded in DNA to build specific proteins. These proteins serve as enzymes, structural components, hormones, antibodies, and much more. The process can be broken down into a few foundational ideas that connect genetics to cell function.

1

The Central Dogma

Genetic information flows in one primary direction: DNA → RNA → Protein. DNA is the master blueprint stored safely in the nucleus, RNA is the disposable copy that travels to the ribosome, and the protein is the final functional product.
2

Transcription

In transcription, the enzyme RNA polymerase reads a gene on the template strand of DNA and synthesises a complementary messenger RNA (mRNA) molecule. This occurs in the nucleus of eukaryotic cells.
3

The Genetic Code

The mRNA is read in groups of three bases called codons. Each codon specifies one amino acid (or a stop signal). The code is universal, degenerate (multiple codons per amino acid), and non-overlapping.
4

Translation

During translation, ribosomes read the mRNA codons and, with the help of transfer RNA (tRNA) molecules, assemble amino acids into a polypeptide chain. This occurs in the cytoplasm at free or rough-ER-bound ribosomes.
5

Post-Translational Modification

The new polypeptide folds into its three-dimensional shape and may be modified (e.g., by adding sugar groups or being cleaved) to become a fully functional protein. Misfolded proteins can lead to disease.
KEY TAKEAWAY
Think of protein synthesis like building a house from blueprints. DNA is the master blueprint that stays locked in the architect's office (the nucleus). mRNA is a photocopy of just one page of the blueprint that gets sent to the construction site (the ribosome). tRNA molecules are the delivery trucks, each carrying a specific building material (amino acid) that matches a line item on the photocopy. The workers at the construction site (the ribosome) read each instruction and attach the materials one by one until the structure (protein) is complete.

Visualising Transcription

Transcription is the first major stage of protein synthesis, occurring inside the nucleus of eukaryotic cells. The diagram below shows how RNA polymerase unwinds a section of the DNA double helix, reads the template (antisense) strand in the 3' → 5' direction, and assembles a complementary mRNA strand in the 5' → 3' direction. Notice that in mRNA, the base uracil (U) replaces thymine (T).

The diagram shows RNA polymerase (violet/pink ellipse) moving along the template strand (blue, 3'→5'), synthesising a complementary mRNA strand (pink, 5'→3'). Notice that uracil (U) appears in mRNA wherever thymine (T) would appear in DNA. The coding strand (green) has the same sequence as the mRNA, except with T instead of U.

During transcription, the enzyme RNA polymerase binds to a specific region of DNA called the promoter. It then unwinds a small section of the double helix and begins reading the template strand, adding complementary RNA nucleotides one by one. The base-pairing rules are similar to DNA replication — adenine pairs with uracil (A-U), and cytosine pairs with guanine (C-G). When RNA polymerase reaches a terminator sequence, it detaches, and the newly formed pre-mRNA is released. In eukaryotes, this pre-mRNA undergoes processing — including the addition of a 5' cap, a poly-A tail, and the removal of introns through splicing — before becoming mature mRNA that exits the nucleus.

The Mechanism of Translation

Once the mature mRNA exits the nucleus through a nuclear pore, it enters the cytoplasm, where translation takes place. Translation is the process by which ribosomes decode the mRNA codons and assemble amino acids into a polypeptide chain. The ribosome has two subunits — a small subunit that reads the mRNA and a large subunit that catalyses peptide bond formation. Translation occurs in three main phases: initiation, elongation, and termination.

Initiation

The small ribosomal subunit binds to the 5' end of the mRNA and scans along until it finds the start codon (AUG), which codes for the amino acid methionine. A special initiator tRNA with the anticodon UAC binds to this start codon, and the large subunit then joins, forming the complete ribosome. The ribosome has three binding sites for tRNA: the A site (aminoacyl), the P site (peptidyl), and the E site (exit).

Elongation

A tRNA carrying the next amino acid enters the A site, where its anticodon pairs with the mRNA codon through complementary base pairing. The ribosome then catalyses the formation of a peptide bond between the amino acid at the P site and the amino acid at the A site. The ribosome shifts (translocates) one codon along the mRNA in the 5' → 3' direction. The empty tRNA exits through the E site, and the process repeats, adding one amino acid at a time. This cycle occurs rapidly — ribosomes can add about 15–20 amino acids per second in eukaryotes.

Termination

Translation continues until the ribosome encounters one of the three stop codons (UAA, UAG, or UGA) on the mRNA. No tRNA has a matching anticodon for these codons. Instead, a release factor protein binds to the A site, causing the completed polypeptide chain to be released. The ribosome then disassembles into its two subunits, ready to be used again.

💡 IB Exam Tip
Remember the base-pairing rules for transcription and translation. DNA template to mRNA: A→U, T→A, C→G, G→C. mRNA codon to tRNA anticodon: A→U, U→A, C→G, G→C. A common exam mistake is forgetting that RNA uses uracil instead of thymine.

The Genetic Code & Codon Table

The genetic code is the set of rules that translates the four-letter nucleotide alphabet (A, U, G, C in mRNA) into the twenty-letter amino acid alphabet of proteins. Because there are 4 possible bases and codons are read in groups of three, there are 4³ = 64 possible codons. Since only 20 amino acids exist (plus start and stop signals), the code is described as degenerate — meaning most amino acids are encoded by more than one codon. The code is also universal, used by nearly all organisms from bacteria to humans, which is powerful evidence for a common ancestor.

A simplified codon table highlighting key codons. The start codon (AUG) always codes for methionine and begins translation. The three stop codons (UAA, UAG, UGA) do not code for any amino acid. Notice that tryptophan (UGG) and methionine (AUG) are the only amino acids with a single codon — all others have two or more.

When reading the codon table on an IB exam, you will typically be given the mRNA sequence and need to identify the corresponding amino acids. Remember to always start from the AUG start codon and read each subsequent group of three bases without skipping or overlapping. The reading frame is set by the start codon — if you shift by even one base, every downstream amino acid changes, which is why frameshift mutations (insertions or deletions) are so damaging.

Worked Example: From DNA to Polypeptide

Let's walk through a complete example of protein synthesis, starting with a DNA template strand and ending with a polypeptide. This is a common IB exam question type.

Transcribe and Translate a Gene Sequence
1
Step 1 — Identify the DNA template strandYou are given the following DNA template strand (read 3'→5'): 3' — T A C A A G G C A T T G A C T — 5' This is the strand that RNA polymerase reads during transcription.
Template strand: 3'-TACAAGGCATTGACT-5'
2
Step 2 — Transcribe to mRNAApply the complementary base-pairing rules, remembering that RNA uses uracil (U) instead of thymine (T). Pair each template base: T→A, A→U, C→G, G→C. Template: 3' — T A C A A G G C A T T G A C T — 5' mRNA: 5' — A U G U U C C G U A A C U G A — 3'
mRNA: 5'-AUGUUCCGUAACUGA-3'
3
Step 3 — Divide mRNA into codonsStarting from the 5' end, group the mRNA sequence into triplets (codons). The first codon should be AUG (the start codon). AUG | UUC | CGU | AAC | UGA
Five codons identified: AUG, UUC, CGU, AAC, UGA
4
Step 4 — Translate each codon using the codon tableUse the mRNA codon table to identify each amino acid: • AUG → Methionine (Met) — also the START signal • UUC → Phenylalanine (Phe) • CGU → Arginine (Arg) • AAC → Asparagine (Asn) • UGA → STOP — translation terminates here
Polypeptide: Met – Phe – Arg – Asn (4 amino acids)
5
Step 5 — Determine the tRNA anticodonsFor each mRNA codon, the corresponding tRNA carries a complementary anticodon: • AUG → tRNA anticodon: UAC • UUC → tRNA anticodon: AAG • CGU → tRNA anticodon: GCA • AAC → tRNA anticodon: UUG (No tRNA binds to the stop codon — a release factor binds instead.)
tRNA anticodons: UAC, AAG, GCA, UUG
🔑 PATTERN TO REMEMBER
The flow is always: DNA template strand → mRNA → tRNA anticodon → amino acid. Each step uses complementary base pairing. If you can reliably apply A-U and C-G pairing (and remember that RNA never contains thymine), you can solve any transcription/translation question on the IB exam.

Transcription vs. Translation

Students often confuse transcription and translation because both involve nucleic acids and base pairing. The table below highlights the key differences and similarities between these two stages of protein synthesis to help you keep them distinct.

Comparison of transcription and translation
FeatureTranscriptionTranslation
LocationNucleus (eukaryotes)Cytoplasm (ribosomes)
TemplateDNA template strand (3'→5')mRNA (5'→3')
ProductmRNA (single-stranded RNA)Polypeptide chain (protein)
Key enzymeRNA polymeraseRibosome (ribozyme)
Monomers usedRNA nucleotides (A, U, G, C)Amino acids (20 types)
Base pairingDNA → RNA (T→A, A→U, C→G, G→C)mRNA codon → tRNA anticodon
Start signalPromoter sequence on DNAAUG start codon on mRNA
Stop signalTerminator sequence on DNAUAA, UAG, or UGA stop codons
KEY TAKEAWAY
A helpful way to remember the difference: Transcription is like rewriting notes from a textbook (DNA) into your own notebook (mRNA) — you're converting information within the same 'language' of nucleic acids. Translation is like translating those notes from one language (nucleotides) into a completely different language (amino acids). That's why it's called translation — there's a genuine change of 'alphabet.'

Connection to Gene Expression & Mutations

Understanding protein synthesis opens the door to many advanced topics in the IB Biology syllabus, including gene expression regulation and mutations. Cells do not transcribe every gene at all times; instead, they regulate which genes are turned on or off in response to signals, ensuring that the right proteins are made in the right cells at the right time. This regulation can occur at the transcription level (e.g., through transcription factors binding to promoters) or at the translation level.

How protein synthesis connects to advanced IB topics
ConceptStandard Protein SynthesisAdvanced Applications
Gene regulationAll genes are potentially transcribedOnly specific genes are expressed in specific cell types (e.g., insulin gene in β-cells only)
Point mutationsWild-type base sequence produces normal proteinA single base substitution may cause a silent, missense, or nonsense mutation (e.g., sickle cell: GAG→GUG)
Frameshift mutationsReading frame set by AUG start codonInsertions or deletions shift all downstream codons, usually producing a non-functional protein
BiotechnologyNatural protein synthesis in cellsGenetic engineering uses the same machinery (e.g., mRNA vaccines instruct ribosomes to make viral spike proteins)

A particularly famous example is sickle cell disease, caused by a single base substitution in the gene for the β-globin chain of haemoglobin. The normal codon GAG (glutamic acid) becomes GUG (valine), changing just one amino acid out of 146. This tiny change causes the haemoglobin molecules to polymerise under low-oxygen conditions, distorting red blood cells into a sickle shape. This example powerfully illustrates how protein synthesis errors can cascade into significant biological consequences.

🔬 Looking Ahead
In higher-level IB Biology and university courses, you'll explore topics like epigenetics (how chemical modifications to DNA affect gene expression without altering the sequence), RNA splicing (how one gene can produce multiple proteins), and CRISPR gene editing (which directly manipulates DNA sequences to alter the proteins a cell produces).

Practice Problems

PROBLEM 1CONCEPTUAL
Explain why the genetic code is described as 'degenerate' but 'unambiguous'. How do these two properties differ, and why is this distinction important for the accuracy of protein synthesis?
PROBLEM 2BASIC CALCULATION
A DNA template strand has the sequence: 3'-TACCCGAAATTT-5'. Determine the mRNA sequence, identify the codons, and use a codon table to find the corresponding amino acid sequence.
PROBLEM 3INTERMEDIATE
An mRNA sequence reads: 5'-AUGGCUUACGAUUGA-3'. (a) How many amino acids will the resulting polypeptide contain? (b) If the third base of the second codon (U in GCU) is mutated to C (making GCC), will the polypeptide change? Explain your answer. (c) If the third base of the second codon is deleted instead, what happens to the rest of the sequence?
PROBLEM 4APPLIED
In sickle cell disease, the sixth codon of the β-globin mRNA changes from GAG to GUG. (a) What amino acid substitution does this cause? (b) Identify the single base change in the DNA template strand that would produce this mRNA mutation. (c) Explain why this single amino acid change has such a dramatic effect on red blood cell shape.
PROBLEM 5CRITICAL THINKING
A researcher discovers that a bacterial gene produces a functional protein of 200 amino acids. She finds that 15% of the codons in the mRNA are for leucine. (a) How many leucine codons are present? (b) Given that leucine has six codons (UUA, UUG, CUU, CUC, CUA, CUG), discuss why the degeneracy of the code for leucine might be advantageous for the organism. (c) If the researcher could somehow replace all leucine codons in the mRNA with the single codon UUA, would the same protein be produced? Justify your answer and consider potential complications.

Protein Synthesis — Summary

Protein synthesis is the two-stage process by which cells convert DNA instructions into functional proteins, following the central dogma: DNA → RNA → Protein. In transcription, RNA polymerase reads the DNA template strand (3'→5') and builds a complementary mRNA strand (5'→3') in the nucleus. The mRNA is then processed and exported to the cytoplasm.

In translation, ribosomes read the mRNA in groups of three bases called codons. Each codon is matched by a tRNA anticodon carrying the corresponding amino acid. Translation begins at the AUG start codon (methionine) and ends at one of three stop codons (UAA, UAG, UGA). The resulting polypeptide folds into a three-dimensional protein whose function depends on its precise amino acid sequence. The genetic code is universal, degenerate, non-overlapping, and unambiguous — properties that ensure accuracy while buffering against mutations.

Varsity Tutors • IB Biology • Understand Protein Synthesis