Historical Context & Motivation
For centuries, scientists knew that living organisms were made of complex substances, but they had no idea how cells actually manufactured them. By the mid-twentieth century, researchers understood that DNA carried hereditary information, yet the mechanism by which this information was converted into functional proteins remained a mystery. Unravelling this process — known as protein synthesis — became one of the greatest achievements in molecular biology and reshaped our understanding of how genes control the traits of every living organism.
These discoveries answered a fundamental question: how does a cell convert a sequence of nucleotide bases in DNA into a precise sequence of amino acids in a protein? The answer involves two major stages — transcription and translation — each with its own molecular machinery. Understanding these stages is essential for the IB Biology curriculum and for grasping how mutations, gene expression, and biotechnology work.
Core Principles of Protein Synthesis
Protein synthesis is the process by which cells use the instructions encoded in DNA to build specific proteins. These proteins serve as enzymes, structural components, hormones, antibodies, and much more. The process can be broken down into a few foundational ideas that connect genetics to cell function.
The Central Dogma
Transcription
The Genetic Code
Translation
Post-Translational Modification
Visualising Transcription
Transcription is the first major stage of protein synthesis, occurring inside the nucleus of eukaryotic cells. The diagram below shows how RNA polymerase unwinds a section of the DNA double helix, reads the template (antisense) strand in the 3' → 5' direction, and assembles a complementary mRNA strand in the 5' → 3' direction. Notice that in mRNA, the base uracil (U) replaces thymine (T).
During transcription, the enzyme RNA polymerase binds to a specific region of DNA called the promoter. It then unwinds a small section of the double helix and begins reading the template strand, adding complementary RNA nucleotides one by one. The base-pairing rules are similar to DNA replication — adenine pairs with uracil (A-U), and cytosine pairs with guanine (C-G). When RNA polymerase reaches a terminator sequence, it detaches, and the newly formed pre-mRNA is released. In eukaryotes, this pre-mRNA undergoes processing — including the addition of a 5' cap, a poly-A tail, and the removal of introns through splicing — before becoming mature mRNA that exits the nucleus.
The Mechanism of Translation
Once the mature mRNA exits the nucleus through a nuclear pore, it enters the cytoplasm, where translation takes place. Translation is the process by which ribosomes decode the mRNA codons and assemble amino acids into a polypeptide chain. The ribosome has two subunits — a small subunit that reads the mRNA and a large subunit that catalyses peptide bond formation. Translation occurs in three main phases: initiation, elongation, and termination.
Initiation
The small ribosomal subunit binds to the 5' end of the mRNA and scans along until it finds the start codon (AUG), which codes for the amino acid methionine. A special initiator tRNA with the anticodon UAC binds to this start codon, and the large subunit then joins, forming the complete ribosome. The ribosome has three binding sites for tRNA: the A site (aminoacyl), the P site (peptidyl), and the E site (exit).
Elongation
A tRNA carrying the next amino acid enters the A site, where its anticodon pairs with the mRNA codon through complementary base pairing. The ribosome then catalyses the formation of a peptide bond between the amino acid at the P site and the amino acid at the A site. The ribosome shifts (translocates) one codon along the mRNA in the 5' → 3' direction. The empty tRNA exits through the E site, and the process repeats, adding one amino acid at a time. This cycle occurs rapidly — ribosomes can add about 15–20 amino acids per second in eukaryotes.
Termination
Translation continues until the ribosome encounters one of the three stop codons (UAA, UAG, or UGA) on the mRNA. No tRNA has a matching anticodon for these codons. Instead, a release factor protein binds to the A site, causing the completed polypeptide chain to be released. The ribosome then disassembles into its two subunits, ready to be used again.
The Genetic Code & Codon Table
The genetic code is the set of rules that translates the four-letter nucleotide alphabet (A, U, G, C in mRNA) into the twenty-letter amino acid alphabet of proteins. Because there are 4 possible bases and codons are read in groups of three, there are 4³ = 64 possible codons. Since only 20 amino acids exist (plus start and stop signals), the code is described as degenerate — meaning most amino acids are encoded by more than one codon. The code is also universal, used by nearly all organisms from bacteria to humans, which is powerful evidence for a common ancestor.
When reading the codon table on an IB exam, you will typically be given the mRNA sequence and need to identify the corresponding amino acids. Remember to always start from the AUG start codon and read each subsequent group of three bases without skipping or overlapping. The reading frame is set by the start codon — if you shift by even one base, every downstream amino acid changes, which is why frameshift mutations (insertions or deletions) are so damaging.
Worked Example: From DNA to Polypeptide
Let's walk through a complete example of protein synthesis, starting with a DNA template strand and ending with a polypeptide. This is a common IB exam question type.
3' — T A C A A G G C A T T G A C T — 5'
This is the strand that RNA polymerase reads during transcription.Template: 3' — T A C A A G G C A T T G A C T — 5'
mRNA: 5' — A U G U U C C G U A A C U G A — 3'AUG | UUC | CGU | AAC | UGATranscription vs. Translation
Students often confuse transcription and translation because both involve nucleic acids and base pairing. The table below highlights the key differences and similarities between these two stages of protein synthesis to help you keep them distinct.
| Feature | Transcription | Translation |
|---|---|---|
| Location | Nucleus (eukaryotes) | Cytoplasm (ribosomes) |
| Template | DNA template strand (3'→5') | mRNA (5'→3') |
| Product | mRNA (single-stranded RNA) | Polypeptide chain (protein) |
| Key enzyme | RNA polymerase | Ribosome (ribozyme) |
| Monomers used | RNA nucleotides (A, U, G, C) | Amino acids (20 types) |
| Base pairing | DNA → RNA (T→A, A→U, C→G, G→C) | mRNA codon → tRNA anticodon |
| Start signal | Promoter sequence on DNA | AUG start codon on mRNA |
| Stop signal | Terminator sequence on DNA | UAA, UAG, or UGA stop codons |
Connection to Gene Expression & Mutations
Understanding protein synthesis opens the door to many advanced topics in the IB Biology syllabus, including gene expression regulation and mutations. Cells do not transcribe every gene at all times; instead, they regulate which genes are turned on or off in response to signals, ensuring that the right proteins are made in the right cells at the right time. This regulation can occur at the transcription level (e.g., through transcription factors binding to promoters) or at the translation level.
| Concept | Standard Protein Synthesis | Advanced Applications |
|---|---|---|
| Gene regulation | All genes are potentially transcribed | Only specific genes are expressed in specific cell types (e.g., insulin gene in β-cells only) |
| Point mutations | Wild-type base sequence produces normal protein | A single base substitution may cause a silent, missense, or nonsense mutation (e.g., sickle cell: GAG→GUG) |
| Frameshift mutations | Reading frame set by AUG start codon | Insertions or deletions shift all downstream codons, usually producing a non-functional protein |
| Biotechnology | Natural protein synthesis in cells | Genetic engineering uses the same machinery (e.g., mRNA vaccines instruct ribosomes to make viral spike proteins) |
A particularly famous example is sickle cell disease, caused by a single base substitution in the gene for the β-globin chain of haemoglobin. The normal codon GAG (glutamic acid) becomes GUG (valine), changing just one amino acid out of 146. This tiny change causes the haemoglobin molecules to polymerise under low-oxygen conditions, distorting red blood cells into a sickle shape. This example powerfully illustrates how protein synthesis errors can cascade into significant biological consequences.
Practice Problems
Protein Synthesis — Summary
Protein synthesis is the two-stage process by which cells convert DNA instructions into functional proteins, following the central dogma: DNA → RNA → Protein. In transcription, RNA polymerase reads the DNA template strand (3'→5') and builds a complementary mRNA strand (5'→3') in the nucleus. The mRNA is then processed and exported to the cytoplasm.
In translation, ribosomes read the mRNA in groups of three bases called codons. Each codon is matched by a tRNA anticodon carrying the corresponding amino acid. Translation begins at the AUG start codon (methionine) and ends at one of three stop codons (UAA, UAG, UGA). The resulting polypeptide folds into a three-dimensional protein whose function depends on its precise amino acid sequence. The genetic code is universal, degenerate, non-overlapping, and unambiguous — properties that ensure accuracy while buffering against mutations.