HIGH SCHOOL BIOLOGY (NEXT GENERATION SCIENCE STANDARDS) • MOLECULES TO ORGANISMS: STRUCTURES AND PROCESSES

Explain how DNA base sequences encode genetic information.

Discover how a four-letter molecular alphabet directs the construction of every protein in your body.

Historical Context & Motivation

For most of human history, the mechanism of heredity remained a mystery. Farmers and breeders understood that offspring resemble their parents, but no one could explain how traits pass from one generation to the next at the molecular level. The search for the molecule responsible for inheritance took more than a century of experiments, controversies, and breakthroughs. Each discovery built on the last, gradually revealing that a single molecule — deoxyribonucleic acid (DNA) — carries the instructions for building and operating every living organism.

1869
Miescher Isolates "Nuclein"
Friedrich Miescher extracted a phosphorus-rich substance from white blood cells in used surgical bandages. He named this material "nuclein" because it came from the cell nucleus. This was the first chemical isolation of what we now call DNA.
1944
Avery–MacLeod–McCarty Experiment
Oswald Avery and colleagues demonstrated that DNA, not protein, is the "transforming principle" that transfers genetic information between bacteria. This experiment provided critical evidence that DNA is the molecule of heredity.
1950
Chargaff's Rules
Erwin Chargaff discovered that in any DNA sample, the amount of adenine equals thymine and the amount of cytosine equals guanine. These base-pairing ratios later became essential clues for determining the structure of DNA.
1953
Watson, Crick, Franklin & Wilkins
James Watson and Francis Crick proposed the double-helix model of DNA, built upon Rosalind Franklin's X-ray crystallography data and Maurice Wilkins's structural work. Their model explained how DNA stores and copies genetic information through complementary base pairing.
1961
Cracking the Genetic Code
Marshall Nirenberg and Heinrich Matthaei deciphered the first codon, showing that the three-base sequence UUU codes for the amino acid phenylalanine. By 1966 all 64 codons had been assigned, completing the genetic code.

These discoveries raised a fundamental question: how does the specific sequence of chemical units in DNA encode the information needed to build thousands of different proteins? This lesson answers that question by tracing the flow of genetic information from the base sequence of DNA to the amino acid sequence of a protein — a process often summarized as the central dogma of molecular biology: DNA → RNA → protein.

Core Principles of the Genetic Code

DNA encodes genetic information through the specific sequence of four chemical subunits called nucleotides. Each nucleotide contains a sugar (deoxyribose), a phosphate group, and one of four nitrogenous bases: adenine (A), thymine (T), cytosine (C), and guanine (G). The order of these bases along a DNA strand forms the language of life. Just as the 26 letters of the English alphabet can be arranged to spell millions of different words, the four DNA bases can be arranged in an essentially limitless number of sequences, each carrying different genetic instructions.

1

Complementary Base Pairing

Adenine always pairs with thymine (A–T) through two hydrogen bonds, and cytosine always pairs with guanine (C–G) through three hydrogen bonds. This specificity ensures accurate DNA replication and transcription.
2

Triplet Codons

Genetic information is read in groups of three bases called codons. Each codon specifies one amino acid (or a stop signal). With four bases and three positions, there are 4³ = 64 possible codons.
3

The Central Dogma

Information flows from DNA to RNA through transcription, and from RNA to protein through translation. This one-directional flow (DNA → RNA → protein) is the central dogma of molecular biology.
4

Universality of the Code

Nearly all organisms — from bacteria to humans — use the same genetic code. The codon AUG means methionine in almost every species, providing strong evidence for the common ancestry of life.
5

Degeneracy (Redundancy)

Because there are 64 codons but only 20 amino acids (plus stop signals), most amino acids are specified by more than one codon. This redundancy helps buffer organisms against the effects of some point mutations.
KEY TAKEAWAY
Think of DNA as a digital storage system. Just as computers encode all information using only two digits (0 and 1), DNA encodes all biological information using only four chemical letters (A, T, C, G). The specific order of these letters determines which proteins a cell produces, much like the specific sequence of 0s and 1s determines whether a computer file is a photo, a song, or a program.

Visual Explanation: DNA Structure and Base Pairing

This diagram shows a segment of DNA "unzipped" to reveal its two complementary strands. The blue sugar-phosphate backbone of Strand 1 runs 5′ → 3′ (left to right), while the purple backbone of Strand 2 runs antiparallel (3′ → 5′). Bases on opposite strands are connected by hydrogen bonds: A–T pairs share two, while C–G pairs share three, making C–G bonds stronger.

The diagram illustrates several key features of DNA structure. First, notice that the two strands run in opposite directions — they are antiparallel. The 5′ end of one strand aligns with the 3′ end of the other. Second, the bases always pair in a specific way: A with T, and C with G. This complementary base pairing means that if you know the sequence of one strand, you can predict the sequence of the other. Third, the bases project inward from the backbone, forming the "rungs" of the DNA ladder, while the sugar-phosphate backbone forms the "rails." It is the sequence of these inner bases — not the backbone — that carries genetic information.

From DNA to Protein: The Central Dogma

The information stored in DNA directs the synthesis of proteins through a two-step process. In the first step, called transcription, the enzyme RNA polymerase reads one strand of DNA (the template strand) and builds a complementary strand of messenger RNA (mRNA). RNA uses the base uracil (U) in place of thymine, so wherever the DNA template has an A, the mRNA will have a U. RNA polymerase reads the template strand in the 3′ → 5′ direction and synthesizes the mRNA in the 5′ → 3′ direction.

In the second step, called translation, the ribosome reads the mRNA three bases at a time. Each group of three mRNA bases is a codon. Transfer RNA (tRNA) molecules carry amino acids to the ribosome, matching each codon with its corresponding amino acid through an anticodon — a complementary three-base sequence on the tRNA. As the ribosome moves along the mRNA, amino acids are linked together by peptide bonds, forming a polypeptide chain that folds into a functional protein.

This diagram traces the flow of information from a DNA template strand through transcription into mRNA and then through translation into a polypeptide. Notice how each three-base codon on the mRNA corresponds to one amino acid. The start codon AUG begins every polypeptide with methionine, and translation halts at the first stop codon (here, UAG).

Follow the information flow in the diagram from top to bottom. The DNA template strand is read 3′ → 5′ by RNA polymerase, and a complementary mRNA strand is produced in the 5′ → 3′ direction. Each template base is matched by its RNA complement: T → A, A → U, C → G, and G → C. The resulting mRNA carries the genetic message from the nucleus to the ribosome in the cytoplasm. During translation, the ribosome reads mRNA codons in order from the start codon (AUG) to the first stop codon, linking amino acids into a polypeptide chain. The specific sequence of amino acids determines how the protein folds and what function it performs.

Reading the Genetic Code: The Codon Table

With 64 possible three-base codons and only 20 amino acids, the genetic code is degenerate — meaning that most amino acids are specified by more than one codon. This redundancy is not random; it tends to occur at the third base of the codon, which scientists call the wobble position. The table below shows selected codons and their corresponding amino acids. In practice, scientists use a standard codon chart that lists all 64 possibilities organized by the first, second, and third mRNA bases.

Selected mRNA codons and their corresponding amino acids
mRNA CodonAmino AcidRole / Notes
AUGMethionine (Met)Start codon — initiates translation
UUU, UUCPhenylalanine (Phe)Two codons — differ at the wobble position
GCU, GCC, GCA, GCGAlanine (Ala)Four codons — highly redundant
GAG, GAAGlutamic acid (Glu)Hydrophilic — important in sickle cell comparison
GUG, GUU, GUC, GUAValine (Val)Hydrophobic — replaces Glu in sickle cell hemoglobin
UAG, UAA, UGA— (no amino acid)Stop codons — signal the ribosome to release the polypeptide

Notice that a single base change in a codon can change the amino acid or have no effect at all. For example, changing GAG (glutamic acid) to GUG (valine) swaps a hydrophilic amino acid for a hydrophobic one — this is the exact mutation responsible for sickle cell disease. In contrast, changing GCU to GCC still codes for alanine, so the protein is unaffected. This relationship between base sequence and amino acid identity is the foundation of how DNA encodes information.

KEY TAKEAWAY
The codon table is like a translation dictionary between two languages. Just as a Spanish-English dictionary converts words from one language to another, the codon table converts three-letter RNA "words" into their amino acid equivalents. Some RNA "words" are synonyms — different codons can code for the same amino acid, just as "happy" and "glad" have the same meaning in English.

Worked Example: From DNA to Amino Acid Sequence

Let's work through a complete example of converting a DNA template strand into an amino acid sequence. This exercise practices the skill of tracing the central dogma from start to finish.

🧬 Problem
A segment of a DNA template strand has the sequence: 3′ — TAC AAA GGA ACT — 5′. Determine the mRNA sequence and the amino acid sequence it encodes.
From DNA Template to Amino Acid Sequence
1
Step 1 — Identify the Template Strand DirectionThe DNA template strand is given as 3′ — TAC AAA GGA ACT — 5′. RNA polymerase reads this strand in the 3′ → 5′ direction (left to right as written here) and synthesizes the mRNA in the 5′ → 3′ direction.
Reading direction: 3′ → 5′ (TAC → AAA → GGA → ACT)
2
Step 2 — Transcribe DNA to mRNAApply the base-pairing rules for transcription: A on the template pairs with U on the mRNA; T pairs with A; C pairs with G; and G pairs with C. Reading the template 3′ → 5′: T→A, A→U, C→G (first codon: AUG). A→U, A→U, A→U (second codon: UUU). G→C, G→C, A→U (third codon: CCU). A→U, C→G, T→A (fourth codon: UGA).
mRNA: 5′ — AUG UUU CCU UGA — 3′
3
Step 3 — Identify Codons and Reading FrameThe mRNA is already grouped into codons of three bases each: AUG, UUU, CCU, and UGA. Translation begins at the start codon AUG and reads in the 5′ → 3′ direction.
Codons: AUG | UUU | CCU | UGA
4
Step 4 — Translate Each CodonUsing the standard genetic code: AUG = Methionine (Met), the start codon. UUU = Phenylalanine (Phe). CCU = Proline (Pro). UGA = Stop — the ribosome releases the polypeptide.
Amino acid sequence: Met – Phe – Pro (STOP)
5
Step 5 — Verify and InterpretThe resulting polypeptide is three amino acids long. The order of amino acids — methionine, phenylalanine, proline — is entirely determined by the order of base triplets in the original DNA template strand. A change to any single base could alter the amino acid sequence and potentially change the protein's structure and function.
Final polypeptide: Met–Phe–Pro (3 amino acids)

When the Code Changes: Types of Mutations

Because protein function depends on the precise sequence of amino acids, any change to the DNA base sequence has the potential to alter the resulting protein. These changes, called mutations, can range from harmless to lethal depending on where they occur and what effect they have on protein structure. Understanding mutation types reinforces how the base sequence encodes information — if it did not matter, mutations would have no consequences.

Common types of DNA mutations and their effects on protein structure
Mutation TypeWhat Happens to DNAEffect on Protein
Silent (synonymous)One base is substituted, but the new codon specifies the same amino acid (e.g., GCU → GCC both code for alanine).No change — the protein is identical. Redundancy in the genetic code provides a buffer.
MissenseOne base is substituted, changing the codon to one that specifies a different amino acid (e.g., GAG → GUG changes Glu to Val).One amino acid changes. May disrupt protein folding and function, or may have little effect depending on location and chemical properties.
NonsenseOne base is substituted, creating a premature stop codon (e.g., UAC → UAG changes Tyr codon to Stop).The protein is truncated (shortened). Often nonfunctional because the full amino acid sequence is needed for proper folding.
Frameshift (insertion/deletion)One or more bases are inserted into or deleted from the sequence (not in multiples of three).All codons downstream of the mutation are altered. Usually produces a completely nonfunctional protein.
KEY TAKEAWAY
Think of a frameshift mutation like removing one letter from the beginning of a sentence and then re-reading it in groups of three. The sentence "THE CAT ATE THE RAT" becomes "_HE CAT ATE THE RAT" — if we shift the reading frame, it reads "HEC ATA TET HER AT" — complete nonsense. Inserting or deleting a single base in DNA has the same catastrophic effect on the codon reading frame, scrambling every amino acid from that point forward.

Beyond the Basics: Gene Expression and Regulation

The central dogma describes the flow of information from DNA to protein, but cells do not simply transcribe all genes all the time. Gene expression is the process by which a specific gene's DNA sequence is read and used to direct protein synthesis. Cells regulate which genes are expressed, when they are expressed, and how much protein is produced. This regulation explains how a single genome can produce hundreds of different cell types — a muscle cell and a neuron contain the same DNA but express very different sets of genes.

FeatureBasic Central Dogma (This Lesson)Advanced Gene Expression (Future Topics)
ScopeDNA → mRNA → protein for a single geneRegulation of thousands of genes across different tissues and developmental stages
mRNA processingmRNA is a direct copy of the coding sequencePre-mRNA is processed: introns are removed, a 5′ cap and poly-A tail are added (eukaryotes)
Regulatory elementsPromoter region signals RNA polymerase to begin transcriptionEnhancers, silencers, transcription factors, and epigenetic modifications control gene activity
Mutation impactChanges in coding sequence alter amino acidsMutations in regulatory regions can change when and how much protein is made without altering the protein itself

As you continue in biology, you will encounter topics like epigenetics, RNA splicing, and gene regulation networks that build on the fundamental concepts covered here. Every one of those advanced topics rests on the principle that the linear sequence of bases in DNA is the primary source of genetic information. Mastering the central dogma gives you the foundation for understanding all of modern molecular biology.

Practice Problems

📐 NGSS Integration Notes
These problems integrate NGSS three-dimensional learning. Each problem engages at least one Science and Engineering Practice (SEP) and one Crosscutting Concept (CCC), in addition to Disciplinary Core Ideas (HS-LS1-1). Look for the tags before each question.
PROBLEM 1CONCEPTUAL
SEP: Constructing Explanations | CCC: Structure and Function DNA molecules in all living organisms are composed of the same four nucleotide bases (A, T, C, G). Which of the following best explains how DNA can encode such a vast diversity of genetic information using only four bases? A) Different organisms have different numbers of total base pairs in their genomes. B) The four bases can be arranged in an enormous number of different sequences, and it is the specific order that encodes information. C) Some organisms use additional bases that have not yet been discovered. D) The sugar-phosphate backbone varies between organisms, providing additional information.
PROBLEM 2BASIC
SEP: Developing and Using Models | CCC: Structure and Function A DNA template strand has the sequence 3′ — TAC CGG ATC — 5′. What is the sequence of the mRNA produced during transcription? A) 5′ — AUG GCC UAG — 3′ B) 5′ — AUG GCC AUC — 3′ C) 5′ — UAC CGG AUC — 3′ D) 5′ — AUG CGG AUC — 3′
PROBLEM 3INTERMEDIATE
SEP: Analyzing and Interpreting Data | CCC: Cause and Effect A researcher compares two versions of a gene. The original DNA coding strand reads 5′ — ATG GAG CTT — 3′. A mutant version reads 5′ — ATG GUG CTT — 3′. Wait — DNA does not contain uracil. Assuming this is a typo and the mutant coding strand actually reads 5′ — ATG GTG CTT — 3′, what type of mutation has occurred and what is its likely effect? A) Silent mutation — the same amino acid is produced. B) Missense mutation — one amino acid is changed, which may alter protein function. C) Nonsense mutation — a premature stop codon is introduced. D) Frameshift mutation — the reading frame is shifted, altering all downstream codons.
PROBLEM 4APPLIED
SEP: Constructing Explanations and Designing Solutions | CCC: Cause and Effect Sickle cell disease results from a single base substitution in the beta-globin gene that replaces glutamic acid (a hydrophilic amino acid) with valine (a hydrophobic amino acid) at position 6 of the hemoglobin protein. Using your understanding of how DNA base sequences encode information, which of the following best explains why this single amino acid change has such a severe effect on red blood cell shape? A) The valine substitution prevents the hemoglobin protein from being translated at all. B) The hydrophobic valine creates a sticky patch on deoxygenated hemoglobin that promotes polymerization into long, rigid protein fibers, which distort the flexible red blood cell membrane into a sickle shape. C) The mutation changes every amino acid downstream of position 6, producing a completely different protein. D) The mutation removes the start codon, so no hemoglobin protein can be made.
PROBLEM 5CRITICAL THINKING
SEP: Engaging in Argument from Evidence | CCC: Cause and Effect, Structure and Function Two patients each have a different point mutation (single base substitution) in a gene that codes for a 300-amino-acid enzyme. Patient X has a missense mutation at codon 12 (near the beginning of the protein), and Patient Y has a missense mutation at codon 299 (near the end of the protein). The mRNA coding region for this protein is 900 nucleotides long (300 amino acids × 3 nucleotides per codon). Both patients produce full-length proteins, but Patient X's enzyme has only 5% of normal activity while Patient Y's enzyme retains 85% of normal activity. Which of the following best explains this difference? A) Patient X's mutation occurred in a region of the protein critical for its three-dimensional folding or active site, whereas Patient Y's mutation occurred in a region less important for overall structure and catalytic function. B) Patient X's mutation is closer to the start codon, so it affects more of the downstream amino acid sequence. C) Patient Y's mutation must be a silent mutation because it is near the end of the gene. D) Patient X's protein is shorter than Patient Y's protein because early mutations always cause truncation.

Summary: How DNA Base Sequences Encode Genetic Information

DNA stores genetic information in the specific sequence of its four nitrogenous bases — adenine, thymine, cytosine, and guanine. The two strands of the double helix are held together by complementary base pairing (A–T and C–G). During transcription, RNA polymerase reads the DNA template strand (3′ → 5′) and produces a complementary mRNA strand (5′ → 3′), substituting uracil for thymine.

During translation, ribosomes read the mRNA in three-base units called codons. Each codon specifies one of 20 amino acids or a stop signal, creating the genetic code (4³ = 64 codons). The code is degenerate — multiple codons can specify the same amino acid, buffering against some mutations. The order of amino acids determines how a protein folds and functions, meaning that the base sequence of DNA ultimately dictates the structure and function of every protein in an organism. This principle — DNA → RNA → Protein — is the central dogma of molecular biology.

Varsity Tutors • High School Biology (Next Generation Science Standards) • Explain how DNA base sequences encode genetic information.