COLLEGE BIOLOGY • BIOCHEMISTRY FOUNDATIONS

Proteins

The molecular machines that catalyze reactions, transmit signals, and build the structural framework of life.

Historical Context & Motivation

The study of proteins traces its origins to the early nineteenth century, when chemists first recognized that a distinct class of nitrogen-containing organic molecules was essential to all living organisms. The Dutch chemist Gerardus Johannes Mulder, working alongside Jöns Jacob Berzelius, introduced the term protein in 1838—from the Greek proteios, meaning "of the first rank"—signaling the conviction that these substances occupied a central position in biology. Over the subsequent two centuries, breakthroughs in crystallography, sequencing, and structural biology transformed proteins from mysterious albuminous substances into the most thoroughly characterized macromolecules in the cell. Understanding the historical arc of protein science reveals how each advance built upon prior discoveries, ultimately connecting molecular structure to biological function in ways that underpin modern medicine, biotechnology, and molecular engineering.

1838
Naming of Proteins
Mulder and Berzelius coin the term "protein" after discovering that animal and plant tissues contain similar nitrogen-rich organic compounds with characteristic elemental compositions.
1902
The Peptide Bond Hypothesis
Emil Fischer and Franz Hofmeister independently propose that amino acids are linked by peptide bonds in linear chains, establishing the covalent architecture that underlies all protein primary structures.
1951
α-Helix and β-Sheet Discovered
Linus Pauling, Robert Corey, and Herman Branson predict the α-helix and β-pleated sheet using model building and X-ray diffraction data, revealing the regular hydrogen-bonding patterns of secondary structure.
1958
First Protein Structure Solved
John Kendrew determines the three-dimensional structure of myoglobin by X-ray crystallography, earning the Nobel Prize and proving that proteins adopt defined tertiary folds.
2020
AlphaFold Revolutionizes Prediction
DeepMind's AlphaFold2 achieves near-experimental accuracy in predicting protein structures from amino acid sequences, demonstrating that AI can solve the protein folding problem at scale.

A central question has driven protein science for over a century: how does a linear sequence of amino acids encode the precise three-dimensional shape—and thereby the biological function—of a protein? This question, often called the protein folding problem, connects chemistry to biology at the most fundamental level and motivates much of the material we explore in this lesson.

Core Principles & Definitions

Proteins are polypeptides—linear polymers of amino acids joined by covalent peptide bonds—that fold into specific three-dimensional conformations dictated by their sequence. The twenty standard amino acids encoded by the genetic code differ in their R groups (side chains), which range from simple hydrogen atoms in glycine to complex indole rings in tryptophan. These side-chain properties—hydrophobic, hydrophilic, charged, or aromatic—collectively determine how a polypeptide chain folds, what ligands it binds, and what reactions it catalyzes. Grasping four foundational principles provides the conceptual scaffold for understanding all of protein biochemistry.

1

Amino Acid Building Blocks

Each amino acid contains an amino group (−NH₂), a carboxyl group (−COOH), an α-carbon, a hydrogen, and a variable R group. The R group determines polarity, charge, and chemical reactivity, enabling the diversity of protein function.
2

The Peptide Bond

A condensation (dehydration synthesis) reaction links the α-carboxyl of one amino acid to the α-amino of the next, releasing water. The resulting C−N peptide bond is planar and rigid owing to partial double-bond character from resonance.
3

Four Levels of Structure

Protein architecture is described at four hierarchical levels: primary (sequence), secondary (local folding motifs), tertiary (overall 3-D shape of one chain), and quaternary (multi-subunit assembly). Each level depends on distinct chemical interactions.
4

Structure Determines Function

A protein's biological role—enzyme catalysis, signal transduction, structural support, or molecular transport—emerges from its three-dimensional conformation. Mutations that alter folding can abolish function and cause disease.
KEY TAKEAWAY
Think of a protein like a sentence composed from a 20-letter alphabet: just as the specific order of letters in a sentence conveys a particular meaning, the precise sequence of amino acids in a polypeptide encodes a unique three-dimensional shape and therefore a unique biological function. Change one critical "letter" and the meaning—and the function—can be entirely lost.

Amino Acid Structure & the Peptide Bond

The diagram below illustrates the general structure of an amino acid at physiological pH (≈ 7.4), where the molecule exists as a zwitterion with the amino group protonated (−NH₃⁺) and the carboxyl group deprotonated (−COO⁻). The α-carbon sits at the center, bonded to the amino group, the carboxylate, a hydrogen atom, and the variable R group that distinguishes each of the twenty standard amino acids. Adjacent to this single-amino-acid depiction, the formation of a peptide bond between two amino acids is shown, emphasizing the loss of a water molecule and the resulting planar amide linkage.

Left: the general amino acid structure at physiological pH, showing the protonated amino group (NH₃⁺), deprotonated carboxylate (COO⁻), hydrogen, and variable R group attached to the central α-carbon. Right: the condensation reaction that forms a peptide bond between two amino acids, releasing one water molecule and creating a planar amide linkage whose resonance constrains rotation.

The planarity of the peptide bond is a crucial constraint: because the C−N bond carries roughly 40% double-bond character due to resonance between the carbonyl oxygen and the nitrogen lone pair, the six atoms of the peptide unit (Cα₁, C, O, N, H, Cα₂) lie in a single plane. Rotation is permitted only about the bonds flanking the peptide bond—specifically the phi (ϕ) angle (N−Cα rotation) and the psi (ψ) angle (Cα−C rotation). The allowed combinations of ϕ and ψ are mapped on a Ramachandran plot, which defines the conformational space accessible to polypeptide backbones and dictates which secondary structures can form.

The Four Levels of Protein Structure

Protein architecture is conventionally described at four hierarchical levels, each stabilized by a characteristic set of chemical forces. Primary structure is the linear sequence of amino acids joined by peptide bonds; it is genetically determined and represents the covalent backbone of the polypeptide. Secondary structure refers to local, regular folding patterns—primarily the α-helix and the β-pleated sheet—stabilized by backbone hydrogen bonds between the carbonyl oxygen of one residue and the amide hydrogen of another. In an α-helix, these hydrogen bonds form between residues separated by four positions (i → i + 4), coiling the backbone into a right-handed helix with 3.6 residues per turn. In β-sheets, hydrogen bonds connect extended backbone segments (strands) running either parallel or antiparallel to each other.

Tertiary structure describes the overall three-dimensional fold of a single polypeptide chain, arising from interactions among the R groups: hydrophobic interactions that drive nonpolar side chains into the protein interior, ionic bonds (salt bridges) between oppositely charged residues, hydrogen bonds between polar side chains, van der Waals forces, and disulfide bonds (covalent S−S linkages between cysteine residues). Quaternary structure exists only in proteins composed of two or more polypeptide subunits and describes their spatial arrangement. Hemoglobin, for instance, is a tetramer of two α and two β subunits whose quaternary interactions mediate cooperative oxygen binding.

PEPTIDE BOND FORMATION (CONDENSATION)
AA₁−COOH + H₂N−AA₂ → AA₁−CO−NH−AA₂ + H₂O
Each peptide bond releases one water molecule; a polypeptide of n amino acids contains (n − 1) peptide bonds.
APPROXIMATE MOLECULAR WEIGHT
MW ≈ n × 110 Da
Where n = number of amino acid residues; 110 Da is the average residue molecular weight (average amino acid MW ≈ 128 Da minus 18 Da for water lost per bond).
α-HELIX PARAMETERS
Rise per residue = 1.5 Å; Residues per turn = 3.6; Pitch = 5.4 Å
The pitch (rise per full turn) equals 3.6 residues × 1.5 Å/residue = 5.4 Å. Hydrogen bonds form between the C=O of residue i and the N−H of residue i + 4.

Amino Acid Classification & Protein Diversity

The twenty standard amino acids can be classified by the chemical properties of their R groups, a classification that directly informs our understanding of protein folding and function. Nonpolar (hydrophobic) side chains—alanine, valine, leucine, isoleucine, methionine, phenylalanine, tryptophan, and proline—tend to cluster in the interior of globular proteins, driven by the hydrophobic effect, which is the dominant thermodynamic force in protein folding. Polar uncharged residues—serine, threonine, asparagine, glutamine, tyrosine, and cysteine—can form hydrogen bonds and are often found on protein surfaces or at active sites. Positively charged residues at physiological pH include lysine, arginine, and histidine, while negatively charged residues include aspartate and glutamate. Glycine, the smallest amino acid with only a hydrogen as its R group, is uniquely flexible and permits conformations inaccessible to other residues.

The four levels of protein structure illustrated side by side. Primary structure is the amino acid sequence joined by peptide bonds. Secondary structure consists of local patterns (α-helix and β-sheet) held by backbone hydrogen bonds. Tertiary structure is the complete 3-D fold stabilized by R-group interactions. Quaternary structure describes the arrangement of multiple polypeptide subunits, as in hemoglobin's α₂β₂ tetramer.
Classification of the 20 standard amino acids by R-group properties
R-Group CategoryRepresentative Amino AcidsKey PropertyTypical Location in Globular Protein
Nonpolar / HydrophobicAla, Val, Leu, Ile, Phe, Trp, Met, ProAvoid water; pack together via van der Waals forcesProtein interior (core)
Polar UnchargedSer, Thr, Asn, Gln, Tyr, CysForm H-bonds with water and other residuesSurface or active sites
Positively Charged (Basic)Lys, Arg, HisCarry +1 charge at pH 7.4 (His conditionally)Surface; DNA-binding domains
Negatively Charged (Acidic)Asp, GluCarry −1 charge at pH 7.4Surface; catalytic sites
Special CasesGly (smallest), Pro (cyclic)Gly: maximal flexibility; Pro: restricts backbone anglesTurns, loops, flexible regions

Worked Example: Analyzing a Polypeptide

Consider a polypeptide with the primary sequence Ala-Lys-Phe-Glu-Cys-Ser-Trp-Arg (8 residues). We will determine its approximate molecular weight, count its peptide bonds, predict which residues are likely buried in the hydrophobic core vs. exposed on the surface, and identify possible intramolecular interactions that could stabilize its tertiary fold.

Polypeptide Analysis: Ala-Lys-Phe-Glu-Cys-Ser-Trp-Arg
1
Step 1 — Count Peptide BondsA polypeptide with n amino acid residues contains (n − 1) peptide bonds. With 8 residues, we have 8 − 1 = 7 peptide bonds.
7 peptide bonds
2
Step 2 — Estimate Molecular WeightUsing the average residue molecular weight of approximately 110 Da (the average amino acid MW of ~128 Da minus 18 Da for the water lost per peptide bond formation): MW ≈ 8 × 110 Da = 880 Da. Note that the actual MW will vary because different amino acids have different molecular weights; this provides a useful first approximation.
MW ≈ 880 Da
3
Step 3 — Classify Residues by PolarityNonpolar (hydrophobic): Ala, Phe, Trp — these would tend to be buried in the interior of a larger protein. Polar uncharged: Cys, Ser — capable of hydrogen bonding; Cys can form disulfide bonds. Positively charged: Lys (+1), Arg (+1). Negatively charged: Glu (−1). Net charge at pH 7.4 = (+2) + (−1) = +1.
Net charge at pH 7.4 = +1
4
Step 4 — Predict Stabilizing InteractionsPotential tertiary interactions include: (1) a salt bridge between Lys(+) and Glu(−); (2) hydrophobic packing among Ala, Phe, and Trp side chains; (3) hydrogen bonds involving Ser and Cys hydroxyl/sulfhydryl groups; (4) possible cation-π interaction between Arg or Lys and the aromatic ring of Phe or Trp. If another Cys were present, a disulfide bond could form under oxidizing conditions, but with only one Cys here, an intramolecular disulfide is not possible.
Salt bridge (Lys–Glu), hydrophobic core (Ala, Phe, Trp), H-bonds (Ser, Cys), possible cation-π

Functional Classes of Proteins

Proteins serve an astonishing diversity of biological roles, and classifying them by function highlights the intimate relationship between molecular structure and physiological purpose. The table below compares major functional categories, illustrating how different structural features suit different biological tasks. Enzymes, for instance, rely on precisely shaped active sites that bind substrates with high specificity and lower activation energies, while structural proteins like collagen derive their tensile strength from repeating triple-helical motifs rich in glycine and proline.

Major functional classes of proteins with representative examples
Functional ClassExample(s)Key Structural FeatureBiological Role
EnzymesDNA polymerase, hexokinase, lysozymeActive site with precise geometry; often includes cofactorsCatalyze specific biochemical reactions, lowering activation energy
StructuralCollagen, keratin, elastinFibrous, repetitive sequences; triple helices or coiled coilsProvide mechanical support, elasticity, and tensile strength
TransportHemoglobin, transferrin, albuminLigand-binding pockets; allosteric conformational changesBind and carry molecules (O₂, Fe³⁺, fatty acids) through blood
SignalingInsulin, growth hormone, G proteinsReceptor-binding surfaces; conformational switchingTransmit chemical signals between or within cells
DefenseImmunoglobulins (antibodies), complementVariable antigen-binding regions; constant Fc domainsRecognize and neutralize pathogens in adaptive immunity
MotorMyosin, kinesin, dyneinATPase domains; lever-arm or walking mechanismsConvert chemical energy (ATP) into mechanical movement
KEY TAKEAWAY
Imagine proteins as specialized tools in a toolkit: a wrench (enzyme) has a shaped jaw that grips a specific bolt (substrate), a steel cable (structural protein) resists pulling forces through its braided architecture, and a delivery truck (transport protein) has compartments fitted to carry particular cargo. The shape of each tool is inseparable from its function—redesign the shape, and the tool either works differently or not at all. This principle extends to disease: misfolded proteins in Alzheimer's, sickle-cell anemia, and cystic fibrosis illustrate how even small structural perturbations have devastating functional consequences.

Connections to Advanced Protein Biochemistry

The foundational concepts covered in this lesson—amino acid chemistry, peptide bond geometry, and the four levels of structure—serve as the gateway to several advanced topics in biochemistry and molecular biology. Enzyme kinetics applies quantitative frameworks (Michaelis-Menten and beyond) to understand how enzymes accelerate reactions. Protein folding thermodynamics examines the free-energy landscape that guides a nascent polypeptide from its unfolded ensemble to its native state, including the roles of chaperones like GroEL/GroES. Post-translational modifications (phosphorylation, glycosylation, ubiquitination) expand the functional repertoire of proteins far beyond what the genetic code alone encodes.

How foundational protein concepts connect to advanced biochemistry
Foundational Concept (This Lesson)Advanced ExtensionWhy It Matters
Amino acid R-group chemistryEnzyme active-site catalysis & mechanismR groups act as general acids/bases, nucleophiles, and electrostatic stabilizers in catalytic mechanisms
Peptide bond planarity & ϕ/ψ anglesRamachandran analysis & computational structure predictionConformational constraints guide algorithms like AlphaFold and Rosetta in predicting 3-D folds
Tertiary structure & folding forcesProtein misfolding diseases (prions, amyloidosis)Misfolded proteins aggregate into toxic fibrils; understanding folding suggests therapeutic targets
Quaternary structure & subunit assemblyAllostery & cooperative bindingHemoglobin's sigmoidal O₂-binding curve arises from subunit cooperativity, a hallmark of quaternary regulation
Primary sequence determines structureProtein engineering & directed evolutionRational design and laboratory evolution exploit sequence-structure relationships to create novel proteins

As you progress through biochemistry, keep in mind that the principles of amino acid chemistry and protein folding recur at every level of biological organization—from the regulation of gene expression by transcription factors to the signaling cascades that control cell fate. Mastering the material in this lesson equips you with the conceptual vocabulary needed to engage with these advanced topics, whether in enzymology, structural biology, pharmacology, or synthetic biology.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain why the peptide bond is described as having "partial double-bond character." What structural consequence does this have for the polypeptide backbone, and how does it constrain protein folding?
PROBLEM 2BASIC CALCULATION
A protein contains 450 amino acid residues. (a) How many peptide bonds does it contain? (b) Estimate its approximate molecular weight in kilodaltons (kDa). (c) If the protein forms an α-helical segment spanning residues 100–136, how long is this helical segment in angstroms?
PROBLEM 3INTERMEDIATE
A researcher mutates a buried leucine residue in the hydrophobic core of a globular enzyme to an aspartate residue. Predict the likely effects of this mutation on (a) the protein's tertiary structure, (b) its thermodynamic stability, and (c) its enzymatic activity. Justify your reasoning by referencing the chemical properties of the original and mutant side chains.
PROBLEM 4APPLIED
Sickle-cell disease results from a single amino acid substitution in the β-globin subunit of hemoglobin: glutamate at position 6 is replaced by valine (E6V). Using your knowledge of amino acid properties and quaternary structure, explain the molecular mechanism by which this single substitution causes hemoglobin molecules to polymerize into long fibers under low-oxygen conditions, deforming red blood cells.
PROBLEM 5CRITICAL THINKING
Anfinsen's classic ribonuclease refolding experiment demonstrated that the primary sequence of a protein contains sufficient information to dictate its native three-dimensional structure. However, many proteins in vivo require molecular chaperones (such as Hsp70 or GroEL/GroES) to fold correctly. Does the requirement for chaperones contradict Anfinsen's dogma? Construct a carefully reasoned argument addressing both perspectives, and discuss how Levinthal's paradox informs the debate.

Proteins — Summary

Proteins are polypeptides built from 20 standard amino acids linked by peptide bonds that are planar due to resonance-derived partial double-bond character. Each amino acid's R group determines its chemical personality—nonpolar, polar, positively charged, or negatively charged—and these side-chain properties collectively govern folding and function. Protein architecture is organized into four hierarchical levels: primary structure (amino acid sequence), secondary structure (α-helices and β-sheets stabilized by backbone hydrogen bonds), tertiary structure (the overall 3-D fold stabilized by R-group interactions including the hydrophobic effect, salt bridges, hydrogen bonds, van der Waals forces, and disulfide bonds), and quaternary structure (multi-subunit assembly).

The central dogma of protein science—structure determines function—is exemplified across all functional classes: enzymes catalyze reactions via precisely shaped active sites, structural proteins derive strength from fibrous architectures, and transport proteins carry cargo through ligand-binding conformational changes. Disruptions to folding—whether from point mutations (as in sickle-cell disease) or environmental stress—underscore the fragility and elegance of the sequence-structure-function relationship that lies at the heart of modern biochemistry.

Varsity Tutors • College Biology • Proteins