COLLEGE BIOLOGY • EVOLUTION & NATURAL SELECTION

Phylogeny

Reconstructing the evolutionary relationships among organisms through shared ancestry and divergence.

Historical Context & Motivation

Long before Darwin set sail on HMS Beagle, naturalists wrestled with a deceptively simple question: how should we organize the staggering diversity of life? Early classification systems, such as the Great Chain of Being popularized during the medieval period, arranged organisms on a linear ladder from 'lower' to 'higher' forms—a scheme that reflected theological assumptions more than biological reality. The critical intellectual shift occurred when scientists began to suspect that the similarities among organisms were not mere design coincidences but rather signatures of shared descent from common ancestors. This realization demanded a fundamentally different way of organizing life—one based on genealogy rather than superficial resemblance. Phylogeny, the study of evolutionary relationships among species and other taxa, arose to meet exactly this need.

1735
Linnaeus Publishes Systema Naturae
Carl Linnaeus introduced the hierarchical classification system (kingdom, class, order, genus, species) and binomial nomenclature. Although Linnaeus did not propose evolution, his nested groupings inadvertently reflected the branching pattern of descent that later biologists would interpret phylogenetically.
1859
Darwin's On the Origin of Species
Charles Darwin provided the theoretical engine—natural selection—and sketched the first explicitly evolutionary tree diagram in his notebook. He argued that Linnaeus's hierarchical groups are hierarchical because life diversifies through branching descent, not because of a divine plan.
1950
Hennig Formalizes Cladistics
German entomologist Willi Hennig published Grundzüge einer Theorie der phylogenetischen Systematik, establishing cladistics—the method of classifying organisms strictly by shared derived characters (synapomorphies). Hennig's framework became the gold standard for modern phylogenetic analysis.
1977
Woese and the Three-Domain System
Carl Woese used ribosomal RNA (rRNA) sequence comparisons to show that prokaryotes comprise two fundamentally distinct lineages—Bacteria and Archaea—alongside Eukarya. This molecular phylogeny revolutionized our understanding of life's deepest branches.
2000s–present
Genomic-Scale Phylogenetics
High-throughput sequencing technologies and computational advances (Bayesian inference, maximum likelihood) now enable phylogenies built from thousands of genes simultaneously, revealing complex phenomena such as horizontal gene transfer and incomplete lineage sorting.

The central question phylogeny addresses is both elegant and profound: given a set of organisms alive today—and, when fossils are available, organisms from the past—what branching pattern of ancestor–descendant relationships best explains the observed distribution of characters across those taxa? Answering this question requires rigorous methods for distinguishing homology (similarity due to shared ancestry) from homoplasy (similarity due to convergent evolution, parallelism, or reversal). The remainder of this lesson explores the principles, methods, and applications of phylogenetic analysis.

Core Principles & Definitions

Phylogenetic analysis rests on several foundational concepts that govern how evolutionary relationships are inferred, represented, and interpreted. A firm grasp of these principles is essential before examining the technical methods used to construct phylogenetic trees. At its core, phylogeny assumes that all life on Earth shares a common ancestor and that lineages diverge over time through processes of speciation and extinction. The patterns left behind by these processes can be recovered—albeit imperfectly—by examining the characters organisms share or differ in.

1

Phylogenetic Tree (Cladogram)

A branching diagram depicting hypothesized evolutionary relationships among taxa. Each node represents a speciation event (a hypothetical common ancestor), each branch represents a lineage evolving through time, and each tip represents a taxon (extant or extinct).
2

Monophyletic Group (Clade)

A clade consists of an ancestor and all of its descendants. Cladistic classification insists that only monophyletic groups receive formal taxonomic names. A group that excludes some descendants is paraphyletic; one that includes unrelated lineages is polyphyletic.
3

Synapomorphy vs. Plesiomorphy

A synapomorphy is a shared derived character that unites a clade. A plesiomorphy (ancestral character) is uninformative for grouping because it was already present before the clade diverged. Only synapomorphies diagnose clades.
4

Outgroup Comparison

To determine which character state is ancestral vs. derived, systematists compare the ingroup (taxa of interest) to one or more outgroups—taxa known to fall outside the ingroup. Characters shared with the outgroup are presumed plesiomorphic.
5

Parsimony Principle

When multiple tree topologies can explain the data, the most parsimonious tree—the one requiring the fewest evolutionary changes—is preferred (all else being equal). While parsimony is not always the best criterion, it remains a foundational heuristic in cladistics.
KEY TAKEAWAY
Think of a phylogenetic tree like a family genealogy chart. Just as cousins share grandparents and siblings share parents, species that cluster together on a phylogeny share a more recent common ancestor than species on distant branches. A clade is analogous to an entire extended family descended from one couple—leave out any one descendant, and you no longer have the complete family (monophyletic group). The key difference is that phylogenies deal with millions of years and require inference from character data rather than birth records.

Anatomy of a Phylogenetic Tree

The diagram below illustrates the essential components of a phylogenetic tree (cladogram). Understanding how to read this diagram is arguably the single most important skill in phylogenetics. Note that the horizontal axis does not necessarily represent time unless the tree is explicitly drawn as a chronogram (with branch lengths scaled to time). In a standard cladogram, the branching pattern—the topology—conveys the information about relationships; branch lengths may be arbitrary.

A simplified phylogenetic tree with four taxa (A–D). The root at the base represents the most recent common ancestor (MRCA) of the entire ingroup. Each internal node marks a speciation event, and each tip represents a taxon. The dashed box highlights the clade comprising Taxa C and D, which share a more recent common ancestor with each other than either does with A or B.

A common misconception is that the order in which taxa appear at the tips of a phylogeny implies some ranking or progression. In reality, the tips can be rotated around any internal node without changing the tree's meaning—only the branching topology (which nodes connect to which) carries phylogenetic information. Taxa A and B in the diagram above are sister taxa because they share an immediate common ancestor that is not shared with C or D. Likewise, the clade (C + D) is the sister group to the clade (A + B). Recognizing these relationships from the branching pattern is the fundamental skill of phylogenetic literacy.

Methods of Phylogenetic Inference

Constructing a phylogeny is fundamentally a problem of inference: given a matrix of character data (morphological traits, DNA sequences, protein sequences, etc.), which tree topology best explains the observed patterns? Several algorithmic frameworks have been developed, each with its own philosophical assumptions, computational demands, and strengths. The three dominant approaches in modern systematics are maximum parsimony, maximum likelihood, and Bayesian inference.

Maximum Parsimony

Under maximum parsimony, the preferred tree is the one that requires the fewest total character-state changes. For a given alignment of DNA sequences, this means counting, for every possible tree topology, the minimum number of nucleotide substitutions needed to explain the data. The tree with the smallest total 'tree length' is selected. While parsimony is intuitive and computationally tractable for small datasets, it can be inconsistent in the statistical sense—particularly when certain branches evolve much faster than others, a situation known as long-branch attraction.

PARSIMONY TREE LENGTH
L(T) = Σᵢ (minimum substitutions at site i given topology T)
Where L(T) is the parsimony score (total tree length) for topology T, and the sum is taken over all informative sites i in the alignment. The topology with the smallest L(T) is the most parsimonious tree.

Maximum Likelihood (ML)

Maximum likelihood methods evaluate the probability of observing the data given a particular tree topology and a model of sequence evolution (e.g., the Jukes–Cantor model or the General Time Reversible model). The tree and branch-length combination that maximizes this likelihood is selected. ML is statistically consistent—given enough data, it will converge on the true tree—but it is computationally intensive, especially as the number of taxa grows.

LIKELIHOOD FUNCTION
L(T, θ | D) = P(D | T, θ) = ∏ᵢ P(Dᵢ | T, θ)
Where D is the observed sequence alignment, T is the tree topology with branch lengths, θ represents the parameters of the substitution model, and the product runs over all alignment sites i. The tree that maximizes this product is chosen.

Bayesian Inference

Bayesian phylogenetics uses Bayes' theorem to compute the posterior probability of a tree given the data and a prior distribution on tree topologies and model parameters. This is implemented via Markov chain Monte Carlo (MCMC) algorithms that sample trees in proportion to their posterior probability. The result is not a single 'best' tree but a distribution of trees, from which a consensus topology and clade-support values (posterior probabilities) can be derived. Bayesian methods share ML's statistical consistency and additionally provide a natural framework for incorporating prior knowledge.

BAYES' THEOREM FOR PHYLOGENY
P(T, θ | D) = P(D | T, θ) × P(T, θ) / P(D)
The posterior probability P(T, θ | D) is proportional to the likelihood P(D | T, θ) multiplied by the prior P(T, θ). The denominator P(D) is the marginal likelihood, which normalizes the posterior but is typically intractable and thus bypassed via MCMC sampling.
⚠️ Number of Possible Trees
For n taxa, the number of possible unrooted bifurcating trees is (2n − 5)!! (double factorial). For just 10 taxa this yields 2,027,025 possible trees; for 20 taxa, over 2 × 1020. Exhaustive search is therefore impossible for most real datasets, and heuristic search strategies (branch-swapping, stepwise addition) are essential.

Data Sources & Character Types

Phylogenetic trees are only as reliable as the data used to build them. Characters used in phylogenetic analysis fall broadly into morphological, molecular, and behavioral categories. In modern systematics, molecular data—particularly DNA and protein sequences—dominate because they provide a large number of discrete, heritable characters that can be objectively compared across distantly related organisms. Nonetheless, morphological characters remain indispensable for placing fossils on the tree of life, since sequence data are rarely recoverable from ancient specimens.

An overview of the phylogenetic analysis pipeline. Data from morphological, molecular, or genomic sources are assembled into a multiple sequence alignment (or character matrix), which is then analyzed by one of the major inference methods to produce a tree with statistical support values.
Comparison of major data sources for phylogenetic analysis
Data TypeExampleStrengthsLimitations
MorphologicalSkeletal features, leaf shape, flower symmetryApplicable to fossils; directly observableProne to convergence; limited number of characters
DNA SequencesNuclear genes, mitochondrial DNA, chloroplast DNAAbundant, discrete, heritable; applicable across all lifeRequires intact DNA; substitution saturation at deep timescales
Protein SequencesCytochrome c, hemoglobin, ribosomal proteinsEvolves more slowly than DNA—useful for ancient divergencesCodon degeneracy reduces resolution at recent timescales
Rare Genomic ChangesSINE insertions, gene fusions, intron positionsVery low homoplasy; nearly unambiguous synapomorphiesRare—too few characters for fine-scale resolution

Worked Example: Building a Simple Cladogram

Suppose we have five vertebrate taxa—lamprey, trout, lizard, pigeon, and mouse—and we wish to construct a cladogram based on the presence or absence of six characters: vertebral column, hinged jaw, four limbs, amnion (amniotic egg), feathers, and hair/fur. We will use the lamprey as the outgroup (it diverged earliest) and apply maximum parsimony to identify the tree requiring the fewest character-state changes.

Constructing a Vertebrate Cladogram by Parsimony
1
Step 1 — Assemble the Character MatrixWe begin by coding each character as 0 (absent) or 1 (present) for each taxon. Lamprey: 1,0,0,0,0,0. Trout: 1,1,0,0,0,0. Lizard: 1,1,1,1,0,0. Pigeon: 1,1,1,1,1,0. Mouse: 1,1,1,1,0,1. The vertebral column (character 1) is shared by all taxa and is therefore uninformative for resolving relationships within the ingroup—it is a plesiomorphy at this level.
2
Step 2 — Identify Synapomorphies via Outgroup ComparisonThe lamprey (outgroup) lacks jaws, limbs, amnion, feathers, and hair. Therefore, the presence of each of these characters is derived within the ingroup. Jaws unite trout + lizard + pigeon + mouse. Four limbs unite lizard + pigeon + mouse. The amnion unites lizard + pigeon + mouse. Feathers are unique to pigeon (an autapomorphy), and hair is unique to mouse (also an autapomorphy).
Four informative synapomorphies identified: jaws, limbs, amnion, feathers/hair as autapomorphies.
3
Step 3 — Build the Tree Using Nested SynapomorphiesWe nest the taxa according to successively more exclusive shared derived characters. Jaws → (Trout, (Lizard, Pigeon, Mouse)). Four limbs → (Lizard, Pigeon, Mouse). Amnion → (Lizard, Pigeon, Mouse)—this is the same grouping as limbs here, so we look for characters that further resolve this clade. The pigeon has feathers (not shared with lizard or mouse); the mouse has hair (not shared with lizard or pigeon). Without additional data, lizard, pigeon, and mouse form an unresolved trichotomy, or we can use additional characters (e.g., endothermy) to resolve pigeon + mouse as a clade with lizard as the outgroup within Amniota.
4
Step 4 — Calculate the Parsimony ScoreCount the minimum number of character-state changes on the tree. Vertebral column: 1 change (at the root). Jaws: 1 change (at the node uniting jawed vertebrates). Limbs: 1 change (at the tetrapod node). Amnion: 1 change (at the amniote node). Feathers: 1 change (on the pigeon branch). Hair: 1 change (on the mouse branch).
Total parsimony score: L(T) = 6 changes. No homoplasy is required—each character changes exactly once, making this the most parsimonious topology.
5
Step 5 — Interpret the ResultThe resulting cladogram shows a nested hierarchy: (Lamprey, (Trout, (Lizard, (Pigeon, Mouse)))). Each branching point corresponds to the origin of a synapomorphy. This tree mirrors the well-established vertebrate phylogeny and demonstrates how a small set of characters can recover major evolutionary relationships when homoplasy is minimal.
The cladogram correctly recovers Gnathostomata (jawed vertebrates), Tetrapoda (four-limbed vertebrates), and Amniota (amniote egg-layers) as nested monophyletic groups.

Strengths & Limitations of Phylogenetic Methods

No single phylogenetic method is universally superior. The choice of method depends on the dataset size, the depth of divergence being investigated, the computational resources available, and the specific biological questions being asked. The table below summarizes the major trade-offs among the three dominant inference frameworks. In practice, many researchers employ multiple methods and look for concordance among them—if parsimony, likelihood, and Bayesian approaches all recover the same topology with strong support, confidence in that topology is high.

Comparison of the three main phylogenetic inference methods
CriterionMaximum ParsimonyMaximum LikelihoodBayesian Inference
Philosophical basisOccam's razor—fewest changes preferredFrequentist probability—maximize P(Data | Model)Bayesian probability—posterior distribution over trees
Statistical consistencyInconsistent under certain conditions (long-branch attraction)Consistent if the model is correctConsistent if the model and priors are reasonable
Model requirementModel-free (no explicit substitution model)Requires explicit substitution modelRequires model + prior distributions
Computational costLow to moderateModerate to highHigh (MCMC chains must converge)
Support measureBootstrap supportBootstrap supportPosterior probability
Best suited forMorphological data, exploratory analysisLarge molecular datasets, model testingDivergence-time estimation, complex models
KEY TAKEAWAY
Phylogenetic inference is analogous to forensic detective work: the 'crime' (speciation events) happened in the deep past, and investigators (systematists) must reconstruct what happened from incomplete evidence (character data). Different analytical methods—parsimony, likelihood, Bayesian—are like different forensic techniques. No single technique is foolproof, but when multiple independent lines of evidence converge on the same conclusion, our confidence soars. The most robust phylogenies are those supported by multiple data types analyzed under multiple methods.

Connections to Advanced Evolutionary Theory

The simple bifurcating tree model, while powerful, represents an idealization. Several phenomena in nature produce patterns that deviate from strictly tree-like evolution, challenging the core assumptions of phylogenetic analysis. Understanding these complications is crucial for interpreting phylogenies in the genomic era and connects phylogeny to cutting-edge research in evolutionary biology.

From basic phylogeny to advanced evolutionary analysis
ConceptBasic Phylogeny (This Lesson)Advanced / Research Frontier
Tree shapeStrictly bifurcating trees (each node splits into two)Phylogenetic networks that accommodate reticulate evolution (hybridization, HGT)
Gene vs. species treeAssumed to be equivalent—one gene tree = species treeGene trees can differ from the species tree due to incomplete lineage sorting (ILS), gene duplication/loss, and HGT
Divergence timingCladograms show topology only (no time axis)Molecular clocks and fossil calibrations produce dated phylogenies (chronograms)
Character evolutionCharacters mapped onto a fixed treeAncestral state reconstruction, rates of trait evolution (comparative methods)
DiversificationTree topology describes relatednessBirth-death models estimate speciation and extinction rates from tree shape

One of the most active areas in modern phylogenetics is the study of gene tree–species tree discordance. When thousands of gene trees are inferred from a genome, they frequently disagree with one another and with the species tree. This discordance is expected under incomplete lineage sorting (ILS), in which ancestral polymorphisms persist through multiple speciation events and sort randomly into descendant lineages. Methods such as the multispecies coalescent model explicitly account for ILS and infer the species tree from a distribution of gene trees. Similarly, horizontal gene transfer (especially prevalent in prokaryotes) means that the history of life is better represented as a network than a strictly bifurcating tree. These complexities do not invalidate the tree model—they refine it, and they make phylogenetics an even richer and more exciting field of study.

🔭 Looking Ahead
Courses in molecular evolution, bioinformatics, and comparative genomics build directly on the phylogenetic foundations covered here. If you pursue research in ecology, conservation biology, epidemiology (tracking pathogen evolution), or even linguistics, phylogenetic thinking will be an indispensable tool in your intellectual toolkit.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain the difference between a monophyletic group, a paraphyletic group, and a polyphyletic group. Why does modern cladistic classification insist that all named taxa be monophyletic?
PROBLEM 2BASIC CALCULATION
How many possible unrooted bifurcating tree topologies exist for 7 taxa? Use the formula (2n − 5)!! where n is the number of taxa and !! denotes the double factorial.
PROBLEM 3INTERMEDIATE
You are given the following DNA alignment for four taxa at a single site: Taxon A = G, Taxon B = G, Taxon C = T, Taxon D = T. The outgroup shows state G. There are three possible unrooted trees for four taxa: Tree 1: ((A,B),(C,D)); Tree 2: ((A,C),(B,D)); Tree 3: ((A,D),(B,C)). What is the minimum number of substitutions (parsimony score) for this site on each tree?
PROBLEM 4APPLIED
A conservation biologist must prioritize one of two endangered species for a limited rescue program. Species X is one of 20 species in a large, recently diversified genus. Species Y is the sole surviving member of an ancient lineage with no close relatives (it sits on a very long branch in the phylogeny). Using the concept of phylogenetic diversity, which species should the biologist prioritize, and why?
PROBLEM 5CRITICAL THINKING
A researcher builds a phylogeny of 50 bacterial species using the 16S rRNA gene and recovers a well-supported tree. A colleague then builds a phylogeny of the same 50 species using a metabolic gene and obtains a significantly different topology. Propose at least three biological explanations for this discordance (not methodological artifacts) and describe how you might determine which explanation is most likely.

Summary

Phylogeny is the study of evolutionary relationships among organisms, reconstructed from morphological, molecular, and genomic data. The central output is a phylogenetic tree, whose branching topology represents hypothesized ancestor–descendant relationships. Valid taxonomic groups must be monophyletic (clades), defined by synapomorphies (shared derived characters), and distinguished from plesiomorphies (ancestral characters) via outgroup comparison.

Three major inference methods dominate modern practice: maximum parsimony (fewest changes), maximum likelihood (best probability given a substitution model), and Bayesian inference (posterior probability distribution over trees). Each has distinct strengths and limitations, and robust phylogenies are those supported by multiple methods and data types. Advanced topics—including gene tree–species tree discordance, horizontal gene transfer, molecular clocks, and phylogenetic diversity in conservation—extend these foundations into the frontiers of evolutionary research.

Varsity Tutors • College Biology • Phylogeny