GENETICS • MOLECULAR GENETICS TECHNIQUES & GENOMICS

Linkage Disequilibrium

Understanding why certain gene variants travel together through generations and what that reveals about our DNA.

Historical Context & Motivation

When scientists first started mapping genes on chromosomes, they noticed something curious. Certain combinations of gene variants seemed to show up together far more often than you would expect by chance. This observation puzzled early geneticists because it seemed to break the rules of random inheritance that Gregor Mendel had established decades earlier.

The story of linkage disequilibrium (often shortened to LD) begins with the discovery of genetic linkage — the idea that genes located close together on the same chromosome tend to be inherited as a group. Over time, researchers realized that LD could reveal the hidden history of populations, help find disease-causing genes, and even trace human migration patterns across the globe.

1905
Bateson & Punnett Discover Linkage
William Bateson and Reginald Punnett noticed that certain pea traits did not sort independently. They found that some trait combinations appeared together far more than expected, hinting that genes could be physically connected.
1913
Sturtevant Maps Genes
Alfred Sturtevant, a student of Thomas Hunt Morgan, created the first genetic map by measuring how often genes separated during crossing over in fruit flies. This confirmed that genes have physical positions on chromosomes.
1960
LD Is Formally Defined
Population geneticists Lewontin and Kojima formally defined linkage disequilibrium as a measurable statistic, giving scientists a precise way to quantify how strongly two gene variants are associated in a population.
2002
The HapMap Project Begins
The International HapMap Project set out to catalog patterns of LD across the human genome. By identifying blocks of DNA that travel together, researchers could search for disease genes far more efficiently.
2010s
GWAS Revolution
Genome-wide association studies (GWAS) used LD maps to link thousands of genetic variants to diseases like diabetes, heart disease, and cancer. LD became an essential tool in modern genomics.

The central question that LD addresses is this: Why do certain gene variants stick together instead of mixing freely? Understanding the answer helps scientists locate genes involved in diseases, trace ancestry, and predict how populations evolve over time.

Core Principles & Definitions

Before diving into linkage disequilibrium, you need to understand a few building blocks. An allele is a specific version of a gene — think of it like choosing between chocolate and vanilla at an ice cream shop. A locus (plural: loci) is the exact spot on a chromosome where a gene sits. When we talk about LD, we are looking at how alleles at two different loci relate to each other in a population.

1

Linkage Equilibrium

When alleles at two loci combine randomly, like shuffling two separate decks of cards. The presence of one allele tells you nothing about the other.
2

Linkage Disequilibrium

When alleles at two loci appear together more or less often than you would predict by chance. They are not combining randomly — something is keeping them connected.
3

Haplotype

A haplotype is a set of alleles inherited together on the same chromosome. Think of it as a 'package deal' of genetic variants that travel as a group from parent to child.
4

Recombination

During the formation of egg and sperm cells, chromosomes can swap segments in a process called crossing over. Over many generations, this shuffles alleles and breaks down LD.
5

D and r² Values

Scientists measure LD using statistics called D and . A value of r² = 1 means two alleles always appear together; r² = 0 means they combine completely at random.
KEY TAKEAWAY
Imagine you have a bag of beads where red beads and square beads always end up on the same string. If you picked a random string and saw a red bead, you could predict there is also a square bead. That predictable pairing is linkage disequilibrium. If bead color and shape were mixed randomly — sometimes red with round, sometimes blue with square — that would be linkage equilibrium. LD means there is a non-random pattern, and that pattern tells us something about the history and structure of the genome.

Visualizing Linkage Disequilibrium

The diagram below shows two loci on a chromosome — Locus 1 and Locus 2. At Locus 1, there are two alleles: A and a. At Locus 2, there are also two alleles: B and b. On the left side, you can see what linkage equilibrium looks like — all four possible haplotypes (AB, Ab, aB, ab) appear at the frequencies you would expect from random mixing. On the right side, linkage disequilibrium is shown — certain combinations appear much more often than expected.

Left: In linkage equilibrium (D = 0), all four haplotypes — AB, Ab, aB, and ab — appear at equal frequencies of 25%. Right: In linkage disequilibrium (D ≠ 0), certain haplotype combinations like AB and ab are over-represented (40% each), while Ab and aB are under-represented (10% each). The bar lengths in the bottom panels show how uneven the frequencies become.

Notice how on the left side, knowing that a chromosome carries allele A tells you nothing about whether it also carries B or b. On the right side, if you see allele A, you can predict with good confidence that allele B is nearby. That predictability is the hallmark of linkage disequilibrium, and it is incredibly useful for geneticists trying to track down important genes.

Mathematical Framework

Scientists need a way to put a number on how strongly two alleles are associated. The core idea is straightforward: compare what you actually observe in a population to what you would expect if the alleles combined randomly.

LD COEFFICIENT (D)
D = f(AB) − f(A) × f(B)
Where f(AB) is the observed frequency of the AB haplotype, f(A) is the frequency of allele A in the population, and f(B) is the frequency of allele B. If D = 0, there is no LD. If D is positive or negative, LD exists.

The value of D depends on allele frequencies, which makes it hard to compare across different gene pairs. To fix this, geneticists often normalize D into two more convenient measures.

NORMALIZED LD (D')
D' = D / D_max
D' ranges from −1 to 1. A value of |D'| = 1 means that at least one of the four possible haplotypes is completely absent from the population. Dmax is the maximum possible value of D given the allele frequencies.
CORRELATION COEFFICIENT (r²)
r² = D² / [f(A) × f(a) × f(B) × f(b)]
r² ranges from 0 to 1 and measures how well you can predict the allele at one locus from the allele at the other. An r² of 1 means perfect prediction — knowing one allele tells you exactly which allele is at the other locus. This is the most commonly used measure in genetic studies.
DECAY OF LD OVER GENERATIONS
D_t = D₀ × (1 − c)^t
Where D₀ is the initial LD, c is the recombination rate between the two loci (a number between 0 and 0.5), and t is the number of generations. LD decreases over time as recombination shuffles alleles, but loci that are very close together (small c) lose LD very slowly.
🔄 Why Does LD Decay?
During meiosis, chromosomes can swap pieces in a process called crossing over. Each generation gives chromosomes another chance to break apart allele combinations. Loci that are far apart on a chromosome recombine frequently (high c), so their LD disappears quickly. Loci that sit right next to each other rarely recombine (low c), so their LD can last for thousands of generations.

Causes of Linkage Disequilibrium

LD does not appear for just one reason. Several different forces can create or maintain it, and understanding these causes helps scientists interpret what LD patterns mean.

Four major forces create linkage disequilibrium: physical linkage (the most common cause), natural selection, genetic drift, and population mixing. Recombination gradually breaks down LD over generations, as shown by the decay equation at the bottom. Loci that are physically close on a chromosome lose LD slowly because recombination between them is rare.

The most important cause of LD in practice is physical linkage. When two loci are very close on the same chromosome, recombination events between them are extremely rare. This means their alleles stay together for many generations, creating strong and lasting LD. In contrast, when two populations that carry different haplotype patterns merge, the resulting LD is usually temporary — recombination will scramble the associations within a handful of generations.

Summary of the four main causes of linkage disequilibrium
Cause of LDDurationDepends On
Physical linkageLong-lasting (thousands of generations)Distance between loci on the chromosome
Natural selectionMaintained as long as selection continuesStrength of selection on the allele combination
Genetic driftVariable; can be permanent in small populationsPopulation size
Population mixing (admixture)Temporary (decays within a few generations)Allele frequency differences between source populations

Worked Example: Calculating LD

Let's walk through a concrete example. Suppose you survey a population of 1,000 chromosomes and observe two loci, each with two alleles. You count the number of chromosomes carrying each haplotype.

Calculating D and r² from Haplotype Data
1
Step 1 — Record Observed Haplotype CountsOut of 1,000 chromosomes you observe: AB = 400, Ab = 100, aB = 100, ab = 400. Convert these to frequencies by dividing each count by 1,000: f(AB) = 0.40, f(Ab) = 0.10, f(aB) = 0.10, f(ab) = 0.40.
f(AB) = 0.40, f(Ab) = 0.10, f(aB) = 0.10, f(ab) = 0.40
2
Step 2 — Calculate Allele FrequenciesThe frequency of allele A is the sum of all haplotypes containing A: f(A) = f(AB) + f(Ab) = 0.40 + 0.10 = 0.50. Similarly, f(a) = 0.50, f(B) = f(AB) + f(aB) = 0.40 + 0.10 = 0.50, and f(b) = 0.50.
f(A) = 0.50, f(a) = 0.50, f(B) = 0.50, f(b) = 0.50
3
Step 3 — Calculate Expected Haplotype FrequencyIf there were no LD, the expected frequency of the AB haplotype would be: f(A) × f(B) = 0.50 × 0.50 = 0.25. We observe f(AB) = 0.40, which is much higher than 0.25. That difference signals LD.
Expected f(AB) = 0.25 (but observed is 0.40)
4
Step 4 — Calculate DUsing the formula D = f(AB) − f(A) × f(B): D = 0.40 − (0.50 × 0.50) = 0.40 − 0.25 = 0.15. Since D is positive and not zero, there is linkage disequilibrium. The AB and ab haplotypes are more common than expected.
D = 0.15
5
Step 5 — Calculate r²Using the formula r² = D² / [f(A) × f(a) × f(B) × f(b)]: r² = (0.15)² / (0.50 × 0.50 × 0.50 × 0.50) = 0.0225 / 0.0625 = 0.36. This means about 36% of the variation at one locus can be predicted from the other locus. This is moderate LD — the alleles are clearly associated, but not perfectly.
r² = 0.36 (moderate LD)
📊 Interpreting r² Values
As a rough guide: r² < 0.1 means very weak LD (alleles are nearly independent). Values of r² between 0.1 and 0.5 indicate moderate LD. Values of r² > 0.8 indicate strong LD, meaning the alleles are highly predictable from each other. In GWAS studies, researchers typically consider r² > 0.8 as 'useful' LD for tagging a nearby variant.

Applications, Strengths, & Limitations

Linkage disequilibrium is not just a theoretical concept — it is one of the most practical tools in modern genetics. Scientists use LD patterns to find disease genes, study population history, and design efficient genetic tests. But like every tool, LD has both strengths and limitations.

Strengths and limitations of using linkage disequilibrium in genetics research
StrengthsLimitations
Allows genome-wide association studies (GWAS) to scan millions of variants using only a fraction of them as 'tags' — saving time and money.LD patterns differ between populations, so results from one ethnic group may not apply directly to another.
Helps narrow down the location of disease-causing mutations by identifying regions that travel together.LD only shows association, not causation. A variant in LD with a disease gene is not necessarily the cause itself.
Reveals population history — bottlenecks, migrations, and admixture events leave distinct LD signatures.LD decays over time due to recombination, so very old associations may be undetectable.
Used in breeding programs to select favorable trait combinations in crops and livestock.Fine-mapping requires very dense marker panels to find the exact causal variant within an LD block.
KEY TAKEAWAY
Think of LD like using footprints in the sand. If you see a set of sneaker prints next to a set of dog paw prints along a beach, you can guess that a person was walking their dog — even if you never actually see the person or the dog. In genetics, a tag SNP (a variant in LD with a disease gene) acts like those sneaker prints. You do not need to test every single variant in the genome — you just need a few well-placed 'footprints' that reliably predict the variants nearby. This is what makes GWAS studies possible and affordable.

Connection to Advanced Topics

Linkage disequilibrium connects to several advanced areas of genetics and genomics. As you continue your studies, you will see LD appear again and again as a foundation for more complex analyses.

Advanced topics that build upon linkage disequilibrium
ConceptHow It Relates to LDLevel
Haplotype blocksRegions of strong LD form 'blocks' separated by recombination hotspots. The human genome can be divided into about 500,000 such blocks.AP Biology / College
ImputationUsing LD patterns to infer (impute) genotypes that were not directly measured, greatly increasing the power of genetic studies without extra cost.College / Graduate
Selective sweepsWhen natural selection rapidly increases the frequency of a beneficial allele, it drags nearby variants along, creating a long stretch of high LD. Scientists use this pattern to detect recent evolution.AP Biology / College
Polygenic risk scoresThese scores combine information from thousands of variants to predict disease risk. LD must be carefully accounted for so that the same signal is not counted multiple times.College / Graduate

As sequencing technology becomes faster and cheaper, scientists are now able to study LD at an incredibly fine resolution — down to individual base pairs. This has opened up new frontiers in precision medicine, where doctors can use a patient's unique LD patterns to predict how they might respond to specific drugs. The concept you learned today is a stepping stone to some of the most exciting areas of 21st-century biology.

Practice Problems

PROBLEM 1CONCEPTUAL
In your own words, explain the difference between linkage equilibrium and linkage disequilibrium. Why would it matter to a scientist whether two gene variants are in LD or not?
PROBLEM 2BASIC CALCULATION
In a population, allele A has a frequency of 0.60 and allele B has a frequency of 0.40. If the observed frequency of the AB haplotype is 0.30, calculate D. Is there linkage disequilibrium?
PROBLEM 3INTERMEDIATE
Using the allele frequencies from Problem 2 (f(A) = 0.60, f(a) = 0.40, f(B) = 0.40, f(b) = 0.60) and D = 0.06, calculate r². How would you describe the strength of LD?
PROBLEM 4APPLIED
A new mutation appeared in a population 50 generations ago. The recombination rate between the mutation and a nearby marker is c = 0.01 per generation. If the initial LD between them was D₀ = 0.25, what is the current LD value? How does this compare to a marker with c = 0.10?
PROBLEM 5CRITICAL THINKING
Two populations have been isolated for thousands of years. Population X shows strong LD between loci 1 and 2, while Population Y shows no LD at the same loci. Both populations have similar allele frequencies. Propose two different biological explanations for this difference. How might you test each hypothesis?

Lesson Summary

Linkage disequilibrium (LD) is the non-random association of alleles at two or more loci in a population. When alleles combine randomly, the population is in linkage equilibrium (D = 0). When certain haplotypes appear more or less often than expected, that is LD (D ≠ 0). Four main forces create LD: physical linkage (close proximity on a chromosome), natural selection, genetic drift, and population mixing.

LD is measured using the D coefficient and the r² statistic, where r² = 1 means perfect association and r² = 0 means no association. Over time, recombination breaks down LD at a rate that depends on the physical distance between loci. This concept is the foundation of genome-wide association studies (GWAS), haplotype mapping, and precision medicine — making LD one of the most important concepts in modern genomics.

Varsity Tutors • Genetics • Linkage Disequilibrium