Historical Context & Motivation
For centuries, people assumed that if two things happened together, one must cause the other. If a rooster crows every morning right before sunrise, does the rooster cause the sun to rise? Of course not — but this kind of mistake has been surprisingly common throughout the history of science. In genetics, the stakes are even higher. When scientists discover that a certain gene variant appears more often in people with a disease, it's tempting to conclude that the gene causes the disease. But that jump from observation to conclusion can lead researchers — and the public — astray.
The concepts of correlation (two things happening together in a pattern) and causation (one thing actually making the other happen) have been debated by scientists and philosophers for hundreds of years. Understanding the difference is one of the most important skills in modern genetics and medicine.
The big question that scientists still wrestle with today is this: when a genetic study finds a link between a gene and a trait or disease, how do we figure out whether the gene truly causes the trait, or whether the two just happen to appear together? This lesson will give you the tools to answer that question.
Core Principles & Definitions
Before we can tell correlation and causation apart, we need to understand exactly what each term means and why confusing them is so dangerous — especially in genetics.
Correlation
Causation
Confounding Variables
Reverse Causation
Linkage Disequilibrium
Visualizing Correlation vs. Causation
The diagram below shows three different relationships that can exist between a gene variant and a disease. Understanding these patterns is the key to interpreting genetic study results correctly.
Notice how in Scenario 1, there's a clear chain of events from gene to protein to disease. That's what real causation looks like — you can trace every step. In Scenarios 2 and 3, the gene variant and the disease appear together, so they show a correlation, but there's no direct pathway connecting them. The challenge of modern genetics is figuring out which scenario you're looking at when you find a statistical link in your data.
How Scientists Measure Correlation in Genetics
In genetic studies, researchers use math to measure how strongly two things are linked. The most common tool is the correlation coefficient, symbolized by the letter r. This number tells you how closely two variables move together.
In genetics, researchers also use a statistic called a p-value to determine whether a correlation is likely real or just due to chance. A p-value tells you the probability that you'd see a result this strong even if there were no real connection between the gene and the trait.
Types of Genetic Studies: From Correlation to Causation
Not all genetic studies are created equal. Some can only find correlations, while others can provide evidence of causation. Understanding the different study types helps you judge how much weight to give a finding.
The key lesson from this ladder is that no single study proves causation. Scientists build a case by combining multiple types of evidence. A GWAS might identify a suspicious gene variant. Then family studies check whether the variant tracks with the disease across generations. Then a lab experiment using CRISPR (a gene-editing tool) might disable that gene in mice to see if the disease appears. When all these lines of evidence agree, scientists grow confident that the relationship is truly causal.
Worked Example: Evaluating a Genetic Study
Let's walk through a realistic scenario. Imagine you're a scientist who just read a study with the following finding: "People who carry the variant rs12345 in the FTO gene are, on average, 3 kg heavier than people without the variant (p = 0.001, OR = 1.4)." Does this gene variant cause weight gain?
Correlation vs. Causation: Side-by-Side Comparison
Let's put correlation and causation side by side so you can clearly see how they differ. Understanding these distinctions will help you evaluate any genetic study you encounter.
| Feature | Correlation | Causation |
|---|---|---|
| Definition | Two things happen together in a pattern | One thing directly makes the other happen |
| Direction | Does not tell you which causes which | Specifies A → B (direction is known) |
| Evidence needed | Statistical test showing a significant pattern (e.g., p < 0.05) | Controlled experiments, biological mechanism, and replication |
| Study type | Observational (GWAS, surveys) | Experimental (gene knockouts, CRISPR) |
| Confounders | Can be explained by hidden third variables | Controlled for through experimental design |
| Genetics example | "Gene variant X is found more often in people with asthma" | "Mutation in CFTR gene produces a faulty protein → thick mucus → cystic fibrosis" |
| Can prove alone? | No — correlation alone never proves a cause | Yes — when mechanism + experiment + replication all agree |
Connection to Advanced Genetics & Genomics
As you continue studying genetics, you'll encounter more advanced methods that blur the line between observational and experimental studies. One important technique is called Mendelian Randomization. This clever approach uses the fact that your gene variants are randomly assigned at conception (like a natural experiment) to test causal claims without actually doing a lab experiment.
| Feature | Standard GWAS (This Lesson) | Mendelian Randomization (Advanced) |
|---|---|---|
| Goal | Find correlations between gene variants and traits | Use gene variants as natural experiments to test causation |
| How it works | Compares genetic markers across large populations | Uses genetic variants as "instruments" to mimic random assignment |
| Can establish causation? | No — only correlation | Potentially yes — if assumptions are met |
| Handles confounders? | Must control for them statistically | Naturally avoids many confounders because genes are randomly inherited |
| Example question | "Is vitamin D level correlated with depression?" | "Do genes that cause higher vitamin D levels also lead to lower depression rates?" |
Another advanced topic you'll encounter is polygenic risk scores. These scores combine the effects of hundreds or thousands of small genetic correlations to predict your overall risk for a disease. While each individual correlation might be weak, together they can be quite powerful. However, because they're built from correlations, they come with all the limitations we've discussed — confounders, population differences, and the inability to prove that any single variant is truly causal.
Practice Problems
Lesson Summary
Correlation means two variables change together in a pattern, while causation means one directly produces a change in the other. In genetic studies, GWAS and other observational studies can identify correlations between gene variants and traits, but they cannot prove causation on their own. Hidden factors called confounding variables (like shared ancestry or environmental differences) and linkage disequilibrium (where nearby genes are inherited together) can create false associations that look meaningful but aren't. Statistical tools like the p-value and odds ratio measure the strength of a correlation but cannot, by themselves, establish a cause.
To move from correlation toward causation, scientists climb a ladder of evidence: from observational studies, to family and twin studies, to functional experiments like gene knockouts and CRISPR editing, and finally to techniques like Mendelian Randomization that use natural genetic variation as a built-in experiment. True causal proof in genetics requires converging evidence — multiple independent lines of study all pointing in the same direction. Remember: correlation is the starting point, not the finish line.