Historical Context & Motivation
Science depends on the idea that results should be repeatable. If one researcher finds that a certain therapy reduces anxiety, other researchers should be able to run the same study and get a similar result. This principle—called replication—has been a cornerstone of the scientific method for centuries. Yet psychology, along with other sciences, has faced serious challenges when famous findings could not be reproduced.
Throughout the twentieth century, psychological research expanded rapidly. New journals published thousands of studies, and many influential findings shaped textbooks and public policy. However, by the early 2010s, researchers began to notice a troubling pattern: when they tried to repeat classic experiments, many results failed to replicate. This sparked what became known as the replication crisis, shaking confidence in psychological science and forcing the field to examine its own practices.
This history raises a central question: Why do some research findings fail to replicate, and what questionable practices contribute to unreliable conclusions? Understanding the answers is essential for evaluating any scientific claim you encounter—in class, in the news, or on social media.
Core Principles & Definitions
Before diving deeper, let's establish the key ideas that underpin this topic. Replication is not just about "doing a study again." It involves specific standards, and the problems that undermine it have specific names.
Replication
Questionable Research Practices (QRPs)
Publication Bias
P-Hacking
Pre-Registration
The Replication Process — A Visual Overview
The diagram below illustrates how replication is supposed to work—and what happens when it breaks down. Follow the flow from an original study through replication attempts, and notice the two possible outcomes: confirmation or failure. The key insight is that a single study is never enough to establish a scientific fact. Only through repeated testing can we build confidence.
Notice that a failed replication does not automatically mean the original researchers did something wrong. Sometimes differences in the participant population, cultural context, or small procedural variations can affect outcomes. However, when many replications fail, it strongly suggests the original effect was either exaggerated or did not exist in the first place.
How Questionable Research Practices Distort Conclusions
To understand why findings fail to replicate, you need to understand the specific practices that can inflate or distort results. These are not always intentional fraud—many questionable research practices (QRPs) happen because researchers face pressure to publish exciting results, and subtle analytical choices can tip the scales toward statistical significance.
The p-Value and Why It Matters
In psychology research, a p-value tells you the probability of getting your result (or something more extreme) if there is actually no real effect. The conventional threshold is p < 0.05, meaning there is less than a 5% chance the result occurred by random chance alone. However, this threshold can be gamed.
Common Questionable Research Practices
- P-hacking: Running many different statistical tests, removing data points, or adding new variables until something reaches p < 0.05. Imagine flipping a coin 20 times and only reporting the sequence that looks non-random.
- HARKing (Hypothesizing After Results are Known): Presenting an unexpected finding as though it was your original prediction. This makes exploratory results look like confirmed hypotheses.
- Selective reporting: Only publishing the analyses that support your hypothesis and hiding the ones that don't. This is also called the "file drawer problem"—negative results stay hidden.
- Small sample sizes: Using too few participants, which makes results unstable and easily influenced by a few unusual data points. Small studies produce dramatic-looking effects that often shrink or vanish in larger studies.
- Flexible stopping rules: Checking results repeatedly as data comes in and stopping data collection once you hit p < 0.05, rather than collecting a pre-determined sample size.
Mapping Questionable Research Practices
The diagram below organizes questionable research practices by the stage of research at which they occur. Notice that QRPs can creep in at every phase of a study—from designing the hypothesis, to collecting data, to analyzing results, to writing up the paper. This is why comprehensive safeguards are necessary.
The spectrum below shows how research practices range from fully transparent and ethical at one end to outright fraud at the other. Most questionable research practices fall in the gray area in between—they are not fabrication, but they are not rigorous science either.
Worked Example — Spotting QRPs in a Study
Let's walk through a hypothetical scenario to practice identifying questionable research practices. Imagine a researcher wants to show that listening to classical music improves test scores.
Safeguards Against Questionable Practices
The replication crisis did not just expose problems—it also sparked a movement toward better science. Researchers, journals, and institutions have developed several safeguards to improve the reliability of published findings. The table below compares the old way of doing things with the new, more transparent approaches.
| Feature | Traditional Practice | Open Science Reform |
|---|---|---|
| Hypothesis | Can be changed after seeing data | Pre-registered before data collection |
| Data | Kept private by researchers | Shared publicly in open repositories |
| Analysis Plan | Flexible; chosen after seeing results | Specified in advance; deviations disclosed |
| Sample Size | Often small; sometimes determined by convenience | Determined by power analysis before study begins |
| Publication | Only significant results published | Registered Reports accepted regardless of results |
| Replication | Rarely attempted; not rewarded | Encouraged and published in dedicated journals |
Connection to Advanced Research & Real-World Impact
The replication crisis is not just an academic concern—it has real consequences for medicine, education, and public policy. When psychological findings cannot be replicated, interventions based on those findings may waste resources or even cause harm. For example, some widely-adopted educational programs and therapeutic techniques were built on research that later failed to replicate.
| Concept | Introductory Level | Advanced Level |
|---|---|---|
| Replication | Repeating a study to check if results hold | Meta-analysis combining dozens of replications to estimate true effect sizes using statistical weighting |
| Statistical Significance | p < 0.05 threshold for claiming a real effect | Bayesian analysis, confidence intervals, and effect-size estimation replacing binary significant/not-significant decisions |
| Bias Detection | Recognizing p-hacking and HARKing in individual studies | Funnel plots, p-curve analysis, and statistical forensics to detect bias across entire literatures |
| Open Science | Pre-registration and data sharing | Reproducible computational pipelines, adversarial collaborations, and many-labs studies with 30+ simultaneous replications |
As you continue studying psychology, you'll encounter concepts like meta-analysis (statistically combining results from many studies to get a clearer picture of an effect) and effect size (a measure of how large or practically meaningful an effect is, beyond just whether it is statistically significant). These tools help researchers move beyond the simple yes-or-no question of significance and toward a more nuanced understanding of how strong and reliable psychological findings really are.
Practice Problems
Lesson Summary
Replication—the process of repeating studies to verify results—is the foundation of trustworthy science. When findings are replicated by independent researchers, our confidence in them grows. The replication crisis revealed that many published psychology findings could not be reproduced, with the landmark 2015 Open Science Collaboration finding that only about 36% of 100 studies replicated successfully. This crisis was driven not primarily by fraud but by questionable research practices (QRPs) such as p-hacking, HARKing, selective reporting, and publication bias.
The field has responded with powerful safeguards including pre-registration (declaring hypotheses before collecting data), open data sharing, larger sample sizes determined by power analysis, and registered reports that are accepted for publication based on methods rather than results. As a critical consumer of research, you should always ask whether a finding has been replicated, whether the study was pre-registered, and whether the sample size was adequate. These skills help you evaluate not just psychology research, but any scientific claim you encounter in everyday life.