PSYCHOLOGY • FOUNDATIONS & RESEARCH METHODS

Replication & Research Practices — I can explain replication and why questionable research practices can distort scientific conclusions.

Understanding why repeating studies matters and how flawed practices can lead science astray.

Historical Context & Motivation

Science depends on the idea that results should be repeatable. If one researcher finds that a certain therapy reduces anxiety, other researchers should be able to run the same study and get a similar result. This principle—called replication—has been a cornerstone of the scientific method for centuries. Yet psychology, along with other sciences, has faced serious challenges when famous findings could not be reproduced.

Throughout the twentieth century, psychological research expanded rapidly. New journals published thousands of studies, and many influential findings shaped textbooks and public policy. However, by the early 2010s, researchers began to notice a troubling pattern: when they tried to repeat classic experiments, many results failed to replicate. This sparked what became known as the replication crisis, shaking confidence in psychological science and forcing the field to examine its own practices.

1620
Francis Bacon & Empiricism
Bacon argued that knowledge must be built on repeated observation and experimentation, laying the groundwork for modern scientific methodology.
1962
Kuhn's The Structure of Scientific Revolutions
Thomas Kuhn challenged the idea that science progresses smoothly, showing that paradigms shift when accumulating anomalies force the field to rethink its assumptions.
2011
Bem's 'Feeling the Future'
Daryl Bem published a study claiming evidence for precognition—that people could sense future events. Peer-reviewed and published in a top journal, the study ignited fierce debate about research methods when replications failed.
2015
Open Science Collaboration
A large team of researchers attempted to replicate 100 published psychology studies. Only about 36% produced results consistent with the originals, drawing worldwide attention to the replication crisis.
2023
Pre-Registration & Open Science Norms
Major psychology journals now encourage or require researchers to pre-register their hypotheses and share their data, marking a cultural shift toward transparency.

This history raises a central question: Why do some research findings fail to replicate, and what questionable practices contribute to unreliable conclusions? Understanding the answers is essential for evaluating any scientific claim you encounter—in class, in the news, or on social media.

Core Principles & Definitions

Before diving deeper, let's establish the key ideas that underpin this topic. Replication is not just about "doing a study again." It involves specific standards, and the problems that undermine it have specific names.

1

Replication

The process of repeating a study using the same methods to see if the original findings can be consistently reproduced. Direct replication uses identical procedures; conceptual replication tests the same idea with different methods.
2

Questionable Research Practices (QRPs)

Shortcuts or grey-area decisions researchers make that inflate the chances of finding a statistically significant result, even when no real effect exists. Examples include p-hacking and selective reporting.
3

Publication Bias

The tendency of journals to publish studies that find positive or significant results while rejecting studies that find no effect. This creates a distorted picture of reality in the published literature.
4

P-Hacking

Manipulating data analysis—such as testing many variables, removing outliers, or stopping data collection early—until a result reaches the p < 0.05 threshold for statistical significance.
5

Pre-Registration

A safeguard where researchers publicly record their hypotheses, methods, and analysis plans before collecting data. This prevents after-the-fact changes that inflate false-positive rates.
KEY TAKEAWAY
Think of replication like a cooking recipe. If a chef claims their dish tastes amazing, you should be able to follow the exact same recipe and get a similar result. If only that one chef can make it work—and nobody else can—then something is off. Either the recipe is wrong, they left out a secret step, or the result was a lucky accident. Replication in science works the same way: if a finding is real, other researchers should be able to reproduce it.

The Replication Process — A Visual Overview

The diagram below illustrates how replication is supposed to work—and what happens when it breaks down. Follow the flow from an original study through replication attempts, and notice the two possible outcomes: confirmation or failure. The key insight is that a single study is never enough to establish a scientific fact. Only through repeated testing can we build confidence.

This flowchart shows the ideal path from original study to replication. When replications succeed (left branch), confidence in the finding grows. When they fail (right branch), researchers must investigate whether questionable research practices, small sample sizes, or genuine differences in conditions caused the discrepancy.

Notice that a failed replication does not automatically mean the original researchers did something wrong. Sometimes differences in the participant population, cultural context, or small procedural variations can affect outcomes. However, when many replications fail, it strongly suggests the original effect was either exaggerated or did not exist in the first place.

How Questionable Research Practices Distort Conclusions

To understand why findings fail to replicate, you need to understand the specific practices that can inflate or distort results. These are not always intentional fraud—many questionable research practices (QRPs) happen because researchers face pressure to publish exciting results, and subtle analytical choices can tip the scales toward statistical significance.

The p-Value and Why It Matters

In psychology research, a p-value tells you the probability of getting your result (or something more extreme) if there is actually no real effect. The conventional threshold is p < 0.05, meaning there is less than a 5% chance the result occurred by random chance alone. However, this threshold can be gamed.

SIGNIFICANCE THRESHOLD
p < 0.05 → "statistically significant"
p = probability of observing the data if the null hypothesis (no effect) is true. A false positive occurs when a result crosses this threshold by chance—expected about 5% of the time even with no real effect.

Common Questionable Research Practices

  • P-hacking: Running many different statistical tests, removing data points, or adding new variables until something reaches p < 0.05. Imagine flipping a coin 20 times and only reporting the sequence that looks non-random.
  • HARKing (Hypothesizing After Results are Known): Presenting an unexpected finding as though it was your original prediction. This makes exploratory results look like confirmed hypotheses.
  • Selective reporting: Only publishing the analyses that support your hypothesis and hiding the ones that don't. This is also called the "file drawer problem"—negative results stay hidden.
  • Small sample sizes: Using too few participants, which makes results unstable and easily influenced by a few unusual data points. Small studies produce dramatic-looking effects that often shrink or vanish in larger studies.
  • Flexible stopping rules: Checking results repeatedly as data comes in and stopping data collection once you hit p < 0.05, rather than collecting a pre-determined sample size.
⚠️ Why Does This Happen?
Researchers face enormous pressure to publish statistically significant results. University tenure decisions, funding, and career advancement often depend on a researcher's publication record. Journals historically preferred exciting, positive findings over null results. This system created incentives—often unconscious—for researchers to use QRPs. Understanding these systemic pressures helps explain why the replication crisis is not just about individual bad actors, but about the structure of the research ecosystem itself.

Mapping Questionable Research Practices

The diagram below organizes questionable research practices by the stage of research at which they occur. Notice that QRPs can creep in at every phase of a study—from designing the hypothesis, to collecting data, to analyzing results, to writing up the paper. This is why comprehensive safeguards are necessary.

This diagram maps QRPs to the four stages of research. Each stage presents opportunities for bias to creep in. The bottom panels show open science solutions (left) and the cultural shifts (right) needed to address these problems.

The spectrum below shows how research practices range from fully transparent and ethical at one end to outright fraud at the other. Most questionable research practices fall in the gray area in between—they are not fabrication, but they are not rigorous science either.

Spectrum of Research Integrity
Best Practices
Acceptable Flexibility
Questionable Practices (QRPs)
Misconduct / Fraud
Pre-registered studies
P-hacking zone
Data fabrication
TransparentFraudulent

Worked Example — Spotting QRPs in a Study

Let's walk through a hypothetical scenario to practice identifying questionable research practices. Imagine a researcher wants to show that listening to classical music improves test scores.

Scenario: Does Classical Music Boost Test Scores?
1
Step 1 — Examine the HypothesisDr. Martinez hypothesizes that students who listen to Mozart before an exam will score higher than a control group. She does not pre-register this hypothesis. After collecting data, she finds no significant difference for Mozart specifically—but students who listened to any classical music scored slightly higher.
Red flag: The hypothesis shifted after seeing data (potential HARKing).
2
Step 2 — Evaluate Sample SizeDr. Martinez tested only 30 students total—15 in each group. With such a small sample, a few unusually high or low scorers could dramatically shift the average. A proper power analysis would have suggested she needed at least 60–80 participants per group to detect a small-to-medium effect reliably.
Red flag: The study is underpowered, making results unstable.
3
Step 3 — Check the AnalysisIn her paper, Dr. Martinez reports testing the effect on overall test score, math subscore, reading subscore, and memory subscore. She only reports the one comparison that was significant: the reading subscore (p = 0.04). The other three comparisons were non-significant but are barely mentioned.
Red flag: This is selective reporting and p-hacking. Testing four outcomes without correction inflates the false-positive rate.
4
Step 4 — Assess the Multiple Comparisons ProblemWhen you test four separate outcomes, the probability that at least one will reach p < 0.05 purely by chance is approximately 1 − (0.95)4 ≈ 0.185, or about 18.5%. That is far higher than the 5% threshold researchers claim. Without correcting for multiple comparisons (such as using a Bonferroni correction), the significant result is likely a false positive.
The adjusted threshold should be p < 0.05 ÷ 4 = 0.0125 per comparison. The finding of p = 0.04 no longer qualifies.
5
Step 5 — Predict Replication OutcomeGiven the small sample size, HARKing, and p-hacking, this study is unlikely to replicate. A well-powered, pre-registered replication with 80+ participants per group, testing only the pre-specified outcome, would probably find no significant effect of classical music on reading scores.
Conclusion: Multiple QRPs stacked together produced a publishable but unreliable finding.

Safeguards Against Questionable Practices

The replication crisis did not just expose problems—it also sparked a movement toward better science. Researchers, journals, and institutions have developed several safeguards to improve the reliability of published findings. The table below compares the old way of doing things with the new, more transparent approaches.

Traditional vs. Open Science Practices
FeatureTraditional PracticeOpen Science Reform
HypothesisCan be changed after seeing dataPre-registered before data collection
DataKept private by researchersShared publicly in open repositories
Analysis PlanFlexible; chosen after seeing resultsSpecified in advance; deviations disclosed
Sample SizeOften small; sometimes determined by convenienceDetermined by power analysis before study begins
PublicationOnly significant results publishedRegistered Reports accepted regardless of results
ReplicationRarely attempted; not rewardedEncouraged and published in dedicated journals
KEY TAKEAWAY
Think of pre-registration like calling your shot in basketball before you shoot—you announce what you expect to happen before you see the outcome. If you make the basket, everyone knows you weren't just claiming after the fact that you meant to do that. Pre-registration prevents researchers from rewriting their predictions to match their results, which is one of the most powerful safeguards against HARKing and p-hacking.

Connection to Advanced Research & Real-World Impact

The replication crisis is not just an academic concern—it has real consequences for medicine, education, and public policy. When psychological findings cannot be replicated, interventions based on those findings may waste resources or even cause harm. For example, some widely-adopted educational programs and therapeutic techniques were built on research that later failed to replicate.

Introductory vs. Advanced Research Methods
ConceptIntroductory LevelAdvanced Level
ReplicationRepeating a study to check if results holdMeta-analysis combining dozens of replications to estimate true effect sizes using statistical weighting
Statistical Significancep < 0.05 threshold for claiming a real effectBayesian analysis, confidence intervals, and effect-size estimation replacing binary significant/not-significant decisions
Bias DetectionRecognizing p-hacking and HARKing in individual studiesFunnel plots, p-curve analysis, and statistical forensics to detect bias across entire literatures
Open SciencePre-registration and data sharingReproducible computational pipelines, adversarial collaborations, and many-labs studies with 30+ simultaneous replications

As you continue studying psychology, you'll encounter concepts like meta-analysis (statistically combining results from many studies to get a clearer picture of an effect) and effect size (a measure of how large or practically meaningful an effect is, beyond just whether it is statistically significant). These tools help researchers move beyond the simple yes-or-no question of significance and toward a more nuanced understanding of how strong and reliable psychological findings really are.

💡 Why This Matters to You
Every time you read a headline claiming "scientists prove that..." or "a new study shows...," you now have the tools to think critically. Ask yourself: Has this been replicated? How large was the sample? Were the researchers transparent about their methods? Understanding replication and QRPs makes you a more informed consumer of scientific information—a skill that matters far beyond psychology class.

Practice Problems

PROBLEM 1CONCEPTUAL
In your own words, explain what replication means in psychological research and why it is considered essential for establishing scientific knowledge.
PROBLEM 2BASIC CALCULATION
A researcher tests 20 different outcome variables at the p < 0.05 significance level. If there is truly no effect, approximately how many of these tests would you expect to produce a statistically significant result by chance alone?
PROBLEM 3INTERMEDIATE
A researcher conducted a study with 25 participants per group and found a significant result (p = 0.03). A second team attempted a direct replication with 200 participants per group and found no significant effect (p = 0.42). Which study should you trust more, and why? Identify at least two factors that support your reasoning.
PROBLEM 4APPLIED
A school district reads a published study claiming that a 10-minute mindfulness exercise before class improves standardized test scores by 15%. They are considering spending $500,000 to implement the program district-wide. Based on what you know about replication and QRPs, what questions should the school board ask before making this decision?
PROBLEM 5CRITICAL THINKING
Some researchers argue that the replication crisis is actually a sign that science is working correctly—catching its own mistakes—while others argue it reveals a fundamental structural problem. Evaluate both perspectives and explain which you find more persuasive, using at least three specific concepts from this lesson to support your argument.

Lesson Summary

Replication—the process of repeating studies to verify results—is the foundation of trustworthy science. When findings are replicated by independent researchers, our confidence in them grows. The replication crisis revealed that many published psychology findings could not be reproduced, with the landmark 2015 Open Science Collaboration finding that only about 36% of 100 studies replicated successfully. This crisis was driven not primarily by fraud but by questionable research practices (QRPs) such as p-hacking, HARKing, selective reporting, and publication bias.

The field has responded with powerful safeguards including pre-registration (declaring hypotheses before collecting data), open data sharing, larger sample sizes determined by power analysis, and registered reports that are accepted for publication based on methods rather than results. As a critical consumer of research, you should always ask whether a finding has been replicated, whether the study was pre-registered, and whether the sample size was adequate. These skills help you evaluate not just psychology research, but any scientific claim you encounter in everyday life.

Varsity Tutors • Psychology • Replication & Research Practices