STATISTICS & PROBABILITY • MATH

Comparing Treatments with Randomized Experiments

How scientists design controlled experiments to determine which treatment actually works best.

The Birth of Scientific Testing

For centuries, people relied on tradition, intuition, and authority to decide which treatments worked best. Doctors prescribed remedies based on ancient texts, farmers chose crops based on family tradition, and educators taught using methods passed down through generations. But this approach had a fundamental flaw: how could anyone know if these methods actually worked better than alternatives?

1747
First Controlled Trial
James Lind conducts the first randomized controlled trial on British sailors, testing six different treatments for scurvy. He discovers that citrus fruits prevent the disease, revolutionizing naval medicine.
1920s
Agricultural Revolution
Ronald Fisher develops randomization principles for agricultural experiments at Rothamsted Station, establishing the mathematical foundation for comparing crop treatments.
1940s
Medical Breakthroughs
The first large-scale randomized clinical trials test streptomycin for tuberculosis, proving that scientific experiment design can save lives on a massive scale.
1950s
Education and Psychology
Researchers begin using randomized experiments to test teaching methods and psychological interventions, expanding the approach beyond medicine and agriculture into social sciences.
Today
Digital Age Testing
Companies like Google and Facebook run thousands of A/B tests daily, using randomized experiments to optimize everything from search algorithms to social media feeds.

The central question that drove this historical evolution remains urgent today: When we have multiple treatments or interventions available, how can we determine which one actually produces the best results? Randomized experiments provide the most reliable answer to this fundamental question of cause and effect.

Core Principles of Treatment Comparison

Comparing treatments through randomized experiments rests on four fundamental principles that work together to eliminate bias and produce trustworthy results. These principles ensure that any differences we observe between groups are truly due to the treatments themselves, not other confounding factors.

1

Random Assignment

Participants are assigned to treatment groups using a random process like flipping a coin or using random number tables. This ensures that each group is similar in both measured and unmeasured characteristics.
2

Control Groups

One group receives the standard treatment or no treatment (control), while others receive new treatments. This provides a baseline for comparison to measure the effect of each intervention.
3

Large Sample Sizes

Using adequate sample sizes ensures that random assignment actually creates similar groups and provides enough data to detect real differences between treatments.
4

Blinding When Possible

Neither participants nor researchers know which treatment is being administered (double-blind). This prevents expectations from influencing behavior and measurement.
KEY TAKEAWAY
Think of randomized experiments like testing different routes to school. You can't just choose the route on sunny days versus rainy days—that would be unfair! Instead, you flip a coin each morning to decide which route to take. After a month, you compare average travel times. The coin flip ensures that each route gets tested under similar conditions, so the results truly show which route is faster, not which route got luckier with traffic or weather.

Visualizing Randomized Experiments

This diagram shows the essential flow of a randomized experiment. The study population is randomly divided into groups, ensuring each group has similar characteristics. After treatment application, researchers compare outcomes using statistical methods to determine which treatment performs better.

The power of this design lies in the random assignment step. Without randomization, groups might differ in important ways that affect outcomes. For example, if researchers let participants choose their treatment, healthier or more motivated people might select the new treatment, making it appear more effective even if it isn't. Random assignment eliminates this selection bias by ensuring that personal characteristics are distributed equally across all groups.

Statistical Framework for Comparison

Randomized experiments generate data that allows us to make statistical inferences about treatment effects. The mathematical framework provides tools to quantify differences between groups and determine whether observed differences are statistically significant or could reasonably be due to chance.

TREATMENT EFFECT
τ = μ₁ − μ₂
where τ (tau) is the treatment effect, μ₁ is the mean outcome for Treatment 1, and μ₂ is the mean outcome for Treatment 2
SAMPLE TREATMENT EFFECT
d = x̄₁ − x̄₂
where d is the observed difference in sample means, x̄₁ is the sample mean for Group 1, and x̄₂ is the sample mean for Group 2
STANDARD ERROR OF DIFFERENCE
SE = √(s₁²/n₁ + s₂²/n₂)
where SE is the standard error of the difference, s₁² and s₂² are sample variances, and n₁ and n₂ are sample sizes
T-STATISTIC
t = d/SE = (x̄₁ − x̄₂)/√(s₁²/n₁ + s₂²/n₂)
The t-statistic measures how many standard errors the observed difference is away from zero. Larger |t| values indicate stronger evidence of a real treatment effect.

The beauty of randomized experiments is that they create conditions where these statistical formulas give us unbiased estimates of treatment effects. Random assignment ensures that any systematic differences between groups are due to the treatments themselves, not confounding variables. This allows us to make causal inferences about which treatment actually causes better outcomes.

Types of Randomized Experiments

Different research questions require different experimental designs. The type of randomized experiment depends on how many treatments are being compared, whether participants can receive multiple treatments, and what kind of control group is most appropriate for the study context.

Three main types of randomized experiments: simple two-group comparison (most common), multi-group comparison for testing several treatments simultaneously, and crossover design where participants receive all treatments in random order.

Each design has specific advantages. Two-group designs are simplest to analyze and interpret. Multi-group designs allow efficient comparison of several treatments in a single study. Crossover designs are powerful because each participant serves as their own control, eliminating between-person variability and requiring smaller sample sizes.

Comparing Online vs. In-Person Tutoring

A tutoring company wants to determine whether online tutoring is as effective as traditional in-person tutoring for improving math test scores. They design a randomized experiment with 80 students who need help with algebra.

RANDOMIZED EXPERIMENT ANALYSIS
1
Step 1 — Design the ExperimentRecruit 80 students with similar baseline math scores. Use random assignment to place 40 students in online tutoring and 40 in in-person tutoring. Both groups receive the same curriculum and number of sessions.
Group 1: 40 students → Online tutoring
2
Step 2 — Collect Outcome DataAfter 8 weeks of tutoring, all students take the same standardized algebra test. Scores are measured on a 100-point scale.
Online: x̄₁ = 78.5, s₁ = 12.3, n₁ = 40
3
Step 3 — Calculate Treatment EffectCompute the difference in sample means to estimate the treatment effect: d = x̄₁ − x̄₂ = 78.5 − 82.1 = −3.6
In-person: x̄₂ = 82.1, s₂ = 11.7, n₂ = 40
4
Step 4 — Calculate Standard ErrorCompute the standard error of the difference: SE = √(s₁²/n₁ + s₂²/n₂) = √(12.3²/40 + 11.7²/40) = √(3.78 + 3.42) = √7.20 = 2.68
SE = 2.68 points
5
Step 5 — Statistical Significance TestCalculate the t-statistic to test if the difference is statistically significant: t = d/SE = −3.6/2.68 = −1.34. With 78 degrees of freedom, this gives p-value ≈ 0.184.
Not statistically significant (p > 0.05)
6
Step 6 — Interpret ResultsThe 3.6-point difference between groups is not statistically significant (p = 0.184). We cannot conclude that online tutoring is less effective than in-person tutoring. The observed difference could reasonably be due to random variation.
Conclusion: No significant difference detected

Strengths and Limitations of Randomized Experiments

Randomized experiments are considered the gold standard for establishing cause-and-effect relationships, but they also have important limitations that researchers must carefully consider when designing studies and interpreting results.

Balancing the powerful advantages of randomized experiments against their practical limitations
StrengthsLimitations
Causal inference: Random assignment eliminates confounding variables, allowing researchers to conclude that differences in outcomes are caused by the treatmentsEthical constraints: Some treatments cannot be randomly assigned for ethical reasons (e.g., smoking, dangerous behaviors). Researchers must rely on observational studies
Control of bias: Randomization prevents selection bias and ensures groups are similar in both measured and unmeasured characteristicsPractical constraints: Random assignment may not be feasible in real-world settings where people have strong preferences or institutional policies prevent randomization
Statistical power: Well-designed experiments provide precise estimates of treatment effects with known levels of uncertaintyArtificial settings: Controlled experimental conditions may not reflect real-world complexity, limiting generalizability of results
Replicability: Clear experimental protocols make studies easy to replicate and verify by independent researchersCost and time: Randomized experiments often require substantial resources and time to recruit participants and follow them over time
⚖️ KEY TAKEAWAY
Think of randomized experiments like testing recipes. If you want to know whether adding chocolate chips really makes cookies taste better, you can't just compare cookies you made yesterday (without chips) to cookies you make today (with chips). Too many things could be different—your mood, the oven temperature, the flour batch. Instead, you make a double batch of dough, flip a coin to decide which half gets chocolate chips, and bake them at the same time in the same oven. Now you can fairly conclude whether the chocolate chips made the difference.

Advanced Experimental Designs

While simple randomized experiments provide the foundation for treatment comparison, real-world research often requires more sophisticated designs to handle complex situations with multiple factors, different population subgroups, or practical constraints.

Evolution from simple to advanced experimental designs for complex research questions
Simple Randomized DesignAdvanced Design
Single factor: Compares one treatment variable (e.g., online vs. in-person tutoring)Factorial design: Tests multiple factors simultaneously (e.g., tutoring method AND session length)
Ignore subgroups: Assumes treatment effects are same for everyone in the study populationStratified randomization: Ensures equal representation of important subgroups (e.g., equal numbers of male/female participants in each treatment group)
Equal group sizes: Assigns equal numbers of participants to each treatment group for maximum statistical powerAdaptive randomization: Adjusts assignment probabilities based on accumulating evidence to assign more participants to promising treatments
Fixed sample size: Determines sample size in advance and runs to completion regardless of interim resultsSequential design: Allows early stopping if one treatment shows clear superiority, reducing time and cost

These advanced designs maintain the core principle of random assignment while adding sophisticated features to handle real-world complexity. Factorial designs can discover interactions between factors, stratified randomization ensures balanced representation of key subgroups, and adaptive designs use accumulating data to make experiments more efficient and ethical.

Practice Problems

PROBLEM 1CONCEPTUAL
A school principal wants to compare two reading programs to see which one improves student comprehension more. Explain why the principal should use random assignment rather than letting teachers choose which program to use in their classrooms.
PROBLEM 2BASIC CALCULATION
In a study comparing two pain medications, Group A (n = 30) had a mean pain reduction of 6.2 points with standard deviation 2.1, while Group B (n = 30) had a mean pain reduction of 4.8 points with standard deviation 2.3. Calculate the difference in means and the standard error of this difference.
PROBLEM 3INTERMEDIATE
A company tests three different website designs (A, B, and control) by randomly assigning visitors. Design A gets 25% of visitors, Design B gets 25%, and the control gets 50%. After one week, Design A has a 12% conversion rate (240 conversions from 2000 visitors), Design B has 15% (300 from 2000), and control has 10% (400 from 4000). Which design(s) appear most promising and what additional information would you need to draw conclusions?
PROBLEM 4APPLIED
A hospital wants to test whether a new checklist procedure reduces surgical complications compared to standard protocol. They plan to randomly assign surgeons to use either the new checklist or continue standard practice. Identify two potential problems with this design and propose solutions that maintain the principles of randomized experimentation.
PROBLEM 5CRITICAL THINKING
A researcher claims that their study proves vitamin supplements improve memory because participants who took vitamins scored 8 points higher on memory tests than those who took a placebo (p < 0.01). However, participants knew which group they were in because the vitamin pills were a different color from placebo pills. Evaluate this study's validity and explain how the lack of proper blinding might have affected the results.

Key Concepts Review

Randomized experiments provide the most reliable method for comparing treatments because random assignment eliminates confounding variables and selection bias. The key principle is that each participant has an equal chance of receiving any treatment, ensuring groups are similar in both measured and unmeasured characteristics. This allows researchers to make causal inferences about which treatment actually produces better outcomes.

The mathematical framework uses differences in sample means, standard errors, and t-statistics to quantify treatment effects and determine statistical significance. While randomized experiments have limitations—ethical constraints, artificial settings, and high costs—they remain the gold standard for establishing cause-and-effect relationships in medicine, education, technology, and social sciences.

Varsity Tutors • Statistics & Probability • Comparing Treatments with Randomized Experiments