PSYCHOLOGY • PSYCHOLOGICAL DISORDERS & TREATMENT

Evaluating Treatment Claims — I can evaluate claims about mental health treatments using evidence standards (placebo, controlled trials) at a basic level.

Learn to separate science-backed treatments from hype using placebos, controlled trials, and critical thinking.

Historical Context & Motivation

Throughout history, people have claimed to cure mental illness with everything from drilling holes in the skull to spinning patients in chairs. Before the rise of modern science, there was no reliable way to tell whether a treatment actually worked or whether people simply felt better because they believed in the cure. The development of evidence-based standards for evaluating treatments transformed psychology from guesswork into a discipline grounded in data. Understanding this history helps you see why we need rigorous testing before accepting any mental health treatment claim.

1747
First Controlled Trial
James Lind tested six treatments for scurvy on groups of sailors, keeping conditions similar across groups. This is considered one of the earliest controlled experiments in medical history.
1799
Haygarth's Placebo Demonstration
John Haygarth showed that fake metal rods ("tractors") relieved pain just as well as the supposedly magical ones, revealing the power of placebo effects — improvement from belief alone.
1943
The Double-Blind Method
Researchers began using double-blind procedures in drug trials, where neither participants nor researchers knew who received the real treatment. This reduced bias dramatically.
1952
Eysenck Challenges Psychotherapy
Hans Eysenck published data suggesting that psychotherapy was no more effective than no treatment at all. This controversial claim pushed the field to demand controlled evidence for therapy's effectiveness.
1998
Evidence-Based Practice Movement
The American Psychological Association formally endorsed evidence-based practice, requiring that treatments be supported by rigorous research, including randomized controlled trials.

This history leads to a central question that drives our lesson: How can we tell the difference between a treatment that truly works and one that only seems to work? To answer that, we need to understand placebos, controlled trials, and the standards scientists use to evaluate mental health treatments.

Core Principles & Definitions

Before you can evaluate any treatment claim, you need to understand a few foundational ideas. These principles form the toolkit that researchers — and informed consumers — use to separate real cures from false promises. Each principle addresses a specific way that our thinking can be tricked into believing something works when it actually does not.

1

Placebo Effect

A placebo is an inactive treatment (like a sugar pill) that has no therapeutic ingredient. The placebo effect occurs when people feel better simply because they expect the treatment to help. This effect is real and measurable — the brain can actually reduce pain or anxiety based on belief alone.
2

Control Group

A control group is a group of participants who do NOT receive the treatment being tested. They might receive a placebo or no treatment at all. By comparing the treatment group to the control group, researchers can isolate whether the treatment itself caused improvement.
3

Random Assignment

In a well-designed study, participants are placed into the treatment or control group by chance — like flipping a coin. Random assignment ensures that the groups are similar at the start, so any differences at the end can be attributed to the treatment rather than pre-existing differences.
4

Double-Blind Procedure

In a double-blind study, neither the participants nor the researchers interacting with them know who is receiving the real treatment. This prevents both groups from unconsciously influencing the results through expectations or body language.
5

Replication

A single study is never enough. Replication means that other researchers repeat the study and get similar results. If a treatment only "works" in one study but fails in ten others, we should be skeptical of the original claim.
KEY TAKEAWAY
Think of evaluating a treatment like judging whether a lucky charm actually helps you score better on tests. If you wear the charm and do well, was it the charm — or did you also study more, sleep better, or feel more confident? A controlled trial is like having a twin take the same test without the charm, after the same preparation, to see if the charm made any real difference.

Visual Explanation — Anatomy of a Controlled Trial

The diagram below shows how a randomized controlled trial (RCT) is designed from start to finish. Follow the flow from participant recruitment through random assignment and finally to the comparison of outcomes. This structure is what makes it possible to draw causal conclusions about whether a treatment truly helps.

This flowchart shows the structure of a randomized controlled trial. Participants are randomly assigned to either the treatment group (green) or control group (gold). After the study period, researchers compare results to determine if the treatment caused real improvement beyond the placebo effect.

Notice several key features in this design. First, random assignment ensures that the two groups start out roughly equal — similar ages, symptom severity, and backgrounds. Second, both groups go through the same process; the only difference is whether the treatment is real. Third, the double-blind note at the bottom reminds us that keeping everyone "in the dark" about who gets what prevents bias from creeping in. When all of these elements are in place, we can be much more confident that any improvement in the treatment group is due to the treatment itself.

How Evidence Standards Work

While evaluating treatment claims is not primarily a mathematical exercise, there are some key concepts from research methodology that help you understand how scientists decide if a treatment works. These ideas involve comparing group outcomes and understanding what counts as a meaningful difference.

The Logic of Comparison

The fundamental logic is simple: if 70 out of 100 people improve with the real treatment, but only 40 out of 100 improve with the placebo, the difference of 30 people suggests the treatment has a genuine effect beyond placebo. Researchers use statistical significance to determine whether a difference this large is likely due to the treatment or could have happened by random chance alone.

IMPROVEMENT RATE COMPARISON
Treatment Effect = Improvement Rate (Treatment Group) − Improvement Rate (Control Group)
If Treatment Group improvement = 70% and Control Group improvement = 40%, then the treatment effect = 70% − 40% = 30 percentage points. A treatment effect near zero suggests the treatment is no better than placebo.

Effect Size — How Much Does It Help?

Beyond asking "does it work at all," researchers want to know how much it helps. An effect size measures the magnitude of improvement. A treatment might produce a statistically significant result but only help people a tiny amount — not enough to matter in their daily lives. Effect sizes are commonly rated as small (0.2), medium (0.5), or large (0.8).

COHEN'S d (SIMPLIFIED)
d = (Mean of Treatment Group − Mean of Control Group) ÷ Pooled Standard Deviation
A d value of 0.8 or higher is considered a large effect, meaning the treatment produces a noticeable, meaningful improvement compared to the control group.

The Hierarchy of Evidence

Not all evidence is created equal. A personal testimonial — "This crystal cured my anxiety!" — is the weakest form of evidence. A single case study is slightly better. A controlled trial is much stronger. And a meta-analysis, which combines results from many controlled trials, sits at the top of the evidence hierarchy. When evaluating any treatment claim, you should ask: what level of evidence supports it?

The Hierarchy of Evidence — From Weak to Strong

The diagram below illustrates the hierarchy of evidence as a pyramid. The base contains the most common but weakest forms of evidence, while the top contains the strongest but rarest forms. When someone makes a claim about a mental health treatment, identifying where their evidence falls on this pyramid tells you how seriously to take it.

The evidence pyramid ranges from anecdotes and testimonials at the bottom (weakest) to meta-analyses at the top (strongest). The higher on the pyramid, the more confidence we can have that a treatment truly works.

At the base of the pyramid, anecdotes are stories from individuals about their personal experiences. These are the most common form of "evidence" you will encounter on social media, in advertisements, and in conversations. While personal stories can be compelling, they tell us nothing about whether the treatment caused the improvement. The person might have gotten better on their own, experienced a placebo effect, or changed other things in their life at the same time.

Moving upward, case studies provide detailed observations of one or a few patients, and expert opinions draw on clinical experience. These are more informative than anecdotes, but they still lack the control groups and random assignment needed to rule out alternative explanations. Controlled experiments represent a major leap in quality because they directly compare a treatment group to a control group. At the very top, meta-analyses pool data from many controlled studies, giving us the most reliable picture of whether a treatment works across different populations and settings.

Worked Example — Evaluating a Treatment Claim

Let's walk through a realistic scenario. Imagine you see the following advertisement online: "New supplement CalmMind reduces anxiety by 80%! Thousands of satisfied customers!" How would you evaluate this claim using the evidence standards we have learned?

Evaluating the "CalmMind" Supplement Claim
1
Step 1 — Identify the ClaimThe claim states that CalmMind reduces anxiety by 80%. This is a specific, testable assertion about a mental health treatment. We need to ask: what evidence is provided?
Claim: CalmMind reduces anxiety by 80%.
2
Step 2 — Check the Evidence TypeThe ad cites "thousands of satisfied customers." On our evidence hierarchy, this is an anecdote/testimonial — the weakest level. Customer satisfaction reports are self-selected (unhappy customers do not write in) and do not include any comparison group.
Evidence level: Anecdotal (bottom of the pyramid).
3
Step 3 — Look for a Control GroupAsk: was there a control group that received a placebo instead of CalmMind? If not, we cannot know whether the 80% improvement was caused by the supplement or by the placebo effect, natural recovery, or other factors. In this case, no control group is mentioned.
No control group identified — major red flag.
4
Step 4 — Check for Random Assignment & BlindingEven if there were a study, we would ask whether participants were randomly assigned to groups and whether the study was double-blind. Without random assignment, people who chose to take CalmMind might be different from those who did not (maybe they were more motivated to change). Without blinding, expectations could inflate results.
No mention of randomization or blinding.
5
Step 5 — Ask About ReplicationHas any independent research team tested CalmMind and found similar results? If only the company selling the product has data, there is a major conflict of interest. Independent replication is essential before we should trust any treatment claim.
Conclusion: This claim lacks adequate evidence. The "80% reduction" is not supported by controlled, replicated research and should be treated with strong skepticism.

Red Flags vs. Green Flags in Treatment Claims

When you encounter a claim about any mental health treatment — whether it is a new therapy app, a medication, a supplement, or an alternative practice — certain features should raise or lower your confidence. The table below summarizes the most important warning signs and encouraging signs to watch for.

Quick-reference guide for evaluating mental health treatment claims
Feature🚩 Red Flag (Be Skeptical)✅ Green Flag (More Trustworthy)
Evidence TypeRelies on testimonials, celebrity endorsements, or "ancient wisdom"Cites peer-reviewed studies, randomized controlled trials, or meta-analyses
Control GroupNo comparison group; everyone in the study received the treatmentIncludes a placebo or waitlist control group with random assignment
Claims"Cures everything," "100% effective," "no side effects"Specific, measured outcomes with acknowledged limitations
SourceThe company selling the product funded the only studyMultiple independent research teams have replicated results
TransparencyVague about methods; hides data; discourages questionsPublishes full methodology; data available; welcomes scrutiny
Professional SupportRejected or ignored by mainstream psychology and psychiatryEndorsed by professional organizations like the APA
KEY TAKEAWAY
Think of evaluating treatment claims like checking reviews before buying a product online. A product with thousands of verified purchases and detailed reviews from different types of buyers (independent replication) is far more trustworthy than one with five glowing reviews that all sound the same (biased testimonials). In psychology, peer-reviewed, replicated research is the equivalent of those verified, diverse reviews.

Connection to Advanced Research Methods

The basic tools you have learned — placebos, control groups, random assignment, blinding, and replication — are the foundation of evidence-based psychology. As you advance, you will encounter more sophisticated methods that build on these principles. The table below previews how your current knowledge connects to these advanced concepts.

How basic evidence standards connect to advanced research methods
Basic Concept (This Lesson)Advanced Extension
Placebo effect — belief causes improvementNocebo effect — negative expectations cause worsening symptoms, even with inactive substances
Single RCT with control groupMeta-analysis — statistically combines results from dozens of RCTs to find overall treatment effects
Effect size (small, medium, large)Clinical significance — determining whether a statistically significant effect is large enough to matter in real life
Random assignment to groupsStratified randomization — ensuring groups are balanced on key variables like age, gender, and severity
Double-blind procedureActive placebo — a placebo that mimics side effects of the real drug to maintain blinding more effectively

Understanding these basic standards is not just academic — it is a life skill. Whether you are reading a news article about a new antidepressant, hearing about a friend's experience with a therapy app, or seeing an ad for a "miracle cure," the critical thinking tools from this lesson will help you make informed decisions. As psychology continues to grow as a science, the demand for evidence-based thinking will only increase.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain in your own words why a testimonial ("This treatment changed my life!") is not strong evidence that a treatment works. What factors other than the treatment could explain the person's improvement?
PROBLEM 2BASIC CALCULATION
In a study of a new therapy for depression, 60 out of 80 people in the treatment group showed improvement, while 35 out of 80 people in the placebo control group showed improvement. Calculate the improvement rate for each group and the treatment effect (difference in improvement rates).
PROBLEM 3INTERMEDIATE
A researcher wants to test whether a meditation app reduces test anxiety in high school students. She asks for 40 volunteers, lets them choose whether they want to use the app or not, and then compares anxiety scores after four weeks. Identify at least two major flaws in this study design and explain how each could be fixed.
PROBLEM 4APPLIED
You read a news headline: "Groundbreaking Study Shows New Drug Cuts PTSD Symptoms in Half." The article mentions that the study was funded by the drug's manufacturer, involved 30 participants with no control group, and has not been replicated. Using the evidence hierarchy and red flag/green flag framework, write a brief evaluation of this claim. Would you recommend this treatment to a friend? Why or why not?
PROBLEM 5CRITICAL THINKING
Some critics argue that randomized controlled trials are not always the best way to evaluate psychotherapy because it is impossible to truly "blind" someone to whether they are receiving therapy (unlike a pill, you know if you are talking to a therapist). Does this limitation mean we should abandon controlled trials for therapy? How might researchers address this challenge while still maintaining scientific rigor? Propose at least two solutions.

Lesson Summary

Evaluating mental health treatment claims requires understanding several key evidence standards. The placebo effect shows that people can improve simply because they believe a treatment will help, which is why every credible study needs a control group for comparison. Random assignment ensures that treatment and control groups start out equivalent, while the double-blind procedure prevents expectations from biasing results. A single study is never conclusive; replication by independent researchers is essential before a treatment can be considered well-supported.

The hierarchy of evidence ranks evidence from weakest (anecdotes and testimonials) to strongest (meta-analyses of multiple randomized controlled trials). When you encounter a claim about any mental health treatment, look for red flags like reliance on testimonials, no control group, conflicts of interest, and absence of replication. Look for green flags like peer-reviewed research, randomized controlled trials, and endorsement by professional organizations. These critical thinking skills do not just apply in psychology class — they help you navigate health information throughout your life.

Varsity Tutors • Psychology • Evaluating Treatment Claims