PSYCHOLOGY • FOUNDATIONS & RESEARCH METHODS

Reliability & Validity — I can explain reliability and validity in psychological measurement and identify threats to each.

Understanding how psychologists ensure their measurements are consistent and actually measure what they claim to measure.

Historical Context & Motivation

Psychology has always faced a unique challenge compared to sciences like physics or chemistry: how do you accurately measure something you cannot directly see or touch, like intelligence, anxiety, or personality? Early psychologists in the late 1800s recognized that if their field was going to be taken seriously as a science, they needed tools that produced trustworthy results. This drive led to the development of two foundational concepts — reliability and validity — that serve as the quality-control standards for every psychological test, survey, and measurement ever created.

1890
Cattell's Mental Tests
James McKeen Cattell coined the term "mental test" and began measuring reaction times and sensory abilities, raising immediate questions about whether these tests were consistent and meaningful.
1904
Spearman's Correlation Work
Charles Spearman developed statistical methods — including the correlation coefficient — that allowed researchers to quantify how consistent a test's scores were, laying the mathematical groundwork for reliability analysis.
1954
APA Technical Standards
The American Psychological Association published its first official technical standards for psychological tests, formally defining types of validity (content, criterion, and construct) that test developers must demonstrate.
1955
Cronbach & Meehl on Construct Validity
Lee Cronbach and Paul Meehl published a landmark paper arguing that construct validity — whether a test truly measures the abstract concept it claims to — is the most important type of validity in psychology.
2014
Modern Standards Update
The latest edition of the Standards for Educational and Psychological Testing was released, reflecting decades of refinement in how psychologists think about measurement quality in an era of computerized and adaptive testing.

The central question that drove all of this progress is deceptively simple: How can we be confident that a psychological measurement is both consistent and accurate? Without reliability, a test gives random, unpredictable results. Without validity, a test might be perfectly consistent but measure the wrong thing entirely. Understanding these two concepts is essential for evaluating any claim in psychology — from an IQ score on a report card to a diagnosis of a mental health condition.

Core Principles & Definitions

At its core, reliability refers to the consistency of a measurement. A reliable test produces similar results under similar conditions. If you step on a bathroom scale three times in one minute, you expect to see roughly the same weight each time. If the scale shows 130 lbs, then 145 lbs, then 118 lbs, that scale is unreliable — you can't trust any single reading. Validity refers to accuracy — whether a measurement tool actually measures what it claims to measure. A clock may be perfectly reliable, ticking at the exact same rate every day, but using it to measure your intelligence would make it completely invalid for that purpose.

1

Reliability = Consistency

A measurement is reliable when it produces stable, repeatable results across time, items, or raters. Think of it as the precision of a dart throw — are the darts landing in the same cluster?
2

Validity = Accuracy

A measurement is valid when it actually measures the concept it is supposed to measure. This is like asking whether the darts are hitting the bullseye — are they landing in the right place?
3

Reliability Is Necessary but Not Sufficient

A test must be reliable before it can be valid, but reliability alone does not guarantee validity. A broken thermometer could consistently read 5° too high — reliable but not valid.
4

Multiple Types of Each

There are several subtypes of reliability (test-retest, internal consistency, inter-rater) and validity (content, criterion, construct). Each assesses a different dimension of measurement quality.
5

Threats Reduce Quality

Many factors can threaten reliability (fatigue, ambiguous items, environmental changes) and validity (biased items, confounding variables, poor sampling). Identifying threats is a critical research skill.
KEY TAKEAWAY
Think of reliability and validity like a GPS device. Reliability is like the GPS giving you the same coordinates every time you stand in the same spot. Validity is whether those coordinates actually match your real location on the map. A GPS that always says you're two miles north of where you really are is reliable (consistent) but not valid (accurate). You need both for the device to be truly useful.

Visual Explanation — The Target Analogy

The most famous way to visualize the relationship between reliability and validity is the target (dartboard) analogy. Imagine four different dartboards, each representing a different combination of high or low reliability and validity. The bullseye represents the true value of whatever you are trying to measure, and each dart represents one measurement attempt.

The three targets show key combinations. The cyan target (left) shows ideal measurement — darts are tightly clustered on the bullseye. The violet target (center) shows a test that is reliable but not valid — consistent results that miss the mark. The pink target (right) shows neither reliability nor validity — random scatter everywhere.

Notice an important logical relationship in the diagram: you cannot have high validity without high reliability. If your darts are scattered all over the board (unreliable), they can't all be on the bullseye (valid). This is why psychologists say reliability is a necessary but not sufficient condition for validity. You must achieve consistency first, then confirm that the consistent results are actually measuring the right thing.

How Reliability & Validity Are Assessed

Assessing Reliability

Psychologists use several methods to evaluate reliability, each targeting a different source of potential inconsistency. The strength of reliability is typically expressed as a correlation coefficient — a number between 0 and 1, where values closer to 1 indicate stronger consistency. A reliability coefficient of 0.80 or higher is generally considered acceptable for psychological tests.

RELIABILITY COEFFICIENT
r = correlation between two sets of scores (range: 0.00 to 1.00)
Where r represents the degree of consistency. An r of 0.00 means no consistency (pure randomness), while an r of 1.00 means perfect consistency. Most well-designed psychological tests achieve reliability coefficients between 0.70 and 0.95.
Four major types of reliability in psychological testing
Type of ReliabilityWhat It MeasuresHow It's Done
Test-RetestConsistency over time — does the test give the same results when administered again later?Give the same test to the same group on two different occasions and correlate the scores.
Internal ConsistencyConsistency across items — do all the questions on the test measure the same underlying concept?Use split-half method (correlate first half with second half) or calculate Cronbach's alpha across all items.
Inter-RaterConsistency across scorers — do different people grading or observing the same behavior assign similar scores?Have two or more raters independently score the same responses and correlate their judgments.
Parallel FormsConsistency across equivalent versions — do two different versions of the test produce similar scores?Create two equivalent test forms, administer both to the same group, and correlate the results.

Assessing Validity

While reliability asks "Is this test consistent?", validity asks the deeper question: "Is this test actually measuring what we think it's measuring?" There are three major categories of validity evidence that psychologists look for when evaluating a measurement tool.

Three major types of validity evidence
Type of ValidityCore QuestionExample
Content ValidityDoes the test cover a representative sample of the topic it's supposed to measure?A history final exam that only covers Chapter 1 lacks content validity — it doesn't represent the whole course.
Criterion ValidityDoes the test predict or correlate with an external criterion (real-world outcome)?SAT scores that correlate with college GPA show criterion validity. Can be concurrent (measured now) or predictive (measured later).
Construct ValidityDoes the test truly capture the abstract psychological concept (construct) it claims to?An anxiety questionnaire should correlate with other anxiety measures (convergent) and NOT correlate with unrelated traits like athletic ability (discriminant).
💡 Face Validity — A Common Misconception
Face validity refers to whether a test looks like it measures what it claims to. While it can help with test-taker cooperation, face validity is not considered true scientific validity because appearances can be deceiving. A test could look relevant but still measure the wrong thing.

Threats to Reliability & Validity

Even well-designed tests can have their reliability or validity undermined by various threats — factors that introduce error or bias into the measurement process. Recognizing these threats is one of the most important skills in evaluating psychological research. The diagram below organizes the major threats into two categories.

The left column (blue) lists five major threats to reliability, all of which introduce random or inconsistent error. The right column (violet) lists five major threats to validity, all of which cause systematic bias or inaccuracy in what the test measures.

A key distinction to remember is that threats to reliability tend to involve random error — unpredictable fluctuations that make scores bounce around. Threats to validity, on the other hand, involve systematic error — consistent biases that push scores in one direction. For example, if a math test includes very complex English vocabulary, non-native English speakers might consistently score lower not because they lack math ability, but because the test is accidentally measuring English proficiency. That's a validity threat — the test is systematically measuring the wrong thing.

Worked Example — Evaluating a New Anxiety Questionnaire

Let's walk through a realistic scenario. Imagine a school psychologist develops a 20-item questionnaire called the "Student Anxiety Scale" (SAS) to measure test anxiety in high school students. She wants to determine if the SAS is both reliable and valid before using it to identify students who might need support.

Evaluating the Student Anxiety Scale (SAS)
1
Step 1 — Check Test-Retest ReliabilityThe psychologist administers the SAS to 100 students, waits two weeks, and then gives the same test to the same students. She calculates the correlation between the two sets of scores.
Correlation r = 0.85. This is above the 0.80 threshold, so the SAS demonstrates good test-retest reliability. Students' scores are consistent over time.
2
Step 2 — Check Internal ConsistencyShe calculates Cronbach's alpha for the 20 items to see if all the questions are measuring the same underlying construct (test anxiety). Items that don't correlate well with the others may need to be revised or removed.
Cronbach's α = 0.91. This is excellent internal consistency, meaning the 20 items are consistently measuring the same concept.
3
Step 3 — Check Content ValidityShe asks a panel of five psychology teachers and counselors to review the 20 items and judge whether they represent the full range of test anxiety symptoms (physical symptoms, worried thoughts, avoidance behaviors, etc.).
The panel notes that the SAS includes items about worry and physical symptoms but has no items about avoidance behaviors (like skipping class before a test). This is a content validity concern — the test underrepresents the construct.
4
Step 4 — Check Criterion ValidityShe compares SAS scores with students' actual exam performance and with observations of anxiety-related behaviors during test-taking (fidgeting, leaving early, asking to use the restroom).
SAS scores correlate r = −0.42 with exam grades (higher anxiety, lower grades) and r = 0.55 with observed anxious behaviors. Both correlations are in the expected direction, supporting moderate criterion validity.
5
Step 5 — Identify Remaining Threats & Recommend ImprovementsThe psychologist considers potential threats. Could social desirability bias (students underreporting anxiety to look "tough") hurt validity? Could the fact that the questionnaire is given right before an exam inflate anxiety scores for the retest, hurting generalizability?
Recommendation: Add items about avoidance behaviors to improve content validity. Include reverse-scored items to reduce response bias. Consider administering the test at a neutral time, not right before exams, to control for situational confounds.

Reliability vs. Validity — Side-by-Side Comparison

Students sometimes confuse reliability and validity because both involve evaluating the quality of a test. The table below provides a direct comparison to help you keep them straight.

Key differences between reliability and validity
FeatureReliabilityValidity
Central question"Does this test give consistent results?""Does this test measure what it claims to?"
AnalogyA scale that gives the same weight each timeA scale that shows your actual weight
Type of errorRandom error (inconsistency)Systematic error (bias)
Can exist without the other?Yes — a test can be reliable but not validNo — a test cannot be valid without being reliable
How it's measuredCorrelation coefficients (test-retest, split-half, inter-rater, Cronbach's alpha)Expert review, correlations with criteria, theoretical analysis
Easier to establish?Yes — mostly statistical, can be calculated directlyNo — requires judgment, multiple sources of evidence over time
KEY TAKEAWAY
Here's a memory trick: Reliability is about repeatability (both start with "re"). Validity is about value — is the test actually valuable because it measures the right thing? A clock that's always five minutes fast is repeatable (reliable) but doesn't give you the true time (not valid).

Connections to Advanced Measurement Theory

The concepts of reliability and validity you've learned here form the foundation of a larger field called psychometrics — the science of psychological measurement. In college-level psychology and statistics courses, these ideas are explored with much greater mathematical depth. The table below previews how these introductory concepts connect to more advanced topics.

From introductory concepts to advanced psychometrics
What You Know NowAdvanced Extension
Reliability as a correlation between two sets of scores (r = 0.85)Classical Test Theory (CTT) formally defines an observed score as the sum of a true score plus error: X = T + E. Reliability equals the ratio of true score variance to total variance.
Internal consistency measured by Cronbach's alphaItem Response Theory (IRT) models each item individually, estimating item difficulty and discrimination parameters. This allows adaptive testing where different students get different questions.
Content, criterion, and construct validity as separate categoriesModern validity theory (Messick, 1989) treats validity as a unified concept — all evidence contributes to one overarching judgment about the appropriateness of test-score interpretations.
Threats like cultural bias and construct-irrelevant varianceDifferential Item Functioning (DIF) analysis statistically detects items that perform differently across demographic groups, even when those groups have equal ability levels.

You don't need to master these advanced topics now, but it's worth knowing that the reliability and validity framework you're building is the gateway to a sophisticated and mathematically rich area of psychology. These tools are used in high-stakes settings every day — from college admissions testing (SAT, ACT) to clinical diagnosis (depression screeners, ADHD assessments) to personnel selection in organizations. The principles remain the same regardless of the complexity of the math: good measurement requires both consistency and accuracy.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain, in your own words, the difference between reliability and validity. Then explain why a test can be reliable without being valid, but cannot be valid without being reliable.
PROBLEM 2BASIC CALCULATION
A researcher gives the same depression questionnaire to 50 participants on two occasions, two weeks apart. The test-retest correlation is r = 0.62. Is this considered acceptable reliability? Justify your answer using the standard threshold discussed in the lesson.
PROBLEM 3INTERMEDIATE
A school develops a new "Critical Thinking Assessment" for 10th graders. The test has high internal consistency (Cronbach's α = 0.89) and good test-retest reliability (r = 0.83). However, students' scores on this assessment correlate r = 0.78 with their reading comprehension scores and only r = 0.25 with their performance on logic puzzles. What type of validity concern does this pattern suggest? Explain your reasoning.
PROBLEM 4APPLIED
A company uses a personality test to screen job applicants for a customer service position. The test was originally developed and validated using a sample of college students in the United States. The company now wants to use it to hire employees in Japan. Identify at least two specific threats to reliability or validity that might arise in this situation, and suggest one way to address each threat.
PROBLEM 5CRITICAL THINKING
Consider this argument: 'Standardized intelligence tests like the IQ test have very high reliability coefficients (often above 0.90), which proves they are excellent measures of intelligence.' Do you agree or disagree with this argument? Use the concepts from this lesson to construct a thorough response that addresses both what this evidence does and does not tell us.

Lesson Summary

Reliability and validity are the two pillars of quality in psychological measurement. Reliability refers to the consistency of a measurement — assessed through test-retest, internal consistency, inter-rater, and parallel forms methods. Validity refers to accuracy — whether the test measures its intended construct — and includes content validity, criterion validity, and construct validity. A test must be reliable before it can be valid, but reliability alone does not guarantee validity.

Common threats to reliability include test-taker variability, ambiguous items, environmental changes, and scorer subjectivity. Threats to validity include construct underrepresentation, construct-irrelevant variance, cultural bias, and confounding variables. Understanding these concepts equips you to critically evaluate any psychological test, research study, or measurement claim you encounter — a skill that extends far beyond the psychology classroom into everyday life.

Varsity Tutors • Psychology • Reliability & Validity