KPEERI • FOUNDATIONAL CONCEPTS

Evaluating Assessment Validity — 4. Evaluate validity of assessment results

Learn how to systematically judge whether assessment outcomes truly measure what they claim to measure.

Historical Context & Motivation

The question of whether a test actually measures what it purports to measure has been central to the fields of education, psychology, and kinesiology for well over a century. Early assessments in physical education and exercise science were often accepted at face value — if a test looked reasonable, practitioners assumed it was adequate. However, as the consequences of assessment decisions grew — from determining student grades to guiding clinical rehabilitation protocols — the need for rigorous validity evaluation became impossible to ignore. The evolution of validity theory reflects a broader movement in the social and health sciences toward evidence-based practice, demanding that professionals justify every inference drawn from assessment data.

1920s
Early Trait-Based Testing
Physical fitness tests emerged in schools and the military, evaluated primarily on whether test items appeared to measure strength, endurance, or agility — an informal precursor to face validity.
1954
APA Technical Recommendations
The American Psychological Association published its first Technical Recommendations, formalizing four types of validity: content, predictive, concurrent, and construct. This taxonomy dominated validity discussions for decades.
1974
Messick's Unified Framework Begins
Samuel Messick began arguing that validity is not a property of a test itself but of the inferences and actions taken from test scores — a paradigm shift that moved the field toward a unified concept of construct validity.
1989
Messick's Seminal Chapter
Messick published his landmark chapter defining validity as 'an integrated evaluative judgment of the degree to which empirical evidence and theoretical rationales support the adequacy and appropriateness of inferences and actions based on test scores.' This unified view became the gold standard.
2014
Standards for Educational & Psychological Testing
The jointly authored Standards (AERA, APA, NCME) codified five sources of validity evidence — test content, response processes, internal structure, relations to other variables, and consequences — establishing the contemporary framework applied across kinesiology and exercise science.

The central question that this lesson addresses is both simple and profound: How do we determine whether the results produced by an assessment are valid for a specific purpose and population? Understanding the answer requires moving beyond the outdated notion that a test is simply 'valid' or 'invalid' and instead recognizing that validity is a matter of degree, context, and the accumulation of multiple lines of evidence.

Core Principles of Validity Evaluation

Evaluating the validity of assessment results rests on several foundational principles that guide professional practice in kinesiology, exercise science, and related fields. These principles reflect the modern understanding that validity is not an inherent quality of the test instrument alone; rather, it is an attribute of the score-based interpretations and the decisions that follow from those interpretations. A single assessment instrument may yield valid inferences for one population or purpose and invalid inferences for another.

1

Validity Is About Inferences, Not Tests

A test score is a number; validity concerns whether the conclusions drawn from that number are justified. The same VO₂max protocol may yield valid conclusions about aerobic capacity in adults but not in children with different physiological profiles.
2

Multiple Sources of Evidence

No single statistic or study establishes validity. Evaluators must gather and integrate evidence from content alignment, response processes, internal structure, relationships with external criteria, and the consequences of score use.
3

Validity Is a Matter of Degree

Validity exists on a continuum. Rather than labeling an assessment as 'valid' or 'invalid,' practitioners evaluate the strength and coherence of the evidence supporting specific interpretations and uses.
4

Context and Population Specificity

Validity evidence gathered in one context does not automatically generalize. An instrument validated for college athletes may require separate validation for older adults, clinical populations, or individuals from different cultural backgrounds.
5

Validity Requires Reliability as a Prerequisite

An unreliable assessment cannot produce valid inferences. If scores fluctuate randomly across repeated administrations, the consistency needed to support meaningful interpretation is absent. Reliability is necessary but not sufficient for validity.
KEY TAKEAWAY
Think of validity evaluation like a courtroom trial. The assessment result is the 'claim,' and each source of validity evidence is a 'witness.' No single witness proves the case — the judge (you, the evaluator) must weigh all the testimony together to decide whether the claim is well-supported. A strong case has multiple, independent witnesses whose stories converge. A weak case relies on one witness or reveals contradictions across testimony.

Visual Framework — Five Sources of Validity Evidence

The contemporary model of validity evaluation, as outlined in the Standards for Educational and Psychological Testing (2014), organizes validity evidence into five interconnected sources. These sources are not separate 'types' of validity but rather complementary categories of evidence that collectively support or undermine the validity argument. The following diagram illustrates how these five sources converge on the central claim — that a particular interpretation or use of assessment results is justified.

The diagram shows five sources of validity evidence — Test Content, Response Processes, Internal Structure, Relations to Other Variables, and Consequences of Testing — all converging on the central validity claim.

When evaluating the validity of assessment results, it is essential to consider which sources of evidence are most relevant for the intended use and to identify any gaps in the available evidence. For example, a new field test of muscular endurance may have strong content evidence (experts agree the exercises sample the endurance domain) and good criterion-related evidence (scores correlate with laboratory measures), but if internal structure has not been examined, one cannot be certain whether the test measures a single unitary construct or multiple distinct dimensions. Each missing or contradictory piece of evidence weakens the overall validity argument.

How Validity Evidence Is Gathered and Quantified

While validity is ultimately a qualitative judgment, the evidence that supports it is often quantitative. Evaluators rely on a range of statistical indices and analytic procedures to build — or challenge — a validity argument. Understanding these tools is critical for test prep because exam questions frequently require you to identify the appropriate type of evidence, interpret validity coefficients, or critique a study's validity claims.

Key Quantitative Indicators

VALIDITY COEFFICIENT (CRITERION-RELATED)
r_xy = Σ(Xᵢ − X̄)(Yᵢ − Ȳ) / √[Σ(Xᵢ − X̄)² × Σ(Yᵢ − Ȳ)²]
Where rxy is the Pearson correlation between test scores (X) and criterion scores (Y). Values closer to ±1.00 indicate stronger evidence that the test is related to the criterion. In kinesiology, a validity coefficient of r ≥ 0.80 is generally considered strong.
COEFFICIENT OF DETERMINATION
r² = (r_xy)²
The coefficient of determination (r²) represents the proportion of variance in the criterion that is explained by the test. For instance, if r = 0.90, then r² = 0.81, meaning 81% of the variance in the criterion is accounted for by test performance.
CONTENT VALIDITY INDEX (CVI)
CVI = (Number of experts rating item as 'relevant') / (Total number of experts)
The Content Validity Index quantifies the degree of expert agreement on whether test items adequately represent the construct domain. A CVI ≥ 0.80 is typically required, though stricter thresholds (≥ 0.78 per item, per Lynn, 1986) are common.

Beyond correlation coefficients, evaluators may use factor analysis to examine internal structure, known-groups comparisons to test whether the assessment discriminates between groups expected to differ (e.g., trained vs. untrained individuals), and convergent and discriminant evidence to determine whether scores correlate strongly with measures of similar constructs and weakly with measures of unrelated constructs. Each of these approaches provides a distinct thread in the larger tapestry of the validity argument.

Detailed Breakdown of the Five Sources of Evidence

Each of the five sources of validity evidence has its own methods, strengths, and typical applications. The table below provides a detailed comparison, and the subsequent diagram illustrates the decision-making process an evaluator follows when determining which sources of evidence to prioritize.

Summary of the Five Sources of Validity Evidence (AERA, APA, NCME, 2014)
Source of EvidenceKey QuestionTypical MethodsKinesiology Example
Test ContentDo the items/tasks adequately represent the construct domain?Expert panel review, content validity index (CVI), blueprint alignmentExperts confirm that a physical literacy assessment includes locomotor, stability, and manipulation skills
Response ProcessesAre examinees engaging the intended cognitive or motor processes?Think-aloud protocols, observation, eye-tracking, video analysisObserving that students perform a balance test using postural control strategies rather than compensatory trunk movements
Internal StructureDo the test components relate to each other consistent with the construct theory?Factor analysis (EFA/CFA), item-total correlations, Rasch modelingCFA confirms a fitness battery loads on two factors (cardiovascular endurance and muscular fitness) as theorized
Relations to Other VariablesDo scores relate to external criteria as predicted by theory?Convergent/discriminant correlations, criterion (concurrent/predictive) studies, known-groups methodA field-based VO₂max estimate correlates r = 0.85 with direct gas exchange measurement
Consequences of TestingDo score-based actions lead to intended outcomes without unintended negative effects?Impact studies, fairness analysis, examination of bias across subgroupsVerifying that a return-to-play protocol does not systematically disadvantage athletes of certain body types
This flowchart outlines the sequential process of gathering validity evidence. An evaluator begins by defining the assessment's purpose and target population, then systematically works through each evidence source. When evidence at any stage is weak, the evaluator may loop back to revise the instrument or narrow the scope of interpretation (see the sidebar note on the left).

Worked Example — Evaluating a Field Test of Cardiorespiratory Fitness

Suppose a university kinesiology department has developed a new 12-minute run test to estimate maximal oxygen uptake (VO₂max) among college-aged students. The department wants to use test results to classify students into fitness categories and make programming recommendations. Your task is to evaluate the validity of the assessment results for this stated purpose.

Evaluating the Validity of a 12-Minute Run Test
1
Step 1 — Identify the Intended Inference and UseThe test developers claim that distance covered in 12 minutes predicts VO₂max and can be used to place students into fitness categories (e.g., below average, average, above average). The inference is that a higher distance reflects a higher level of cardiorespiratory fitness.
Claim: 12-min run distance → VO₂max estimate → fitness classification
2
Step 2 — Evaluate Content EvidenceA panel of five exercise physiologists reviewed the test protocol and agreed that a sustained 12-minute run primarily taxes the aerobic energy system, which is the physiological basis of VO₂max. The item-level CVI was calculated: all five experts rated the test as 'relevant' or 'very relevant' to cardiorespiratory fitness.
CVI = 5/5 = 1.00 (strong content evidence)
3
Step 3 — Evaluate Criterion-Related EvidenceA concurrent validity study was conducted: 80 college students performed both the 12-minute run and a laboratory-based VO₂max test (the gold standard) within the same week. The Pearson correlation between 12-minute run distance and measured VO₂max was computed.
r = 0.87, r² = 0.76 → 76% of variance in VO₂max is accounted for by the run distance (strong criterion-related evidence)
4
Step 4 — Evaluate Known-Groups EvidenceThe researchers compared 12-minute run scores between varsity endurance athletes (n = 25) and sedentary students (n = 25). An independent-samples t-test showed a statistically significant difference (p < 0.001, d = 2.3), confirming that the test discriminates between groups known to differ in cardiorespiratory fitness.
Large effect size (d = 2.3) supports construct validity via known-groups comparison
5
Step 5 — Identify Gaps and Form an Integrated JudgmentContent evidence is strong (CVI = 1.00). Criterion-related evidence is strong (r = 0.87). Known-groups evidence supports the test's ability to discriminate. However, no response-process evidence has been gathered (e.g., Were students pacing appropriately or sprinting and walking?), and consequence evidence is absent (e.g., Do the fitness classifications lead to fair and effective programming for all subgroups, including students with disabilities or different BMI profiles?).
Overall judgment: Moderate-to-strong validity for estimating VO₂max in healthy college students, but additional evidence (response processes, consequences) is needed before using results for high-stakes placement decisions.

Strengths, Limitations, and Common Pitfalls

Evaluating validity of assessment results is a powerful professional competency, but it comes with inherent challenges. Understanding both the strengths of the modern validity framework and its limitations will help you approach exam questions — and real-world practice — with appropriate nuance.

Strengths and Limitations of the Modern Validity Framework
StrengthsLimitations
The unified framework prevents over-reliance on a single statistic (e.g., a high correlation) as 'proof' of validity.Gathering all five sources of evidence is time-consuming and expensive; many published assessments lack complete validity evidence.
Emphasizing inferences rather than instruments focuses attention on the real-world consequences of assessment decisions.The concept of 'consequential validity' remains controversial; some scholars argue that social consequences should not be considered part of validity proper.
Context-specificity forces practitioners to validate assessments for each new population, preventing inappropriate generalizations.Context-specificity also means validity evidence is never 'settled' — new contexts always require new investigation.
Quantitative indices (r, r², CVI) provide objective benchmarks for evaluating evidence quality.Over-reliance on correlations can be misleading; a high r does not guarantee the test measures the construct correctly (e.g., confounded variables).
The framework is applicable across disciplines — education, clinical practice, sport science, public health.Qualitative evidence (e.g., expert panels, think-alouds) introduces subjectivity that can be difficult to standardize.
COMMON EXAM PITFALL
A frequently tested misconception is treating validity as a binary property — 'the test is valid' or 'the test is not valid.' On exams, choose answers that describe validity as a degree of support for a specific interpretation or use. Also beware of answer choices that equate high reliability with high validity; reliability is necessary but not sufficient. An assessment can produce highly consistent (reliable) scores that consistently measure the wrong construct.

Connection to Advanced Validity Theory

The foundational concepts covered in this lesson connect directly to more advanced frameworks that you may encounter in graduate-level coursework or advanced test prep. Understanding these connections will strengthen your ability to answer higher-order questions and to appreciate the broader intellectual landscape of assessment theory.

From Foundational to Advanced Validity Concepts
Foundational ConceptAdvanced ExtensionKey Idea
Five sources of validity evidenceKane's Argument-Based ApproachValidity is structured as an explicit interpretive argument (IUA) with assumptions that must be tested. Each source of evidence corresponds to a link in the argumentative chain.
Content validity index (CVI)Generalizability Theory (G Theory)Extends classical reliability into multiple facets (raters, items, occasions) and provides variance components that inform how well a test domain is sampled.
Factor analysis for internal structureItem Response Theory (IRT)IRT models provide item-level validity information (discrimination, difficulty) and allow for adaptive testing and differential item functioning (DIF) analysis.
Consequences of testingFairness & Equity FrameworksAdvanced work examines measurement invariance across groups, DIF analysis, and systemic bias to ensure that score-based decisions do not disproportionately harm underrepresented populations.

Kane's argument-based approach to validation is particularly influential in contemporary assessment theory. Rather than simply listing evidence, Kane asks the evaluator to articulate an explicit interpretive/use argument (IUA) that maps out every inference from observed performance to the final decision. Each inference — from scoring, to generalization, to extrapolation, to the decision itself — requires its own supporting evidence. If any link in this chain is unsupported or falsified, the entire validity argument is weakened. This structured approach is increasingly reflected in certification and licensure exams in kinesiology and related health professions.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher states, 'Our balance assessment is valid.' According to modern validity theory, what is fundamentally wrong with this statement, and how should the claim be rephrased?
PROBLEM 2BASIC CALCULATION
A panel of 8 experts reviewed a new flexibility assessment. For a particular test item, 6 experts rated it as 'relevant' or 'very relevant' to the flexibility construct, and 2 rated it as 'not relevant.' Calculate the item-level Content Validity Index (CVI) and determine whether the item meets the commonly accepted threshold.
PROBLEM 3INTERMEDIATE
A new agility test yields a validity coefficient of r = 0.72 when correlated with an established criterion measure. The test also has a test-retest reliability coefficient of r = 0.78. (a) Calculate the coefficient of determination for the validity coefficient. (b) What does this value tell you about the test's predictive accuracy? (c) How does the relatively low reliability limit the potential validity?
PROBLEM 4APPLIED
You are a kinesiology professional tasked with selecting a pre-participation screening tool for a youth sports league. You are choosing between two assessments: Assessment A has a validity coefficient of r = 0.90 with a gold standard, but it was validated only on adult athletes. Assessment B has a validity coefficient of r = 0.75, but it was validated on youth aged 10–14. Both have adequate reliability. Which assessment would you recommend, and what sources of validity evidence inform your decision?
PROBLEM 5CRITICAL THINKING
A university publishes a study showing that their new functional movement screen (FMS variant) has strong content evidence (CVI = 0.95), strong criterion-related evidence (r = 0.85 with injury incidence over one season), and strong internal structure (CFA supports a single-factor model). Based on these three sources alone, a colleague argues that the assessment is 'fully validated.' Construct a counterargument identifying at least two sources of validity evidence that remain unaddressed and explain why these gaps matter for high-stakes decisions like return-to-play clearance.

Lesson Summary

Evaluating the validity of assessment results requires understanding that validity is a property of score-based inferences, not of the test itself. The modern framework identifies five sources of evidencetest content, response processes, internal structure, relations to other variables, and consequences of testing — that must be integrated to form a coherent validity argument. No single source is sufficient on its own.

Key quantitative tools include the validity coefficient (r), the coefficient of determination (r²), and the content validity index (CVI). Remember that reliability is necessary but not sufficient for validity, that validity is always a matter of degree, and that evidence gathered in one context does not automatically generalize to another population or purpose. For advanced applications, Kane's argument-based approach extends these foundations by structuring validity as an explicit interpretive argument with testable assumptions at each inferential step.

Varsity Tutors • KPEERI • Evaluating Assessment Validity — 4. Evaluate validity of assessment results