KPEERI • FOUNDATIONAL CONCEPTS

Relating Assessment Results to Difficulties — 5. Relate norm-referenced and informal assessment results to literacy difficulties

Synthesizing standardized and classroom-based data to identify and address specific literacy difficulties in learners.

Historical Context & Motivation

The practice of connecting assessment data to specific literacy difficulties has evolved significantly over the past century. Early reading instruction relied almost exclusively on teacher observation and subjective judgment, offering little in the way of systematic diagnosis. As norm-referenced tests emerged in the mid-twentieth century, educators gained a standardized lens through which to compare individual student performance to age- or grade-level peers. However, these tests alone could not reveal why a student struggled—only that a gap existed. The subsequent rise of informal assessment methods—running records, miscue analyses, writing samples, and curriculum-based measures—provided the granular, qualitative detail that standardized instruments lacked. Together, these two streams of data now form the diagnostic backbone of modern literacy intervention frameworks, including Response to Intervention (RTI) and Multi-Tiered Systems of Support (MTSS).

1916
Stanford-Binet Intelligence Scale Published
Lewis Terman's adaptation of the Binet-Simon test established the norm-referenced paradigm in American education, comparing individuals against statistically derived population norms. This framework would later extend to reading and literacy assessment.
1960s
Informal Reading Inventories Emerge
Educators such as Emmett Betts formalized the concept of instructional, independent, and frustration reading levels determined through informal reading inventories (IRIs), providing classroom-based diagnostic data that complemented standardized scores.
1969
Goodman's Miscue Analysis
Kenneth Goodman introduced miscue analysis, a method for examining the types of reading errors students make. This qualitative approach illuminated the cognitive and linguistic processes underlying reading difficulties rather than just quantifying deficits.
2000s
RTI / MTSS Frameworks Adopted
Federal legislation and educational policy promoted Response to Intervention and Multi-Tiered Systems of Support, which require educators to integrate norm-referenced screening data with ongoing informal progress-monitoring to make diagnostic and instructional decisions.
2015–Present
Science of Reading Integration
The Science of Reading movement has reinforced the need for structured diagnostic assessment across the five pillars of literacy—phonemic awareness, phonics, fluency, vocabulary, and comprehension—using both formal and informal tools in concert.

The central question this lesson addresses is both practical and conceptual: How do educators synthesize quantitative norm-referenced scores with qualitative informal assessment data to pinpoint, classify, and address specific literacy difficulties? Answering this question is essential not only for test preparation but for any professional role that involves diagnosing or remediating reading and writing problems in learners.

Core Principles & Definitions

Before one can relate assessment results to literacy difficulties, it is necessary to understand the fundamental nature of each assessment type and the kinds of evidence they generate. Norm-referenced assessments compare a student's performance to a norming sample—a large, representative group whose scores establish percentile ranks, stanines, grade equivalents, and standard scores. In contrast, informal assessments are criterion-referenced or descriptive in nature, measuring what a student can or cannot do against a defined set of skills or benchmarks rather than against other students. Both data types are indispensable: norm-referenced data tells you where a student stands relative to peers, while informal data tells you what specific skills need attention.

1

Norm-Referenced Scores Indicate Relative Standing

Percentile ranks, standard scores (mean = 100, SD = 15 on many tests), and stanines tell the diagnostician how far a student's performance deviates from the normative average. A percentile rank below the 25th percentile generally signals a need for further diagnostic investigation.
2

Informal Assessments Reveal Skill-Level Detail

Tools such as running records, phonics inventories, oral reading fluency probes, cloze procedures, and writing rubrics expose the precise areas of breakdown—whether in decoding, fluency, vocabulary, or comprehension strategies—that norm-referenced scores only hint at.
3

Triangulation Strengthens Diagnostic Accuracy

No single data source is sufficient. By triangulating norm-referenced data, informal assessment results, and classroom observations, educators reduce the risk of misidentification and construct a more complete profile of a student's literacy strengths and weaknesses.
4

Pattern Analysis Links Data to Difficulties

Relating results to difficulties requires pattern recognition: consistent weaknesses across multiple measures in a particular domain (e.g., phonological awareness) constitute convergent evidence for a specific literacy difficulty, warranting targeted intervention.
5

Assessment Drives Instruction, Not Just Classification

The ultimate purpose of relating assessment results to difficulties is instructional decision-making—selecting evidence-based interventions, setting measurable goals, and monitoring progress. Classification alone is insufficient without actionable next steps.
KEY TAKEAWAY
Think of norm-referenced assessments as a satellite view and informal assessments as a ground-level survey. The satellite image (norm-referenced data) reveals that a region is experiencing drought—analogous to a low reading score. But you need the ground-level survey (informal data) to determine whether the cause is depleted aquifers, damaged irrigation channels, or soil erosion. Only by combining both perspectives can you prescribe the right remedy.

Visual Explanation — The Diagnostic Synthesis Model

The following diagram illustrates how norm-referenced and informal assessment data converge in a diagnostic synthesis model. Notice that the process begins with universal screening (a norm-referenced step), moves into targeted informal assessment for students who fall below benchmarks, and culminates in an integrated profile that maps directly onto specific literacy difficulties and corresponding interventions.

The model begins with norm-referenced universal screening to flag at-risk students, then branches into norm-referenced diagnostic tests and informal assessments that converge on an integrated literacy profile, ultimately driving targeted intervention.

The flowchart emphasizes a critical principle: data from both formal and informal sources must be integrated rather than used in isolation. A norm-referenced screening score of the 15th percentile in reading comprehension signals concern, but it does not specify whether the difficulty lies in decoding, vocabulary knowledge, background knowledge, inferencing, or text structure awareness. Informal assessments then provide the diagnostic specificity needed to create an actionable intervention plan. The convergence of both data streams into an integrated profile is what distinguishes competent diagnostic practice from surface-level score reporting.

How It Works — Linking Scores to Literacy Domains

To relate assessment results to literacy difficulties effectively, the diagnostician must understand the statistical framework underlying norm-referenced scores and how those scores map onto qualitative performance levels. Most norm-referenced literacy assessments report standard scores with a mean of 100 and a standard deviation (SD) of 15. Performance is then classified into descriptive bands that correspond to varying degrees of concern.

STANDARD SCORE FORMULA
Standard Score = ((X − M) / SD) × 15 + 100
Where X = student's raw score, M = mean raw score of the norming sample, and SD = standard deviation of the norming sample. Scores below 85 (−1 SD) typically warrant diagnostic follow-up; scores below 70 (−2 SD) indicate significant impairment.
PERCENTILE RANK INTERPRETATION
Percentile Rank = (Number of scores below X / Total N) × 100
A percentile rank of 10 means the student scored higher than only 10% of the norming sample. In RTI frameworks, students at or below the 25th percentile on universal screeners are typically flagged for Tier 2 intervention, while those at or below the 10th percentile may require Tier 3 intensive support.

While these quantitative indices place students along a continuum, they do not by themselves reveal the component literacy skills driving the deficit. That is where informal assessment data enters the diagnostic equation. Consider a student whose Woodcock-Johnson IV Passage Comprehension subtest yields a standard score of 78 (8th percentile). This norm-referenced result signals a comprehension difficulty, but it does not explain the mechanism. An informal reading inventory might then reveal that the student reads text at an independent level with high accuracy but cannot retell key ideas—pointing toward a comprehension-monitoring or inference-making weakness rather than a decoding issue. Alternatively, a running record might show frequent substitutions of visually similar words, suggesting that poor decoding accuracy is undermining comprehension from the bottom up.

🔍 Key Diagnostic Principle
Always ask: Does the norm-referenced score confirm or contradict the informal data? Concordant data (both sources pointing to the same area of weakness) strengthens diagnostic confidence. Discordant data (contradictory findings) signals the need for additional assessment to resolve ambiguity.

Mapping Assessments to the Five Literacy Domains

Effective diagnosis requires mapping both norm-referenced and informal assessment data onto the five pillars of literacy identified by the National Reading Panel: phonemic awareness, phonics (decoding), fluency, vocabulary, and comprehension. The following diagram and table detail which assessment tools typically address each domain and how their results relate to specific difficulties within that domain.

The top row maps each literacy pillar to its norm-referenced (NR) and informal (Inf) tools. The bottom section demonstrates how the same low norm-referenced comprehension score can reflect entirely different root causes—one a true comprehension difficulty (left), the other a decoding/fluency difficulty masquerading as a comprehension problem (right). Only informal assessment resolves this ambiguity.
Assessment tools and difficulty indicators mapped to the five literacy domains
Literacy DomainNorm-Referenced ToolsInformal ToolsTypical Difficulty Indicator
Phonemic AwarenessCTOPP-2 (Elision, Blending Words)Yopp-Singer, PAST, phoneme segmentation tasksSS < 85; cannot isolate, segment, or blend individual phonemes
Phonics / DecodingWJ-IV Word Attack, WIAT-4 Pseudoword DecodingPhonics survey, nonsense word fluency, spelling inventoriesSS < 85; high error rate on CVC/CVCe patterns; reliance on guessing
FluencyGORT-5, TOWRE-2CBM-ORF probes, timed reading, NAEP fluency scaleORF below 50th percentile; word-by-word reading; flat prosody
VocabularyPPVT-5, EVT-3, WJ-IV Oral VocabularyCloze tasks, vocabulary matching, word sorts, context-use checksSS < 85; limited expressive/receptive vocabulary; difficulty with Tier 2 words
ComprehensionWJ-IV Passage Comprehension, WIAT-4 Reading ComprehensionQRI-7, retelling rubrics, think-alouds, written responsesSS < 85; poor retelling, inability to infer, weak main-idea identification

Worked Example — Diagnosing a Third-Grader's Literacy Difficulty

The following scenario demonstrates how a reading specialist integrates norm-referenced and informal data to identify a specific literacy difficulty and plan intervention.

Case Study: Mia, Age 8, Third Grade
1
Step 1 — Review Norm-Referenced Screening DataMia's fall universal screening on the aimswebPlus yielded the following results: Oral Reading Fluency (ORF) = 52 words correct per minute (WCPM), placing her at the 12th percentile for third graders nationally. Her comprehension screener placed her at the 18th percentile. Both scores fall below the 25th percentile benchmark, flagging Mia for Tier 2 diagnostic follow-up.
ORF: 12th percentile; Comprehension: 18th percentile → Tier 2 flag
2
Step 2 — Administer Norm-Referenced Diagnostic BatteryThe reading specialist administered selected subtests of the Woodcock-Johnson IV Tests of Achievement. Mia's scores were: Letter-Word Identification SS = 92 (30th percentile), Word Attack SS = 79 (8th percentile), Passage Comprehension SS = 84 (14th percentile), and Oral Vocabulary SS = 105 (63rd percentile). The pattern reveals a notable discrepancy: her sight-word reading is low-average, but her phonics-based decoding of novel words is significantly below average, while her oral vocabulary is solidly average.
Word Attack SS = 79 (significantly low); Oral Vocabulary SS = 105 (average) → Decoding deficit suspected
3
Step 3 — Conduct Informal AssessmentsThe specialist administered a running record on a grade-level narrative passage. Mia read at 91% accuracy (instructional level), with errors concentrated on multisyllabic words and words with vowel teams (e.g., reading 'bead' as 'bed,' 'trainer' as 'tran-er'). A phonics skills inventory confirmed mastery of initial consonants, short vowels, and CVC patterns but revealed significant weaknesses in vowel teams, r-controlled vowels, and syllable division rules. A retelling rubric showed that when Mia was read to, she could summarize main ideas and identify character motivations effectively, confirming that her comprehension skills are intact when decoding demands are removed.
Error pattern: vowel teams, r-controlled vowels, multisyllabic words; comprehension intact when listening
4
Step 4 — Triangulate and Identify the DifficultyThe norm-referenced data (low Word Attack, low ORF, low-average comprehension, average oral vocabulary) converges with the informal data (errors on advanced phonics patterns, intact listening comprehension) to produce a clear diagnostic picture. Mia's primary difficulty is in phonics/decoding, specifically in applying advanced vowel patterns and decoding multisyllabic words. Her comprehension deficit is secondary—a downstream consequence of effortful, inaccurate decoding rather than a primary comprehension impairment.
Primary difficulty: Advanced phonics/decoding deficit; Secondary: comprehension impacted by poor decoding
5
Step 5 — Plan Intervention and Set Monitoring GoalsBased on the integrated profile, the specialist prescribes a structured phonics intervention targeting vowel teams and syllable division (e.g., Wilson Reading System or a comparable systematic program) at Tier 2 intensity (3–4 sessions per week, 30 minutes each). The measurable goal is: Mia will increase her ORF from 52 WCPM to 75 WCPM (≥ 25th percentile) by the spring benchmark, and her Word Attack standard score will improve to ≥ 85. Progress monitoring via weekly CBM-ORF probes and biweekly nonsense word fluency checks will track her response to intervention.
Intervention: Structured phonics (vowel teams, syllable division); Goal: ORF ≥ 75 WCPM by spring; Monitoring: Weekly CBM-ORF

Strengths & Limitations of Each Assessment Type

Neither norm-referenced nor informal assessments are sufficient in isolation. Understanding the strengths and limitations of each type is essential for determining when and how to use them in diagnostic decision-making. The following table organizes these considerations side by side.

Comparison of norm-referenced and informal assessment types across key dimensions
DimensionNorm-Referenced AssessmentsInformal Assessments
StrengthsStandardized administration ensures consistency; percentile ranks enable comparison to national norms; psychometric properties (reliability, validity) are well-documented; results are recognized for eligibility decisions (e.g., IEPs, 504 plans)Flexible and responsive to individual students; reveal specific skill-level strengths/weaknesses; can be embedded in daily instruction; provide immediate, actionable data for intervention planning
LimitationsDo not specify root cause of difficulty; may lack cultural or linguistic sensitivity; snapshot in time—do not capture day-to-day variability; expensive and time-consuming to administer fullyLack standardized norms; inter-rater reliability may vary; results may not be accepted for formal eligibility decisions; quality depends heavily on the examiner's expertise
Best Used ForScreening, quantifying severity, benchmarking against peers, eligibility determination, pre/post outcome measurementDiagnosing specific skill deficits, planning instructional interventions, ongoing progress monitoring, confirming or clarifying norm-referenced findings
Example ScenarioA GORT-5 Fluency Score of 4 (SS = 70) tells you the student is at the 2nd percentile in oral reading fluency—far below peers—but does not explain whether the cause is poor decoding, limited sight-word automaticity, or an underlying language disorder.A running record on the same student reveals 85% accuracy with most errors on vowel-team words, no self-corrections, and word-by-word phrasing—pointing to a specific phonics deficit rather than a broader language issue.
KEY TAKEAWAY
Think of norm-referenced assessments as the lab blood work in a medical examination and informal assessments as the physician's physical exam and patient interview. The blood work yields precise numerical values that flag abnormalities against population baselines, but the doctor must palpate, observe, and ask questions to interpret those numbers in context and arrive at an accurate diagnosis. Neither the lab results nor the physical exam alone produces a reliable diagnosis—both are necessary for competent clinical judgment, and the same principle applies to literacy assessment.

Connection to Advanced Diagnostic Frameworks

The process of relating assessment results to literacy difficulties does not exist in a vacuum—it feeds into broader diagnostic and instructional frameworks that shape how educational professionals serve struggling readers. Two particularly influential frameworks deserve mention: the Simple View of Reading (SVR) and the Multi-Tiered Systems of Support (MTSS) model. The SVR posits that reading comprehension is the product of decoding and linguistic comprehension (R = D × LC), which means that a deficit in either component—or both—can produce a comprehension difficulty. Norm-referenced and informal assessments help diagnosticians determine whether the breakdown resides in decoding, language comprehension, or both, directly informing the type of intervention required. Within the MTSS framework, assessment data drives tier placement and movement: universal screening (Tier 1), targeted diagnostic assessment and intervention (Tier 2), and intensive, individualized evaluation and support (Tier 3).

Linking lesson concepts to the Simple View of Reading and MTSS frameworks
Concept in This LessonAdvanced Framework Connection
Norm-referenced screening identifies students below benchmarkMTSS Tier 1 → Tier 2 decision point; SVR helps determine if screening deficit is in D, LC, or both
Informal assessments reveal specific skill deficitsSVR component analysis: decoding difficulties → phonics intervention; linguistic comprehension difficulties → vocabulary/language intervention; MTSS Tier 2 intervention design
Integrated literacy profile with convergent evidenceComprehensive evaluation for special education eligibility (Tier 3); differential diagnosis of dyslexia vs. language-based learning disability vs. reading difficulty due to limited English proficiency
Progress monitoring with informal CBMsData-based individualization (DBI) within MTSS; rate of improvement compared to norm-referenced benchmarks determines whether to intensify, maintain, or fade intervention

As you advance in your study of literacy assessment, you will encounter increasingly nuanced diagnostic models, including the Dual-Route Cascaded Model of word reading, Scarborough's Reading Rope, and Ehri's phases of word reading development. Each of these theoretical frameworks adds layers of specificity to the diagnostic process, but all depend fundamentally on the ability to relate norm-referenced and informal assessment results to the observable manifestations of literacy difficulty—the very competency this lesson targets.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain why a norm-referenced reading comprehension score at the 10th percentile, by itself, is insufficient for diagnosing the specific nature of a student's literacy difficulty. What additional information is needed, and why?
PROBLEM 2BASIC CALCULATION
A student receives a standard score of 73 on the WJ-IV Word Attack subtest (M = 100, SD = 15). Calculate how many standard deviations below the mean this score falls, and determine the approximate percentile rank using the normal distribution. What level of concern does this suggest?
PROBLEM 3INTERMEDIATE
A fourth-grade student's assessment profile shows: WJ-IV Letter-Word Identification SS = 95 (37th percentile), Word Attack SS = 76 (5th percentile), GORT-5 Fluency SS = 80 (9th percentile), Passage Comprehension SS = 82 (12th percentile), and PPVT-5 SS = 110 (75th percentile). An informal running record reveals 93% accuracy at grade level with errors concentrated on unfamiliar multisyllabic words, and no self-corrections. Retelling of a passage read aloud to the student was detailed and accurate. What specific literacy difficulty is indicated, and what is the evidence for your conclusion?
PROBLEM 4APPLIED
You are a reading specialist reviewing data for two fifth-graders. Student A has the following scores: WIAT-4 Word Reading SS = 70, Pseudoword Decoding SS = 68, Reading Comprehension SS = 72, Oral Language SS = 102. Student B's scores are: Word Reading SS = 97, Pseudoword Decoding SS = 95, Reading Comprehension SS = 72, Oral Language SS = 78. Both have identical comprehension scores. Design an informal assessment plan for each student that would confirm or clarify the hypothesized root cause of their comprehension difficulty, and explain how the informal data would guide different intervention plans.
PROBLEM 5CRITICAL THINKING
A colleague argues that informal assessments are subjective and unreliable, and therefore only norm-referenced test scores should be used to identify and classify literacy difficulties. Construct a systematic rebuttal that addresses issues of validity, diagnostic utility, cultural/linguistic bias, and instructional relevance. Under what conditions, if any, might informal assessment data actually be more valid than norm-referenced data for identifying a specific student's literacy needs?

Summary

Relating assessment results to literacy difficulties requires the systematic integration of two complementary data streams. Norm-referenced assessments—instruments like the WJ-IV, GORT-5, CTOPP-2, and PPVT-5—provide standardized scores (percentile ranks, standard scores, stanines) that quantify the severity of a student's difficulty relative to peers and satisfy requirements for eligibility decisions. Informal assessments—running records, miscue analyses, phonics inventories, retelling rubrics, and CBM probes—reveal specific skill-level strengths and weaknesses that norm-referenced instruments cannot isolate, enabling precise intervention targeting.

The diagnostic process follows a logical sequence: screen with norm-referenced tools, diagnose with informal measures, triangulate for convergent evidence, and build an integrated literacy profile that maps directly onto the five pillars of literacy (phonemic awareness, phonics, fluency, vocabulary, comprehension). The identical norm-referenced score can reflect different root causes—a principle that underscores why informal data is indispensable. Competent diagnosticians use pattern analysis across multiple data sources to ensure that assessment drives instruction and that every struggling reader receives the specific, evidence-based support they need.

Varsity Tutors • KPEERI • Relating Assessment Results to Difficulties — 5. Relate norm-referenced and informal assessment results to literacy difficulties