EPPP: Part 1, Knowledge Quiz: Test Construction
20 questions · exam conditions
0:00
Test ConstructionQuestion 1 of 20

A normative sample overrepresents high-income students. A low-income student's percentile will likely appear

Higher than it should be
Unaffected if scores are valid
Lower than it should be
Higher only on speeded tests
← Back to quizzes

EPPP: Part 1, Knowledge Quiz

EPPP: Part 1, Knowledge Quiz: Test Construction

Practice Test Construction in EPPP: Part 1, Knowledge with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Test Construction, giving you a quick way to practice the rules, question types, and explanations that matter most for EPPP: Part 1, Knowledge.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A normative sample overrepresents high-income students. A low-income student's percentile will likely appear

  1. Higher than it should be
  2. Unaffected if scores are valid
  3. Lower than it should be (correct answer)
  4. Higher only on speeded tests
Explanation: Because the norm group contains too many high-income students, their higher scores push the percentile benchmarks up. A low-income student's score is then compared against an inflated standard, so their percentile looks lower than it would be with a representative sample. The tempting error is thinking overrepresentation inflates the student's percentile, but it actually raises the comparison group, not the student's score.

Question 2

A norm-referenced item has item p = 0.90 and discrimination index = 0.05. The best action is to

  1. Retain; mastery is useful
  2. Delete; item is too difficult
  3. Retain; it maximizes variance
  4. Revise; easy, poor discrim. (correct answer)
Explanation: With p = 0.90, the item is very easy, and a discrimination index of 0.05 means it barely separates high and low scorers, so it adds little to a norm-referenced test. Retaining it for mastery is tempting, but mastery is a criterion-referenced purpose, not the goal here. The item should be revised or replaced.

Question 3

A test has 90% sensitivity and 95% specificity; prevalence is 2%. What is the approximate positive predictive value?

  1. 27% (correct answer)
  2. 90%
  3. 95%
  4. 2%
Explanation: In 10,000 people, 200 have the disease and 180 test positive; 9,800 do not, and 5% of them, or 490, test false positive. So positive tests total 670, and 180/670 is about 27%. The tempting 95% mistakes specificity for predictive value, ignoring that false positives overwhelm true positives when prevalence is low.

Question 4

In a normal distribution, a raw score at the 84th percentile corresponds to which T-score?

  1. 40
  2. 50
  3. 60 (correct answer)
  4. 70
Explanation: The 84th percentile is 1 standard deviation above the mean. T-scores have a mean of 50 and a standard deviation of 10, so 1 SD above the mean is 50 + 10 = 60. A T-score of 70 would be 2 SDs above the mean, about the 97.7th percentile, not the 84th. The tempting error is choosing 50, but that is the mean, which is the 50th percentile.

Question 5

Item A: top p = 0.80, bottom p = 0.55. Item B: top p = 0.70, bottom p = 0.65. Which discriminates better?

  1. Item B; D = 0.05
  2. Item A; D = 0.25 (correct answer)
  3. Item A; D = 0.55
  4. Item B; D = 0.70
Explanation: Discrimination is the difference between top and bottom p values. Item A: 0.80 - 0.55 = 0.25. Item B: 0.70 - 0.65 = 0.05. The larger the gap, the better the item discriminates, so Item A wins. A tempting wrong choice is Item B because its two p values look more balanced, but balance means less separation, not better discrimination.

Question 6

During item analysis, which statistic reflects the proportion answering correctly?

  1. Standard error of measurement
  2. Test-retest reliability
  3. Content validity evidence
  4. Item difficulty index (correct answer)
Explanation: This question tests the application of item analysis, standardization, norming, sensitivity, and specificity in psychological assessments. These concepts are crucial for ensuring the reliability and validity of psychological tests. Item analysis, for example, evaluates items based on difficulty and discrimination to refine test quality. In this question, understanding the role of item difficulty index is key, as it ensures the proportion of correct responses is quantified for test refinement. The correct answer works because it accurately identifies the item difficulty index and its importance in test construction. A common distractor fails because it confuses item analysis with reliability measures like test-retest, which is a frequent misunderstanding. Teaching strategies include emphasizing the differences between these concepts through examples and encouraging practice in applying them to real-world scenarios, which helps in identifying subtle distinctions and preventing common errors.

Question 7

In establishing norms for a new personality inventory, researchers stratify their sample based on age, education, and geographic region, then recruit participants within each stratum. What type of sampling strategy does this represent?

  1. Systematic random sampling with predetermined intervals to ensure representative demographic coverage
  2. Stratified random sampling designed to ensure proportional representation across key demographic variables (correct answer)
  3. Cluster sampling utilizing naturally occurring groups to efficiently gather normative data from target populations
  4. Convenience sampling enhanced with demographic quotas to approximate population representativeness
Explanation: Stratified random sampling involves dividing the population into subgroups (strata) based on important characteristics (age, education, geographic region), then randomly sampling within each stratum. This ensures representation across key demographic variables for norming purposes. A: Incorrect - systematic sampling uses predetermined intervals, not stratification. C: Incorrect - cluster sampling uses naturally occurring groups, not artificial strata. D: Incorrect - the description indicates systematic stratification, not convenience sampling with quotas.

Question 8

An item analysis reveals that for Item 25, 80% of high-scoring examinees answered correctly while 45% of low-scoring examinees answered correctly. What is the discrimination index, and how should this item be evaluated?

  1. D=0.35D = 0.35; the item shows acceptable discrimination and should be retained in the test (correct answer)
  2. D=0.625D = 0.625; the item shows excellent discrimination and represents an optimal test item
  3. D=1.25D = 1.25; the item shows exceptional discrimination exceeding typical measurement standards
  4. D=0.45D = 0.45; the item shows good discrimination though difficulty level may need adjustment
Explanation: Discrimination index D = (proportion correct in high group) - (proportion correct in low group) = 0.80 - 0.45 = 0.35. This indicates acceptable discrimination (threshold typically 0.30+), meaning the item effectively differentiates between high and low ability examinees. Items with D ≥ 0.30 are generally retained. B: Incorrect calculation (0.80 × 0.45 rather than subtraction). C: Incorrect calculation and impossible value (D cannot exceed 1.0). D: Incorrect - uses low group proportion rather than calculating difference.

Question 9

A diagnostic test for anxiety disorders yields a sensitivity of 0.72 and specificity of 0.88 in validation studies. When implemented in a clinical setting where anxiety prevalence is 25%, what does this suggest about the test's clinical utility?

  1. The test will miss a substantial portion of anxiety cases (28% false negative rate), which poses significant clinical risks (correct answer)
  2. The test demonstrates balanced accuracy with equal rates of false positive and false negative errors across populations
  3. The test shows optimal diagnostic accuracy with minimal classification errors across both affected and unaffected individuals
  4. The test will generate excessive false positives leading to overdiagnosis and unnecessary treatment in this clinical setting
Explanation: With 25% prevalence and 72% sensitivity, the false negative rate is 28% (1 - 0.72), meaning more than 1 in 4 individuals with anxiety will be missed by the test. This represents a clinically significant limitation that could delay appropriate treatment. While false positives also occur (12% of non-anxious individuals), the primary concern is the substantial number of missed cases. B: Incorrect - error rates aren't equal. C: Incorrect - missing 28% of cases isn't optimal accuracy. D: Incorrect - false negatives are the primary clinical concern here.

Question 10

During test construction, researchers establish norms using a sample that includes 40% White, 25% Hispanic/Latino, 20% Black/African American, 10% Asian American, and 5% other ethnicities. If this differs from census proportions, what is the most likely reason for this sampling strategy?

  1. To ensure adequate statistical power for detecting ethnic differences in test performance across groups
  2. To create representative norms that accurately reflect the current U.S. population demographic distribution
  3. To oversample minority groups for bias analysis while planning to weight data for final norms (correct answer)
  4. To establish separate ethnic norms that can be applied independently for each demographic group
Explanation: Deliberate oversampling of minority groups (if this differs from census data) is commonly done to ensure sufficient sample sizes for bias analyses, item functioning studies, and validity research across ethnic groups. The data can then be statistically weighted to match population proportions for final norm development. A: While power is a consideration, the primary purpose is bias analysis. B: If this differs from census data, it's not representative without weighting. D: The goal is typically unified norms with bias analysis, not separate ethnic norms.

Question 11

During test construction, an item analysis shows that Item 18 has a difficulty index of 0.25 and a discrimination index of 0.35. What can be concluded about this item's psychometric properties?

  1. The item is too difficult and poorly discriminating, requiring significant revision or elimination from the test
  2. The item is appropriately difficult with acceptable discrimination, suitable for inclusion in the final test version
  3. The item shows optimal difficulty and excellent discrimination, representing an ideal test item
  4. The item is moderately difficult with good discrimination, though difficulty could be reduced slightly (correct answer)
Explanation: A difficulty index of 0.25 means the item is moderately difficult (75% got it wrong), while a discrimination index of 0.35 indicates good discrimination (acceptable threshold is typically 0.30+). The item functions well psychometrically, though items with difficulty closer to 0.50 often provide maximum information. A: Incorrect - 0.35 discrimination is acceptable. B: Incorrect - difficulty is not optimal for maximum discrimination. C: Incorrect - difficulty isn't optimal and discrimination, while good, isn't excellent.

Question 12

A test developer wants to establish age-based norms for a cognitive battery. The standardization sample includes 100 participants each in age ranges 20-29, 30-39, 40-49, 50-59, and 60-69. What potential limitation does this sampling approach present?

  1. The total sample size is insufficient to detect meaningful age-related differences in cognitive performance
  2. The age ranges are too broad to capture gradual developmental changes within each decade
  3. The equal sample sizes across age groups don't reflect actual population proportions in these ranges (correct answer)
  4. The age-based stratification fails to account for cohort effects and generational differences in test performance
Explanation: Using equal sample sizes (100 per group) across age ranges doesn't reflect actual population demographics, where younger age groups are typically larger than older ones. This creates norms that may not appropriately represent the population structure. For accurate norm-referenced interpretation, sample proportions should reflect population proportions. A: Incorrect - 500 total participants with 100 per group is generally adequate. B: Incorrect - 10-year ranges are standard for many cognitive batteries. D: Incorrect - while cohort effects exist, the primary issue here is sample proportion representation.

Question 13

A test constructor finds that the normative sample for a new intelligence test overrepresents individuals with college degrees compared to census data. What is the most appropriate corrective action during the norming process?

  1. Apply statistical weighting procedures to adjust the sample data to match population education demographics (correct answer)
  2. Increase the overall sample size proportionally while maintaining the current education distribution patterns
  3. Create separate norms for college-educated and non-college-educated populations to address the bias
  4. Report the limitation in the manual while proceeding with the current sample for norm development
Explanation: Statistical weighting adjusts the contribution of overrepresented groups to match population proportions, correcting the bias without requiring new data collection. This is a standard and appropriate method for addressing demographic imbalances in normative samples. B: Incorrect - increasing sample size doesn't fix the demographic imbalance. C: Incorrect - separate norms aren't necessary when weighting can correct the issue. D: Incorrect - this doesn't address the bias, just acknowledges it.

Question 14

A test developer conducts item analysis and finds that Item 47 has a point-biserial correlation of -0.15 with the total test score. What does this finding most likely indicate about this item?

  1. The item is functioning well and should be retained in its current form for future administrations
  2. The item is negatively discriminating and should be revised or eliminated from the test (correct answer)
  3. The item has optimal difficulty level and contributes appropriately to overall test reliability
  4. The item demonstrates acceptable criterion validity and meets psychometric standards for inclusion
Explanation: A negative point-biserial correlation (-0.15) indicates that examinees who scored higher on the overall test tended to get this item wrong, while those who scored lower on the overall test tended to get it right. This suggests the item is negatively discriminating and is likely flawed (e.g., keyed incorrectly, poorly written, or measuring something different from the rest of the test). Such items should be revised or eliminated. A: Incorrect - negative discrimination indicates poor item functioning. C: Incorrect - point-biserial correlation measures discrimination, not difficulty level. D: Incorrect - negative discrimination suggests poor validity contribution.

Question 15

A cognitive assessment item has a difficulty index of 0.15 and discrimination index of 0.42. What does this pattern of statistics suggest about the item's characteristics and utility?

  1. The item is very difficult but discriminates well, making it useful for identifying high-ability examinees (correct answer)
  2. The item has optimal difficulty with excellent discrimination, representing an ideal test component
  3. The item is too easy with poor discrimination, requiring revision to improve psychometric properties
  4. The item shows moderate difficulty with acceptable discrimination, suitable for inclusion without changes
Explanation: A difficulty index of 0.15 means only 15% answered correctly (very difficult), while a discrimination index of 0.42 indicates good discrimination (well above 0.30 threshold). Very difficult items with good discrimination are valuable for assessing high-ability examinees and extending the test's ceiling. They provide information at the upper end of the ability distribution. B: Incorrect - difficulty isn't optimal (0.50 would be). C: Incorrect - this describes easy items, not difficult ones. D: Incorrect - 0.15 isn't moderate difficulty.

Question 16

A psychological test manual reports coefficient alpha of 0.91, test-retest reliability of 0.85, and standard error of measurement of 3.2 points. What do these statistics collectively suggest about the test's measurement precision?

  1. The test demonstrates excellent internal consistency but shows concerning temporal instability over repeated administrations
  2. The test exhibits strong reliability characteristics with good precision, though some measurement error remains (correct answer)
  3. The test shows acceptable reliability but the standard error indicates poor precision for individual score interpretation
  4. The test demonstrates optimal psychometric properties with minimal measurement error affecting score accuracy
Explanation: All three statistics indicate good reliability: coefficient alpha of 0.91 shows excellent internal consistency, test-retest of 0.85 shows good temporal stability, and SEM of 3.2 indicates reasonable precision (exact interpretation depends on scale, but these values collectively suggest strong psychometric properties with acknowledgment that measurement error exists). A: Incorrect - 0.85 test-retest is good, not concerning. C: Incorrect - these reliability coefficients are strong. D: Incorrect - no test has minimal measurement error; some error always exists.

Question 17

A screening test for learning disabilities has been validated with sensitivity of 0.84 and specificity of 0.76. In a school district where learning disabilities have a 12% prevalence rate, what is the most significant limitation of using this test?

  1. The low specificity will result in numerous false positives, potentially mislabeling many typical learners (correct answer)
  2. The moderate sensitivity will miss a substantial proportion of students who actually have learning disabilities
  3. The test lacks sufficient discriminative power to differentiate between different types of learning disabilities
  4. The validation sample may not adequately represent the demographic characteristics of this school district
Explanation: With 12% prevalence and 76% specificity, 24% of students without learning disabilities (88% of population) will test positive: 0.24 × 0.88 = 21.1% false positive rate. This means over 1 in 5 typical learners will be incorrectly identified, creating significant resource allocation problems and potential mislabeling. While the test will miss some true cases (16% false negative rate), the false positive problem is more substantial. B: While 16% will be missed, the false positive rate is the bigger issue. C: The question asks about screening, not differential diagnosis. D: While possible, this isn't directly indicated by the given information.

Question 18

A test development team creates norms based on a sample recruited through online panels, resulting in overrepresentation of individuals with internet access and technology familiarity. What statistical approach would best address this sampling limitation?

  1. Calculate separate reliability coefficients for technology-familiar and technology-unfamiliar subgroups
  2. Apply post-stratification weighting to match known population characteristics on relevant demographics (correct answer)
  3. Increase the total sample size while maintaining the current recruitment methodology and procedures
  4. Develop alternative norms specifically designed for populations with limited technology access and familiarity
Explanation: Post-stratification weighting adjusts the sample to match known population characteristics (age, education, income, geographic distribution, etc.) that correlate with internet access and technology familiarity. This statistical correction helps address the bias introduced by the recruitment method without requiring complete re-norming. A: Incorrect - reliability coefficients don't address norm bias. C: Incorrect - larger samples with the same bias don't solve the representativeness problem. D: Incorrect - alternative norms aren't necessary when weighting can address the bias.

Question 19

A test manual reports that during standardization, the research team collected data from 1,500 participants across 15 different geographic regions, with recruitment continuing until each region contributed exactly 100 participants. What potential bias might this approach introduce into the normative data?

  1. Urban-rural bias due to differential accessibility and recruitment challenges across different geographic regions
  2. Temporal bias resulting from extended data collection periods with potential practice effects across regions
  3. Population density bias by giving equal weight to regions with vastly different population sizes (correct answer)
  4. Socioeconomic bias due to systematic differences in volunteer participation rates across geographic boundaries
Explanation: Equal sample sizes from each region (100 per region) don't reflect actual population distributions. Large metropolitan regions with millions of residents receive the same weight as small rural regions with thousands of residents, creating bias where less populated areas are overrepresented in the norms. A: While possible, this isn't the primary bias introduced by equal regional sampling. B: Temporal bias isn't necessarily created by geographic stratification. D: While socioeconomic differences may exist, the main issue is population proportionality.

Question 20

During item analysis, a test developer calculates that Item 32 has a point-biserial correlation of 0.45 with the total test score. What does this statistic primarily indicate about the item's performance?

  1. The item has moderate difficulty level and is answered correctly by approximately 45% of examinees
  2. The item demonstrates good discrimination, with higher-scoring examinees more likely to answer correctly (correct answer)
  3. The item shows acceptable test-retest reliability over the specified time interval for this assessment
  4. The item exhibits strong concurrent validity when compared against established criterion measures in this domain
Explanation: A point-biserial correlation of 0.45 indicates good discrimination - there's a moderate-to-strong positive relationship between getting this item correct and scoring higher on the total test. This means examinees who score higher overall are more likely to get this item right, which is desirable for test items. A: Incorrect - point-biserial correlation measures discrimination, not difficulty. C: Incorrect - this statistic doesn't assess reliability over time. D: Incorrect - this measures internal consistency, not external criterion validity.