All questions
Question 1
A researcher wants to create the most reliable test possible from a pool of 100 items all measuring a single psychological construct. Which of the following strategies would most likely achieve this goal?
- Select the 20 items that have the highest average correlation with all other items in the pool. (correct answer)
- Select the 20 items that show the greatest variability in scores across the population.
- Select the 20 items that have the highest face validity according to a panel of students.
- Select a diverse set of 20 items that show low correlations with each other to cover the construct broadly.
Explanation: When you encounter questions about test reliability, remember that reliability refers to the consistency and stability of measurement. A reliable test produces similar results when administered repeatedly or when items measure the same underlying construct.
Answer A is correct because selecting items with the highest average correlation with all other items maximizes internal consistency. When test items correlate strongly with each other, they're all tapping into the same psychological construct effectively. This creates what psychometricians call high "inter-item correlation," which directly increases Cronbach's alpha—the most common measure of internal consistency reliability. Items that correlate well with the total pool are proven to measure the target construct reliably.
Answer B is wrong because high variability in scores doesn't guarantee reliability. While you need some variability to discriminate between people, items with extreme variability might be measuring random error or multiple constructs rather than your target construct consistently.
Answer C is incorrect because face validity (whether items appear to measure what they claim) doesn't ensure reliability. An item can look relevant to students but still produce inconsistent measurements. Face validity is about appearance, not psychometric quality.
Answer D represents a common misconception. While you want broad coverage of a construct, selecting items that show low correlations with each other actually decreases reliability. Low inter-item correlations suggest the items are measuring different things, which reduces internal consistency.
Remember this key principle: for reliability, you want items that "hang together" statistically. High inter-item correlations indicate your items are consistently measuring the same construct.
Question 2
A clinical psychologist measures a client's level of depression before and after a course of therapy using a standardized inventory. To quantify the therapy's effect, the psychologist calculates a 'change score' by subtracting the post-therapy score from the pre-therapy score. If the depression inventory itself has a reliability coefficient of .80, what can be concluded about the reliability of the calculated change score?
- The change score will be more reliable than the scores from the individual administrations.
- The reliability of the change score will be lower than the reliability of the inventory. (correct answer)
- The reliability of the change score will be equal to the reliability of the inventory.
- The change score is a direct measure of validity, so its reliability cannot be determined.
Explanation: The correct answer is B. A well-known issue in psychometrics is that change (or difference) scores are less reliable than the scores from which they are calculated. This is because the measurement error present in both the pre-test and the post-test scores is compounded in the difference score. The error from the first measurement and the error from the second measurement both contribute to the error in the change score, reducing its overall reliability. Thus, the reliability of the change score will be less than .80.
- A and C are incorrect because they contradict this principle; change scores are systematically less reliable.
- D is incorrect because the change score is itself a measurement that has psychometric properties, including reliability. Its validity as a measure of therapeutic change is dependent on its reliability.
Question 3
A new scale to measure anxiety shows high internal consistency (Cronbach's α = .92) and high test-retest reliability over two weeks (r = .88). However, scores on the scale correlate weakly with scores on a well-established, validated measure of anxiety (r = .20) but correlate strongly with scores on a measure of social desirability (r = .75). What is the most appropriate conclusion about the new scale?
- The scale is valid as a measure of anxiety but is not a reliable instrument.
- The scale is reliable, but its construct validity as a measure of anxiety is questionable. (correct answer)
- The scale demonstrates high predictive validity but suffers from low content validity.
- The scale has poor inter-rater reliability, which accounts for the low correlation with other measures.
Explanation: The correct answer is B. The high internal consistency and test-retest coefficients indicate that the scale is reliable; it consistently measures something. However, its weak correlation with an established anxiety measure suggests poor convergent validity, and its strong correlation with an unrelated construct (social desirability) suggests poor discriminant validity. Both convergent and discriminant validity are critical components of construct validity, which is the extent to which a test measures the theoretical construct it purports to measure. Therefore, the scale's construct validity as a measure of anxiety is highly questionable.
- A is incorrect because the evidence clearly points to high reliability, not low reliability. Validity cannot exist without reliability.
- C is incorrect because the provided data does not contain information about the scale's ability to predict future outcomes (predictive validity) or how well its content represents the domain of anxiety (content validity).
- D is incorrect because the scenario does not involve different raters or observers, so inter-rater reliability is not relevant to the data presented.
Question 4
A test can be without being , but it cannot be without being . Which pair of terms correctly fills the blanks?
- reliable; valid; valid; reliable (correct answer)
- valid; reliable; reliable; valid
- fair; valid; valid; fair
- reliable; fair; fair; reliable
Explanation: This question tests your understanding of the fundamental relationship between reliability and validity in psychological testing. These are two crucial psychometric properties that determine test quality, but they have a specific hierarchical relationship.
Reliability refers to consistency—whether a test produces the same results when administered repeatedly under similar conditions. Validity refers to accuracy—whether a test actually measures what it claims to measure. The key insight is that reliability is a prerequisite for validity, but not vice versa.
A test can be reliable without being valid because it can consistently measure the wrong thing. For example, a scale that consistently reads 5 pounds too heavy is reliable (consistent) but not valid (inaccurate). However, a test cannot be valid without being reliable because if measurements are inconsistent, they cannot be accurately measuring the intended construct. You can't hit the bullseye if your shots are scattered randomly.
Looking at the wrong answers: Option B reverses this relationship, incorrectly suggesting reliability requires validity. Option C introduces "fairness," which relates to bias and cultural sensitivity—a different concept entirely from the reliability-validity relationship. Option D also uses fairness incorrectly and reverses the proper relationship between reliability and validity.
Remember this hierarchy: reliability is necessary but not sufficient for validity. Think of reliability as the foundation—without it, validity cannot exist. This relationship appears frequently on psychology exams, so always ask yourself: "What's the prerequisite?" when evaluating test quality concepts.
Question 5
An applicant scores 112 on a selection test that has a Standard Error of Measurement (SEM) of 4. The company has a strict cutoff score of 115 for hiring. Based on these psychometric data, which is the most appropriate and fair interpretation of the applicant's performance?
- The applicant's true score is exactly 112, so they are not qualified for the position.
- The SEM of 4 indicates the test is too unreliable to be used for hiring decisions.
- The applicant's score is definitively below the cutoff, and they should not be hired.
- The applicant's true score could plausibly be at or above the cutoff due to measurement error. (correct answer)
Explanation: The correct answer is D. The SEM represents the average amount of error in an individual's score. It is used to create a confidence interval around the observed score to estimate the range in which the individual's 'true' score likely falls. A 95% confidence interval is typically calculated as the observed score ± (1.96 * SEM). In this case, approximately 112 ± (2 * 4), which is 104 to 120. Since this range includes and surpasses the cutoff of 115, it is plausible that the applicant's true ability level is high enough to qualify. Using a strict cutoff without considering the SEM is often considered unfair because it ignores the inherent imprecision of psychological testing.
- A and C are incorrect because they treat the observed score as a perfectly precise measure, ignoring the reality of measurement error quantified by the SEM.
- B is incorrect because an SEM of 4 does not, by itself, indicate that a test is too unreliable. The reliability depends on the variance of the total test scores (SEM = SD * sqrt(1-reliability)). Without the standard deviation, we cannot determine the reliability, but the key issue is how to interpret the score given the known error.
Question 6
A developer of a new scale measuring 'Grit' finds that scores are strongly correlated with the number of hours participants voluntarily spend on a frustrating puzzle (r = .65) and weakly correlated with scores on an IQ test (r = .10). These two findings, respectively, offer primary support for which aspects of the scale's validity?
- Content validity and test-retest reliability
- Predictive validity and concurrent validity
- Convergent validity and discriminant validity (correct answer)
- Internal consistency and face validity
Explanation: The correct answer is C. These findings relate to construct validity. Convergent validity is demonstrated when a test correlates highly with other variables or measures with which it should theoretically be related. The strong correlation with a behavioral measure of persistence (hours on a puzzle) supports the scale's convergent validity. Discriminant (or divergent) validity is demonstrated when a test has a low correlation with measures of constructs that are theoretically different. The weak correlation with IQ, a distinct construct, supports the scale's discriminant validity, showing it isn't just another measure of intelligence.
- A is incorrect because content validity is assessed by expert judgment of item content, and test-retest reliability is assessed by administering the test at two different times.
- B is incorrect because the puzzle task is likely measured at the same time as the grit scale, making it evidence for concurrent, not predictive, validity. More importantly, the pair of findings perfectly illustrates the convergent/discriminant distinction.
- D is incorrect because internal consistency and face validity are different psychometric properties not addressed by these correlational findings.
Question 7
A researcher develops a 30-item survey to measure 'Intellectual Humility.' To provide evidence for the survey's construct validity, they conduct a factor analysis. The analysis reveals that the items load onto three distinct, weakly correlated factors ('Openness to Revision,' 'Respect for Others' Viewpoints,' and 'Awareness of Fallibility'). What is the most significant implication of this finding for the survey?
- The survey lacks construct validity because 'Intellectual Humility' may be multidimensional, not unitary. (correct answer)
- The survey demonstrates high internal consistency because all items relate to a broader theme.
- The survey has strong evidence of convergent validity with other established personality traits.
- The survey's items should be revised to improve their test-retest reliability over time.
Explanation: When evaluating construct validity through factor analysis, you're examining whether your survey actually measures what it claims to measure as a unified concept. Factor analysis reveals the underlying structure of your data by showing how items cluster together.
The key finding here is that the 30 items loaded onto three distinct, weakly correlated factors rather than one strong factor. This suggests that "Intellectual Humility" isn't a single, unitary construct but rather a multidimensional one with separate components. When factors are weakly correlated, it indicates these dimensions operate somewhat independently of each other. This challenges the assumption that all items measure the same underlying trait, which is problematic for construct validity if you intended to measure intellectual humility as a single concept.
Looking at the wrong answers: Choice B misunderstands internal consistency—having three separate factors actually suggests lower internal consistency, not higher, since items aren't all measuring the same thing. Choice C confuses convergent validity (correlation with similar measures) with factor structure—the analysis tells us nothing about how this survey relates to other established traits. Choice D incorrectly focuses on test-retest reliability, which concerns stability over time, not factor structure.
The correct answer is A because discovering multiple weakly correlated factors indicates the survey may lack construct validity for measuring intellectual humility as a unitary trait.
Study tip: When you see factor analysis questions, focus on what the factor structure reveals about the construct being measured. Multiple factors suggest multidimensionality, which can challenge construct validity if you expected a unitary trait.
Question 8
A school district creates a final exam for its mandatory world history course. The committee that writes the exam questions is composed entirely of specialists in 19th-century European history. The resulting exam is found to be highly reliable. However, some teachers argue the exam is unfair to students. What is the primary psychometric basis for this concern about fairness?
- The test's internal consistency is likely low because the items are too similar.
- The test likely lacks content validity because it does not sample the entire curriculum domain. (correct answer)
- The test has poor predictive validity for success in college-level history courses.
- The test is not standardized, making comparisons between students unreliable.
Explanation: The correct answer is B. Content validity refers to the extent to which a test's items adequately represent the entire content domain it is intended to measure. Since the test is for a world history course but was written exclusively by specialists in European history, it is very likely that the test oversamples European history and undersamples history from other parts of the world. This lack of representative sampling is a deficit in content validity. It is unfair to students because their grades depend on a narrow slice of the curriculum, penalizing those who mastered other required material.
- A is incorrect because if the items are very similar (e.g., all about 19th-century Europe), the internal consistency (a measure of reliability) is likely to be high, not low. The stem even states the exam is reliable.
- C is incorrect because there is no information provided to assess the test's predictive validity.
- D is incorrect because the problem described is about the content of the test, not the procedures for administration and scoring (standardization).
Question 9
An aptitude test is used for university admissions. It is discovered that the test consistently under-predicts the first-year college GPA for students from rural backgrounds; that is, their actual GPAs are higher than what the test scores would predict. For students from urban backgrounds, the test accurately predicts GPA. This differential prediction pattern is a clear example of:
- adverse impact, because it leads to lower admission rates for rural students.
- low test-retest reliability, because scores are not stable over time for rural students.
- test bias, because the test's predictive validity is different for the two groups. (correct answer)
- poor content validity, because the test items are irrelevant to the college curriculum.
Explanation: The correct answer is C. Test bias, from a psychometric standpoint, is most clearly demonstrated by differential predictive validity. When a test's scores predict an outcome (like GPA) differently for different subgroups, the test is considered biased. In this case, the test is less valid for predicting the performance of rural students than for urban students, systematically underestimating their potential. This is a classic example of predictive bias, a threat to test fairness.
- A is incorrect because adverse impact refers to differences in selection rates, which might result from this situation but is not the name for the psychometric property itself. The core issue is the differential validity, which is bias.
- B is incorrect because the issue described is about the test's accuracy in prediction (validity), not its consistency over time (reliability).
- D is incorrect because content validity concerns the relevance of the test's content to the domain being measured, not its ability to predict future outcomes for different groups.
Question 10
A researcher creates a 100-item test of historical knowledge. To evaluate the test, they calculate the correlation between the total scores on the odd-numbered items and the total scores on the even-numbered items. This procedure is performed to gather evidence about the test's:
- predictive validity, by using one half of the test to forecast scores on the other.
- content validity, by ensuring both halves of the test cover similar material.
- test-retest reliability, by treating the two halves as separate testing sessions.
- internal consistency reliability, by assessing the homogeneity of the test items. (correct answer)
Explanation: The correct answer is D. This procedure describes the split-half method (specifically, an odd-even split), which is a way to measure a test's internal consistency. Internal consistency reliability refers to the degree to which all the items on a test measure the same underlying construct. By correlating two halves of the test, the researcher is checking if the items are consistent with one another.
- A is incorrect because predictive validity involves correlating test scores with a future criterion, not with another part of the same test.
- B is incorrect because content validity is established through expert review of the items against a content blueprint, not through statistical correlation of item subsets.
- C is incorrect because test-retest reliability requires administering the entire test at two different points in time to the same group of people.
Question 11
An elite graduate program admits only students with very high scores on a standardized admissions test. Program administrators conduct a study to assess the test's predictive validity by correlating the admitted students' test scores with their first-year grades. They find a correlation close to zero. What is the most likely reason for this unexpectedly low validity coefficient?
- The low correlation proves that the admissions test is fundamentally invalid.
- The admissions test has low test-retest reliability for this high-scoring population.
- The range of scores on the predictor variable has been severely restricted. (correct answer)
- The grading system for the first-year courses lacks inter-rater reliability.
Explanation: The correct answer is C. This scenario is a classic example of the effect of 'restriction of range.' The correlation coefficient's magnitude depends on the variability of both variables. By only including students with very high test scores, the program has restricted the range of the predictor variable (test scores). This lack of variability will statistically attenuate, or weaken, the observed correlation, even if a strong relationship exists in the full population of applicants. The low correlation is likely an artifact of the selected sample, not an indication that the test is invalid in general.
- A is incorrect because the finding is likely a statistical artifact, not proof of the test's invalidity for the broader population.
- B is incorrect because major standardized tests are known to have high reliability, and there is no reason to assume it would be lower for a specific subgroup.
- D is incorrect because while it's a possibility in any validity study, restriction of range is a more direct and powerful explanation for the specific situation described (a validity study on a highly selective group).
Question 12
A state implements a new high-stakes exam that all students must pass to graduate. The exam is shown to be a psychometrically reliable and valid measure of the state's official curriculum standards. However, critics argue its use is unfair because schools in low-income districts are under-resourced and cannot adequately prepare their students to meet these standards. This criticism highlights a conflict concerning:
- the test's predictive validity, as it fails to predict life success for all students.
- the test's content validity, because the curriculum standards must be flawed.
- the social consequences of using a valid test in a system with unequal opportunities. (correct answer)
- the test's reliability, because scores from students in low-income districts are inconsistent.
Explanation: The correct answer is C. This question addresses the important distinction between the technical, psychometric properties of a test and the broader ethical and social justice implications of its use. The stem stipulates that the test is reliable and valid—it accurately measures what it's supposed to measure. The criticism is not about the test itself, but about the fairness of applying its results as a high-stakes hurdle in a context of systemic inequality. This is a core issue in educational policy and testing ethics: even a 'good' test can produce unfair outcomes if the opportunity to learn the material is not equally distributed.
- A, B, and D are incorrect because they challenge the psychometric properties of the test (validity and reliability), which the premise of the question asks you to accept as sound. The criticism described is about the application and consequences of the test, not its technical quality.
Question 13
A school district creates a final exam for its mandatory world history course. The committee that writes the exam questions is composed entirely of specialists in 19th-century European history. The resulting exam is found to be highly reliable. However, some teachers argue the exam is unfair to students. What is the primary psychometric basis for this concern about fairness?
- The test's internal consistency is likely low because the items are too similar.
- The test likely lacks content validity because it does not sample the entire curriculum domain. (correct answer)
- The test has poor predictive validity for success in college-level history courses.
- The test is not standardized, making comparisons between students unreliable.
Explanation: The correct answer is B. Content validity refers to the extent to which a test's items adequately represent the entire content domain it is intended to measure. Since the test is for a world history course but was written exclusively by specialists in European history, it is very likely that the test oversamples European history and undersamples history from other parts of the world. This lack of representative sampling is a deficit in content validity. It is unfair to students because their grades depend on a narrow slice of the curriculum, penalizing those who mastered other required material.
- A is incorrect because if the items are very similar (e.g., all about 19th-century Europe), the internal consistency (a measure of reliability) is likely to be high, not low. The stem even states the exam is reliable.
- C is incorrect because there is no information provided to assess the test's predictive validity.
- D is incorrect because the problem described is about the content of the test, not the procedures for administration and scoring (standardization).
Question 14
A researcher develops a new personality test to measure conscientiousness. An initial study finds that scores on the test are highly stable when participants take it on two occasions six months apart. Based only on this finding, which of the following conclusions is psychometrically justified?
- The test is a valid measure of the trait of conscientiousness.
- The test is a reliable measure, but its validity remains to be determined. (correct answer)
- The test has high construct validity but may lack sufficient reliability.
- The test is a reliable measure, which is sufficient evidence for its validity.
Explanation: The correct answer is B. The finding that scores are stable over a six-month period is direct evidence of high test-retest reliability. Reliability refers to the consistency of a measure. However, reliability is a necessary but not sufficient condition for validity. A test can be highly reliable (i.e., consistently measure the same thing) without being valid (i.e., measuring what it is intended to measure). Therefore, based only on this finding, one can conclude the test is reliable, but no conclusion about its validity is justified yet.
- A is incorrect because reliability alone does not establish validity.
- C is incorrect because the finding directly supports reliability, not construct validity. In fact, a test cannot be valid if it is not reliable.
- D is incorrect because it makes the common error of equating reliability with validity. Reliability is a prerequisite for validity, not sufficient evidence of it.
Question 15
An aptitude test is used for university admissions. It is discovered that the test consistently under-predicts the first-year college GPA for students from rural backgrounds; that is, their actual GPAs are higher than what the test scores would predict. For students from urban backgrounds, the test accurately predicts GPA. This differential prediction pattern is a clear example of:
- adverse impact, because it leads to lower admission rates for rural students.
- low test-retest reliability, because scores are not stable over time for rural students.
- test bias, because the test's predictive validity is different for the two groups. (correct answer)
- poor content validity, because the test items are irrelevant to the college curriculum.
Explanation: The correct answer is C. Test bias, from a psychometric standpoint, is most clearly demonstrated by differential predictive validity. When a test's scores predict an outcome (like GPA) differently for different subgroups, the test is considered biased. In this case, the test is less valid for predicting the performance of rural students than for urban students, systematically underestimating their potential. This is a classic example of predictive bias, a threat to test fairness.
- A is incorrect because adverse impact refers to differences in selection rates, which might result from this situation but is not the name for the psychometric property itself. The core issue is the differential validity, which is bias.
- B is incorrect because the issue described is about the test's accuracy in prediction (validity), not its consistency over time (reliability).
- D is incorrect because content validity concerns the relevance of the test's content to the domain being measured, not its ability to predict future outcomes for different groups.
Question 16
An applicant scores 112 on a selection test that has a Standard Error of Measurement (SEM) of 4. The company has a strict cutoff score of 115 for hiring. Based on these psychometric data, which is the most appropriate and fair interpretation of the applicant's performance?
- The applicant's true score is exactly 112, so they are not qualified for the position.
- The SEM of 4 indicates the test is too unreliable to be used for hiring decisions.
- The applicant's score is definitively below the cutoff, and they should not be hired.
- The applicant's true score could plausibly be at or above the cutoff due to measurement error. (correct answer)
Explanation: The correct answer is D. The SEM represents the average amount of error in an individual's score. It is used to create a confidence interval around the observed score to estimate the range in which the individual's 'true' score likely falls. A 95% confidence interval is typically calculated as the observed score ± (1.96 * SEM). In this case, approximately 112 ± (2 * 4), which is 104 to 120. Since this range includes and surpasses the cutoff of 115, it is plausible that the applicant's true ability level is high enough to qualify. Using a strict cutoff without considering the SEM is often considered unfair because it ignores the inherent imprecision of psychological testing.
- A and C are incorrect because they treat the observed score as a perfectly precise measure, ignoring the reality of measurement error quantified by the SEM.
- B is incorrect because an SEM of 4 does not, by itself, indicate that a test is too unreliable. The reliability depends on the variance of the total test scores (SEM = SD * sqrt(1-reliability)). Without the standard deviation, we cannot determine the reliability, but the key issue is how to interpret the score given the known error.
Question 17
A researcher develops a scale to predict job burnout in nurses. They administer the scale to 500 nurses and find a strong correlation with a criterion measure of burnout. To ensure this finding was not specific to that particular sample, they administer the scale and criterion measure to a new, independent sample of 500 nurses and re-calculate the correlation. This second step is an example of:
- assessing test-retest reliability.
- performing a meta-analysis.
- establishing content validity.
- conducting a cross-validation. (correct answer)
Explanation: The correct answer is D. Cross-validation is the process of confirming that the findings from one study (e.g., a validity coefficient) can be replicated in an independent sample drawn from the same population. The initial validation study might have capitalized on chance characteristics of the first sample, leading to an inflated validity coefficient. By checking the result in a new sample, the researcher can have more confidence that the scale's predictive power is genuine and not a statistical fluke.
- A is incorrect because test-retest reliability involves re-administering the test to the same sample at a later time.
- B is incorrect because a meta-analysis is a statistical technique for combining the results of multiple different studies, not a step within a single research project.
- C is incorrect because content validity is about ensuring the test items represent the content domain, typically done through expert judgment before data collection.
Question 18
Two clinical psychologists independently score a projective test for a group of 30 patients to assess their level of aggression. An analysis reveals a high correlation (r = .85) between their scores. This finding provides strong evidence for the test's:
- inter-rater reliability. (correct answer)
- test-retest reliability.
- convergent validity.
- internal consistency.
Explanation: The correct answer is A. Inter-rater reliability (also called inter-observer reliability) is the degree of agreement between two or more independent raters or observers. When a test requires subjective scoring, as projective tests do, it is crucial to demonstrate that different scorers will arrive at similar conclusions. The high correlation between the two psychologists' scores indicates that the scoring system is consistent across different raters.
- B is incorrect because test-retest reliability would require administering the test to the same patients at two different times.
- C is incorrect because convergent validity would require correlating the test scores with another, different measure of aggression.
- D is incorrect because internal consistency refers to how well the items within the test correlate with each other, which is not what was assessed here.
Question 19
A researcher creates a 100-item test of historical knowledge. To evaluate the test, they calculate the correlation between the total scores on the odd-numbered items and the total scores on the even-numbered items. This procedure is performed to gather evidence about the test's:
- predictive validity, by using one half of the test to forecast scores on the other.
- content validity, by ensuring both halves of the test cover similar material.
- test-retest reliability, by treating the two halves as separate testing sessions.
- internal consistency reliability, by assessing the homogeneity of the test items. (correct answer)
Explanation: The correct answer is D. This procedure describes the split-half method (specifically, an odd-even split), which is a way to measure a test's internal consistency. Internal consistency reliability refers to the degree to which all the items on a test measure the same underlying construct. By correlating two halves of the test, the researcher is checking if the items are consistent with one another.
- A is incorrect because predictive validity involves correlating test scores with a future criterion, not with another part of the same test.
- B is incorrect because content validity is established through expert review of the items against a content blueprint, not through statistical correlation of item subsets.
- C is incorrect because test-retest reliability requires administering the entire test at two different points in time to the same group of people.
Question 20
A researcher wants to create the most reliable test possible from a pool of 100 items all measuring a single psychological construct. Which of the following strategies would most likely achieve this goal?
- Select the 20 items that have the highest average correlation with all other items in the pool. (correct answer)
- Select the 20 items that show the greatest variability in scores across the population.
- Select the 20 items that have the highest face validity according to a panel of students.
- Select a diverse set of 20 items that show low correlations with each other to cover the construct broadly.
Explanation: When you encounter questions about test reliability, remember that reliability refers to the consistency and stability of measurement. A reliable test produces similar results when administered repeatedly or when items measure the same underlying construct.
Answer A is correct because selecting items with the highest average correlation with all other items maximizes internal consistency. When test items correlate strongly with each other, they're all tapping into the same psychological construct effectively. This creates what psychometricians call high "inter-item correlation," which directly increases Cronbach's alpha—the most common measure of internal consistency reliability. Items that correlate well with the total pool are proven to measure the target construct reliably.
Answer B is wrong because high variability in scores doesn't guarantee reliability. While you need some variability to discriminate between people, items with extreme variability might be measuring random error or multiple constructs rather than your target construct consistently.
Answer C is incorrect because face validity (whether items appear to measure what they claim) doesn't ensure reliability. An item can look relevant to students but still produce inconsistent measurements. Face validity is about appearance, not psychometric quality.
Answer D represents a common misconception. While you want broad coverage of a construct, selecting items that show low correlations with each other actually decreases reliability. Low inter-item correlations suggest the items are measuring different things, which reduces internal consistency.
Remember this key principle: for reliability, you want items that "hang together" statistically. High inter-item correlations indicate your items are consistently measuring the same construct.