All questions
Question 1
A fire department uses a physical ability test for selecting new recruits. A lawsuit alleges the test is biased against female applicants, who pass at a lower rate than male applicants. According to established psychometric and legal principles, which finding would provide the strongest defense for the fire department?
- Evidence that the score difference between men and women on the test reflects national averages in physical strength.
- A demonstration that the test has high test-retest reliability for both male and female applicants.
- Proof that test scores are significantly and equally correlated with on-the-job performance metrics for both men and women. (correct answer)
- An expert review concluding that the test items appear to be directly related to the tasks of a firefighter.
Explanation: The strongest defense against a claim of test bias, particularly in a legal context (e.g., Uniform Guidelines on Employee Selection Procedures), is to demonstrate the test's validity for the job. Specifically, showing that the test is free of predictive bias—that it predicts job performance equally well for all groups—is the key. Option C directly addresses this by mentioning that the scores are equally correlated with job performance for both genders. Option A (reflecting national averages) doesn't prove the test is valid for the job. Option B (reliability) is necessary but not sufficient for validity. Option D addresses content validity, which is good, but predictive validity evidence (C) is typically considered stronger evidence against bias claims.
Question 2
A team of psychologists is tasked with creating a new intelligence test that is as culturally fair as possible. Which of the following strategies would be the LEAST effective method for reducing cultural bias in the test?
- Using non-verbal, abstract reasoning items like matrices and figure analogies.
- Having the test items reviewed by a diverse panel of experts from various cultural backgrounds.
- Statistically removing any item that shows a significant difference in average performance between cultural groups. (correct answer)
- Ensuring that the test is carefully translated and adapted for the cultural context of each group taking it.
Explanation: While it may seem intuitive, simply removing all items that show group differences is a poor strategy. These differences might reflect true, valid differences in the construct being measured. Over-zealously removing such items can damage the test's construct validity, resulting in a test that measures the construct less accurately for everyone. The other options are all established and recommended practices for developing culturally fair tests. Using non-verbal items (A), having diverse review panels (B), and careful translation/adaptation (D) are all methods aimed at minimizing the influence of irrelevant cultural knowledge on test scores.
Question 3
A large corporation uses a validated cognitive ability test to screen candidates for management positions. The test has been shown to be free of predictive bias; it accurately predicts job performance for both male and female employees. Nevertheless, data shows that 80% of the employees promoted to management are male. An external audit raises concerns about fairness.
Based on the information in the passage, what is the most accurate analysis of the situation?
- The test is inherently unfair because it results in a disparate impact on female candidates.
- The test itself is psychometrically unbiased, but its application or other factors in the promotion process may be unfair. (correct answer)
- The finding that the test is free of predictive bias must be incorrect, as evidenced by the promotion disparity.
- The corporation should replace the cognitive ability test with one that produces equal outcomes for men and women.
Explanation: This question highlights the crucial distinction between test bias and test fairness. Test bias is a technical, psychometric property. The passage states the test is free of predictive bias. Test fairness is a broader social and ethical concept about how test scores are used. The disparity in promotion rates, despite the unbiased test, suggests that other factors in the promotion process (e.g., interviewer bias, different opportunities) might be unfair. Therefore, the test can be technically sound (unbiased) while the overall system is unfair. Distractor A incorrectly equates disparate impact with unfairness of the test itself. Distractor C wrongly assumes that unequal outcomes must mean the test is biased. Distractor D proposes a solution focused on equal outcomes, which might compromise the validity of the selection process.
Question 4
A researcher administers a Western-developed self-report measure of 'assertiveness' to participants in both Japan and the United States. They find that in the U.S., scores correlate strongly with measures of independence and leadership. In Japan, scores correlate more strongly with measures of social insensitivity and disruptiveness. This finding suggests the 'assertiveness' test may have what kind of bias?
- Predictive bias
- Construct bias (correct answer)
- Item bias
- Intercept bias
Explanation: Construct bias occurs when a test measures different psychological constructs or traits across different groups or cultures. In this case, the network of correlations (its nomological net) is different in the two cultures. What is interpreted as positive 'assertiveness' in the U.S. is interpreted as a negative trait in Japan. This indicates that the test is tapping into different underlying constructs in the two cultures. Predictive bias (A) and intercept bias (D) relate to how well the test predicts an external criterion. Item bias (C) refers to specific items being problematic, whereas this scenario suggests a problem with the overall construct being measured by the test.
Question 5
The 'Cleary model' of test bias, which is based on regression analysis, defines a test as unbiased if the relationship between the test score and the criterion is the same for all subgroups. This model is primarily designed to detect which specific type of bias?
- Predictive bias (correct answer)
- Construct bias
- Content bias
- Cultural bias
Explanation: When you encounter questions about test bias models, focus on what specific aspect of testing validity each model is designed to measure. The Cleary model uses regression analysis to examine the relationship between test scores and criterion performance across different groups.
The Cleary model specifically targets predictive bias by analyzing whether a test predicts criterion performance equally well for all subgroups. If the regression lines (showing the relationship between test scores and actual performance) are the same across groups, the test is considered unbiased. When these relationships differ significantly between groups, it indicates the test may systematically over- or under-predict performance for certain subgroups, which is the essence of predictive bias.
Looking at the incorrect options: (B) Construct bias refers to whether a test measures the same psychological construct across different groups, which requires factor analysis rather than regression. (C) Content bias involves examining whether test items themselves are inappropriate or unfamiliar to certain groups, typically assessed through expert review panels. (D) Cultural bias is a broader concept encompassing various ways culture might influence test performance, but it's not what the Cleary model specifically measures.
The key distinction is that the Cleary model doesn't examine test content or underlying constructs—it focuses solely on whether the predictive relationship holds consistent across groups. Remember that different bias detection methods target different aspects of test fairness: content review for item appropriateness, factor analysis for construct equivalence, and regression analysis (like Cleary's model) for predictive validity across groups.
Question 6
A university admissions officer is concerned about predictive bias in their entrance exam. They find that the exam systematically overpredicts the first-year grades of students from wealthy backgrounds and underpredicts the grades of students from less-wealthy backgrounds. This is a classic example of which type of predictive bias?
- Intercept bias (correct answer)
- Selection bias
- Slope bias
- Response bias
Explanation: When you encounter questions about predictive bias in psychological testing, focus on how the test performs differently across groups. Predictive bias occurs when a test systematically over- or underpredicts outcomes for certain populations.
This scenario describes intercept bias, where the test's baseline prediction differs systematically between groups. The exam overpredicts wealthy students' grades (suggesting they'll perform better than they actually do) and underpredicts less-wealthy students' grades (suggesting they'll perform worse than they actually do). This creates parallel prediction lines with different starting points (intercepts) for each group, even though the relationship between test scores and grades might be equally strong for both groups.
Looking at the wrong answers: (B) Selection bias refers to problems in how participants are chosen for a study, not how a test predicts outcomes across groups. (C) Slope bias would occur if the test's predictive strength differed between groups—for example, if test scores were highly predictive for one group but weakly predictive for another. (D) Response bias involves systematic errors in how people answer questions, like social desirability bias or acquiescence bias.
Study tip: Remember that intercept bias is about systematic over/underprediction (the "starting point" differs), while slope bias is about different predictive relationships (the "strength" of prediction differs). Look for keywords like "overpredicts" and "underpredicts" when the same test systematically errs in opposite directions for different groups—that's your clue for intercept bias.
Question 7
A large university uses a standardized test for admissions. Analysis reveals that the mean score for applicants from suburban high schools is 15 points higher than for applicants from urban high schools. However, further investigation shows that for any given test score, students from both groups have an equal probability of achieving a first-year GPA of 3.0 or higher. Based on these findings, which conclusion is most appropriate?
- The test exhibits significant predictive bias against students from urban high schools.
- The test lacks evidence of predictive bias, despite the observed difference in mean scores. (correct answer)
- The difference in mean scores is definitive proof of construct bias in the test.
- The test should be considered unfair because it perpetuates existing educational inequalities.
Explanation: Predictive bias refers to a test's accuracy in predicting an outcome for different groups. The key finding is that the test predicts first-year GPA equally well for both urban and suburban students (i.e., a given score means the same thing in terms of future performance for both groups). Therefore, despite the difference in average scores, there is no evidence of predictive bias. Distractor A is incorrect because the evidence points to a lack of predictive bias. Distractor C is incorrect because mean differences alone do not prove construct bias; they could reflect true differences in preparation. Distractor D confuses the technical concept of bias with the broader, value-laden concept of fairness; while one might argue the test's use is unfair, the test itself is not psychometrically biased in its prediction.
Question 8
A researcher is evaluating a new scholastic aptitude test. They find that for Group 1, the correlation between test scores and college GPA is r = .50. For Group 2, the correlation is r = .20. Furthermore, the regression lines predicting GPA from test scores for the two groups have significantly different slopes. This pattern of results is a clear indication of what form of test bias?
- Content bias, because the test material is more relevant to Group 1.
- Intercept bias, because the starting point for prediction differs between the groups.
- Slope bias, because the test has different levels of predictive validity for the two groups. (correct answer)
- Construct bias, because the test measures different underlying traits in the two groups.
Explanation: Slope bias (also known as differential validity) occurs when the correlation between test scores and a criterion is significantly different for different groups. This is reflected in different slopes of the regression lines. The scenario explicitly states that the correlation coefficients (a measure of validity) and the slopes are different for Group 1 and Group 2, which is the definition of slope bias. Intercept bias (B) involves parallel regression lines with different starting points. While content bias (A) or construct bias (D) might be the underlying cause of the slope bias, the statistical evidence presented (different correlations and slopes) directly defines slope bias.
Question 9
A company's HR department wants to ensure their pre-employment test is free from predictive bias regarding gender. Which of the following procedures is the most direct and appropriate way to investigate this?
- Calculate the average test scores for male and female applicants to see if they are significantly different.
- Conduct a factor analysis of the test items separately for men and women to see if the structure is the same.
- Analyze the relationship between test scores and a measure of job performance separately for men and women. (correct answer)
- Have a panel of male and female employees review the test content for potentially offensive or stereotypical language.
Explanation: Predictive bias is specifically about whether a test's predictions of a criterion (like job performance) are systematically inaccurate for different groups. The only way to test this is to collect both test score data and criterion data (job performance) and then analyze the relationship (typically using regression) separately for each group to see if the predictive relationship is the same. Option A checks for mean differences, which is not the same as bias. Option B investigates construct bias, not predictive bias. Option D is a good practice for ensuring content fairness but does not address the statistical issue of predictive bias.
Question 10
A university is using an entrance exam. To increase diversity, it abandons its previous policy of using a single cutoff score for all applicants. Instead, it decides to admit the top 20% of scorers from each major ethnic group. From a psychometric perspective, this change in policy reflects a shift from a fairness model based on to one based on .
- regression; equal opportunity
- equal outcomes; meritocracy
- classical test theory; item response theory
- individual prediction; group parity (correct answer)
Explanation: The original policy, using a single cutoff score, is based on a model of fairness where every individual's score is treated the same in predicting success (individual prediction, also related to the regression model). The new policy, which ensures proportional representation from different groups by using different effective standards (top 20% within each group), is a model based on achieving group parity or equal outcomes. This question tests the understanding of different philosophical models of test fairness. The distractors use related but incorrect terminology for this specific contrast.
Question 11
A psychologist is evaluating an item on a new depression inventory. An analysis reveals that male and female participants who have the same total score on the inventory show a significantly different likelihood of endorsing this specific item. This statistical finding is known as:
- a Type I error.
- differential item functioning. (correct answer)
- low criterion-related validity.
- a lack of internal consistency.
Explanation: The scenario describes the exact definition of differential item functioning (DIF). DIF occurs when subgroups of test-takers (here, males and females) who are matched on the underlying trait being measured (overall depression score) have a different probability of answering a particular item in a certain way. It's a form of item bias. A Type I error (A) is a general statistical concept of a false positive. Low criterion-related validity (C) refers to the test as a whole not predicting an outcome well. A lack of internal consistency (D) means the items on the test do not correlate well with each other.
Question 12
A company finds that its hiring test accurately predicts which employees will be successful. It also finds that applicants from Group X score, on average, lower than applicants from Group Y. If the test is used with a single cutoff score, fewer applicants from Group X will be hired. What is the most precise psychometric implication of this situation?
- The test is invalid because it leads to disparate impact against Group X.
- The test is biased because the average scores of the two groups are different.
- The lower scores for Group X prove that the test measures a construct that is irrelevant to the job.
- The test may be valid and unbiased, but its use creates a conflict between maximizing productivity and achieving demographic representation. (correct answer)
Explanation: When you encounter questions about employment testing and group differences, you need to distinguish between three key psychometric concepts: validity (does the test predict job performance?), bias (does the test predict equally well for different groups?), and disparate impact (does the test result in different hiring rates across groups?).
The correct answer is D because this scenario describes a valid, unbiased test that creates disparate impact. The test accurately predicts job success (validity) for all applicants, and there's no indication it predicts differently for Groups X and Y (no bias). However, using a single cutoff score will result in fewer Group X hires due to their lower average scores, creating a tension between selecting the most qualified candidates and maintaining demographic diversity.
Option A incorrectly assumes disparate impact automatically makes a test invalid. Disparate impact is a legal concept about hiring outcomes, not a measure of test validity. Option B confuses group differences in scores with test bias. Bias occurs when a test predicts differently for different groups, not when groups have different average scores. Option C makes an unfounded logical leap. Group differences in scores don't prove the test measures job-irrelevant constructs—the differences could reflect real differences in the skills being measured.
Remember that validity, bias, and disparate impact are independent concepts in psychometrics. A test can be valid and unbiased while still producing disparate impact, which creates practical and legal challenges for employers who must balance merit-based selection with diversity goals.
Question 13
A physical fitness test is used to select candidates for a demanding job. Research establishes a single regression equation for all candidates: Predicted Performance = (0.8 * Test Score) + 10. A new study examines two groups. For Group A members who score 50 on the test, their average actual performance is 50. For Group B members who also score 50 on the test, their average actual performance is 58. What does this finding indicate?
- The test underpredicts performance for Group B and is therefore biased against them. (correct answer)
- The test lacks content validity because it does not adequately sample the job's requirements.
- The test has slope bias, as the value of performance gained per test point differs between groups.
- The test overpredicts performance for Group A and is therefore biased against them.
Explanation: Test bias in psychological assessment occurs when a test systematically over- or underpredicts performance for different groups, even when they have the same test scores. This is a critical fairness issue that can lead to discrimination in selection processes.
Let's analyze what's happening here. The regression equation predicts that candidates scoring 50 should perform at: (0.8 × 50) + 10 = 50. For Group A, this prediction is accurate—they actually perform at 50. However, Group B members also score 50 on the test but actually perform at 58, which is 8 points higher than predicted. This means the test systematically underestimates Group B's true capabilities.
Answer A correctly identifies this as underprediction bias against Group B. When a test underpredicts performance for a group, it creates unfair disadvantage because qualified members may be rejected based on artificially low predicted scores.
Answer B is incorrect because content validity refers to whether test items represent the job domain, not prediction accuracy differences between groups. Answer C misidentifies slope bias, which occurs when the relationship between test scores and performance differs in strength (slope) between groups—but we only have one data point per group here. Answer D incorrectly claims overprediction for Group A, but the test actually predicts Group A's performance accurately.
Remember: Test bias isn't about different groups having different average scores—it's about the test making systematically wrong predictions for certain groups. Always compare predicted versus actual performance to identify bias.
Question 14
A university discovers that its admissions test exhibits intercept bias, underpredicting the college success of applicants from a specific minority group. To remedy this, the university adds 50 points to the test score of every applicant from that group. Which of the following is the most likely psychometric consequence of this action?
- The adjustment will eliminate both intercept and slope bias for the group, making the test completely fair.
- This method will correct the underprediction issue but may be inaccurate if slope bias also exists. (correct answer)
- The adjustment will decrease the test's overall predictive validity by introducing systematic error.
- This action will correct the bias but is considered psychometrically equivalent to using separate cutoff scores.
Explanation: Adding a constant number of points to the scores of one group is a statistical method for correcting intercept bias. It effectively raises the regression line for that group without changing its slope. Therefore, it will correct the underprediction issue if the problem is only one of intercept bias (i.e., the regression lines should be parallel). However, if the slopes are also different (slope bias), this simple adjustment will not be accurate across the full range of scores and may even worsen prediction for some individuals. A is too strong a claim. C is possible but not the most precise consequence; the goal is to reduce error. D is incorrect; adding points is a different statistical approach than setting separate cutoffs, though both are strategies to address fairness concerns.
Question 15
A large corporation uses a validated cognitive ability test to screen candidates for management positions. The test has been shown to be free of predictive bias; it accurately predicts job performance for both male and female employees. Nevertheless, data shows that 80% of the employees promoted to management are male. An external audit raises concerns about fairness.
Based on the information in the passage, what is the most accurate analysis of the situation?
- The test is inherently unfair because it results in a disparate impact on female candidates.
- The test itself is psychometrically unbiased, but its application or other factors in the promotion process may be unfair. (correct answer)
- The finding that the test is free of predictive bias must be incorrect, as evidenced by the promotion disparity.
- The corporation should replace the cognitive ability test with one that produces equal outcomes for men and women.
Explanation: This question highlights the crucial distinction between test bias and test fairness. Test bias is a technical, psychometric property. The passage states the test is free of predictive bias. Test fairness is a broader social and ethical concept about how test scores are used. The disparity in promotion rates, despite the unbiased test, suggests that other factors in the promotion process (e.g., interviewer bias, different opportunities) might be unfair. Therefore, the test can be technically sound (unbiased) while the overall system is unfair. Distractor A incorrectly equates disparate impact with unfairness of the test itself. Distractor C wrongly assumes that unequal outcomes must mean the test is biased. Distractor D proposes a solution focused on equal outcomes, which might compromise the validity of the selection process.
Question 16
A researcher is evaluating a new scholastic aptitude test. They find that for Group 1, the correlation between test scores and college GPA is r = .50. For Group 2, the correlation is r = .20. Furthermore, the regression lines predicting GPA from test scores for the two groups have significantly different slopes. This pattern of results is a clear indication of what form of test bias?
- Content bias, because the test material is more relevant to Group 1.
- Intercept bias, because the starting point for prediction differs between the groups.
- Slope bias, because the test has different levels of predictive validity for the two groups. (correct answer)
- Construct bias, because the test measures different underlying traits in the two groups.
Explanation: Slope bias (also known as differential validity) occurs when the correlation between test scores and a criterion is significantly different for different groups. This is reflected in different slopes of the regression lines. The scenario explicitly states that the correlation coefficients (a measure of validity) and the slopes are different for Group 1 and Group 2, which is the definition of slope bias. Intercept bias (B) involves parallel regression lines with different starting points. While content bias (A) or construct bias (D) might be the underlying cause of the slope bias, the statistical evidence presented (different correlations and slopes) directly defines slope bias.
Question 17
A researcher administers a Western-developed self-report measure of 'assertiveness' to participants in both Japan and the United States. They find that in the U.S., scores correlate strongly with measures of independence and leadership. In Japan, scores correlate more strongly with measures of social insensitivity and disruptiveness. This finding suggests the 'assertiveness' test may have what kind of bias?
- Predictive bias
- Construct bias (correct answer)
- Item bias
- Intercept bias
Explanation: Construct bias occurs when a test measures different psychological constructs or traits across different groups or cultures. In this case, the network of correlations (its nomological net) is different in the two cultures. What is interpreted as positive 'assertiveness' in the U.S. is interpreted as a negative trait in Japan. This indicates that the test is tapping into different underlying constructs in the two cultures. Predictive bias (A) and intercept bias (D) relate to how well the test predicts an external criterion. Item bias (C) refers to specific items being problematic, whereas this scenario suggests a problem with the overall construct being measured by the test.
Question 18
A university is using an entrance exam. To increase diversity, it abandons its previous policy of using a single cutoff score for all applicants. Instead, it decides to admit the top 20% of scorers from each major ethnic group. From a psychometric perspective, this change in policy reflects a shift from a fairness model based on to one based on .
- regression; equal opportunity
- equal outcomes; meritocracy
- classical test theory; item response theory
- individual prediction; group parity (correct answer)
Explanation: The original policy, using a single cutoff score, is based on a model of fairness where every individual's score is treated the same in predicting success (individual prediction, also related to the regression model). The new policy, which ensures proportional representation from different groups by using different effective standards (top 20% within each group), is a model based on achieving group parity or equal outcomes. This question tests the understanding of different philosophical models of test fairness. The distractors use related but incorrect terminology for this specific contrast.
Question 19
A fire department uses a physical ability test for selecting new recruits. A lawsuit alleges the test is biased against female applicants, who pass at a lower rate than male applicants. According to established psychometric and legal principles, which finding would provide the strongest defense for the fire department?
- Evidence that the score difference between men and women on the test reflects national averages in physical strength.
- A demonstration that the test has high test-retest reliability for both male and female applicants.
- Proof that test scores are significantly and equally correlated with on-the-job performance metrics for both men and women. (correct answer)
- An expert review concluding that the test items appear to be directly related to the tasks of a firefighter.
Explanation: The strongest defense against a claim of test bias, particularly in a legal context (e.g., Uniform Guidelines on Employee Selection Procedures), is to demonstrate the test's validity for the job. Specifically, showing that the test is free of predictive bias—that it predicts job performance equally well for all groups—is the key. Option C directly addresses this by mentioning that the scores are equally correlated with job performance for both genders. Option A (reflecting national averages) doesn't prove the test is valid for the job. Option B (reliability) is necessary but not sufficient for validity. Option D addresses content validity, which is good, but predictive validity evidence (C) is typically considered stronger evidence against bias claims.
Question 20
A company finds that its hiring test accurately predicts which employees will be successful. It also finds that applicants from Group X score, on average, lower than applicants from Group Y. If the test is used with a single cutoff score, fewer applicants from Group X will be hired. What is the most precise psychometric implication of this situation?
- The test is invalid because it leads to disparate impact against Group X.
- The test is biased because the average scores of the two groups are different.
- The lower scores for Group X prove that the test measures a construct that is irrelevant to the job.
- The test may be valid and unbiased, but its use creates a conflict between maximizing productivity and achieving demographic representation. (correct answer)
Explanation: When you encounter questions about employment testing and group differences, you need to distinguish between three key psychometric concepts: validity (does the test predict job performance?), bias (does the test predict equally well for different groups?), and disparate impact (does the test result in different hiring rates across groups?).
The correct answer is D because this scenario describes a valid, unbiased test that creates disparate impact. The test accurately predicts job success (validity) for all applicants, and there's no indication it predicts differently for Groups X and Y (no bias). However, using a single cutoff score will result in fewer Group X hires due to their lower average scores, creating a tension between selecting the most qualified candidates and maintaining demographic diversity.
Option A incorrectly assumes disparate impact automatically makes a test invalid. Disparate impact is a legal concept about hiring outcomes, not a measure of test validity. Option B confuses group differences in scores with test bias. Bias occurs when a test predicts differently for different groups, not when groups have different average scores. Option C makes an unfounded logical leap. Group differences in scores don't prove the test measures job-irrelevant constructs—the differences could reflect real differences in the skills being measured.
Remember that validity, bias, and disparate impact are independent concepts in psychometrics. A test can be valid and unbiased while still producing disparate impact, which creates practical and legal challenges for employers who must balance merit-based selection with diversity goals.