All questions
Question 1
A psychologist is designing a study that will use a behavioral observation system to code the frequency of aggressive behaviors among children on a playground. The coding system requires the observer to make judgments about whether an action (e.g., a push) was playful or aggressive. Which type of reliability is most essential to establish for this coding system?
- Split-half reliability.
- Inter-rater reliability. (correct answer)
- Test-retest reliability.
- Alternate-forms reliability.
Explanation: Inter-rater reliability refers to the degree of agreement between two or more independent raters or observers. When an assessment involves subjective judgment, as in this behavioral coding system, it is crucial to demonstrate that different observers will arrive at the same conclusions when observing the same behavior. The other forms of reliability are less relevant for an observational system of this type.
Question 2
A psychologist is tasked with selecting a pre-employment screening tool for air traffic controllers. The primary goal is to identify candidates most likely to succeed in the high-stress work environment. The selection committee wants the most effective instrument for predicting future on-the-job performance. Which psychometric property is the most critical for the psychologist to prioritize in the selected instrument?
- High test-retest reliability over a six-month interval to ensure trait stability.
- High internal consistency (Cronbach's alpha) to ensure the items measure a single, coherent construct.
- Strong predictive validity, demonstrated by a significant correlation with future performance ratings. (correct answer)
- Strong concurrent validity, established by correlating the test with the performance of current employees.
Explanation: Predictive validity is the extent to which a test score predicts a future outcome, which directly aligns with the stated goal of identifying candidates likely to succeed. While other properties like reliability (A, B) are necessary for a good test, they do not directly address its predictive power. Concurrent validity (D) assesses the relationship with current performance, which is less relevant than predicting the performance of new hires.
Question 3
During a re-evaluation for special education services, a school psychologist considers using the same version of an IQ test administered to a student 12 years prior. The test manual is now two editions out of date. If the psychologist proceeds with using the older version, what is the most likely psychometric consequence due to the Flynn effect?
- The student's obtained IQ score will likely be artificially inflated relative to current population norms. (correct answer)
- The student's obtained IQ score will likely be artificially deflated due to outdated test content.
- The test's reliability will be compromised, leading to a larger standard error of measurement.
- The resulting score will be invalid because of the significant practice effects from the prior administration.
Explanation: The Flynn effect refers to the observed trend of rising scores on intelligence tests over time. Using outdated norms means comparing the student to a lower-performing normative sample from the past. This results in an artificially inflated score, which could mask an intellectual disability or lead to an inaccurate picture of the student's abilities compared to their current peers.
Question 4
A 9-year-old child is referred for a comprehensive psychological evaluation. The referral from the pediatrician notes, 'The patient presents with academic difficulties, social withdrawal, and frequent irritability at home. Please clarify diagnosis and provide treatment recommendations.' The psychologist has access to a wide range of assessment tools.
Given the diffuse nature of the referral concerns, which instrument selection strategy is most appropriate for the initial phase of this assessment?
- Administer a narrowband measure of childhood depression to directly assess a potential mood disorder.
- Administer a broadband behavior rating scale completed by parents and teachers to screen for a wide range of potential problems. (correct answer)
- Administer a specific test of reading comprehension to investigate the primary academic complaint.
- Administer a projective test, such as a sentence completion task, to explore underlying emotional conflicts.
Explanation: When the referral question is broad and symptoms are diffuse, the most effective initial strategy is to use a broadband instrument. This allows the psychologist to screen for a wide array of potential issues (e.g., internalizing, externalizing, attention problems) and then narrow the focus with more specific, narrowband measures as indicated. The other options are too narrow or less empirically supported for an initial, broad-based assessment.
Question 5
An 8-year-old child receives an age-equivalent score of 5 years, 0 months on a vocabulary test. A psychologist reports to the parents that the child's 'vocabulary is like that of a typical 5-year-old.' What is the most significant psychometric flaw in this interpretation?
- The interpretation is likely to cause undue distress to the parents and should be phrased more positively.
- Age-equivalent scores are ordinal and do not account for the increasing variability of performance with age. (correct answer)
- The normative sample for 5-year-olds on this test may be too small to provide a stable comparison point.
- The test's test-retest reliability is likely to be lower for younger children in the normative sample.
Explanation: The primary psychometric problem with age-equivalent scores is that they are ordinal and do not represent an equal-interval scale. A three-year gap has a very different meaning for an 8-year-old than it does for a 5-year-old. Because the standard deviation of scores typically increases with age, a raw score difference that seems large may be statistically less significant for an older child. Standard scores are preferred as they provide information about a child's standing relative to their same-age peers.
Question 6
A psychologist is evaluating a 12-year-old student for a gifted program. The student consistently scores at the maximum possible score on all school-based achievement tests. The goal of the evaluation is to differentiate the student's abilities at the very high end of the spectrum. Which psychometric feature is most crucial when selecting an intelligence test for this student?
- An adequate test ceiling to allow for variability among high-ability individuals. (correct answer)
- An adequate test floor to ensure the instrument can measure lower levels of functioning.
- High convergent validity with the school's achievement tests to confirm existing data.
- High inter-rater reliability to ensure scoring consistency across different examiners.
Explanation: A test ceiling is the highest score that can be obtained on a test. For a gifted evaluation, it is critical to select an instrument with a high enough ceiling to capture the full range of the student's abilities and differentiate them from other high-scoring individuals. A low ceiling would result in the student 'topping out' and their true ability level being underestimated. A low floor (B) is important for assessing disability, not giftedness. Convergent validity with school tests (C) is not helpful when those tests are already being maxed out. Inter-rater reliability (D) is less of a concern for objectively scored cognitive tests.
Question 7
A psychologist is conducting a neuropsychological assessment of a 45-year-old male who sustained a moderate traumatic brain injury (TBI) one year ago. The selected memory test offers norms for the general adult population as well as norms for a TBI-specific sample matched on injury severity. What is the most clinically sophisticated way to use these two sets of norms?
- Use only the general population norms to determine the absolute level of impairment for legal and disability purposes.
- Use only the TBI-specific norms to determine if the client's memory is worse than expected for someone with his injury.
- Compare the client's scores to both sets of norms to create a comprehensive picture of his functioning. (correct answer)
- Average the percentile ranks from both norm groups to obtain a single, more stable estimate of his ability.
Explanation: Using both norm sets provides the most complete clinical picture. Comparing the client to the general population norms reveals the degree of overall impairment relative to healthy peers. Comparing him to the TBI-specific norms helps determine if his performance is typical or atypical for someone with a similar injury, which can inform prognosis and treatment planning. Relying on only one set of norms provides an incomplete view.
Question 8
A student's score on a test of quantitative reasoning improved from the 55th percentile to the 65th percentile over the course of a school year. During the same period, their score on a verbal reasoning test improved from the 89th percentile to the 99th percentile. Based on the properties of the percentile rank scale, what can the psychologist conclude?
- The improvements were equal, as both represent a gain of 10 percentile points.
- The improvement in quantitative reasoning was more meaningful as it crossed the average range.
- The improvement in verbal reasoning represents a larger increase in underlying ability. (correct answer)
- No conclusion can be drawn without converting the scores to T-scores or z-scores.
Explanation: Percentile ranks are an ordinal scale and are not spaced equally. The distribution of scores is densest in the middle and more spread out at the extremes. Therefore, a 10-point jump near the mean (55 to 65) represents a smaller change in the underlying raw score or standard score than a 10-point jump at the upper extreme (89 to 99). The verbal reasoning improvement is more substantial.
Question 9
A psychologist is retained for a highly contentious child custody evaluation. The psychologist plans to assess parental personality and psychopathology. In selecting an instrument for this purpose, which psychometric feature is of paramount importance given the forensic context?
- Brevity of the instrument, to minimize the burden on the parents during a stressful time.
- Availability of robust validity scales designed to detect impression management and random responding. (correct answer)
- Norms that are specific to parents involved in custody disputes to provide a relevant comparison.
- High face validity, so the parents perceive the questions as relevant and are more cooperative.
Explanation: In forensic contexts like custody evaluations, there is a high potential for dissimulation (i.e., faking good or faking bad). Therefore, the most critical feature of a personality instrument is the presence of well-validated scales to detect invalid response styles, such as defensiveness, exaggeration, or inconsistency. While other factors are considerations, ensuring the validity of the obtained data is the primary psychometric concern.
Question 10
A psychologist uses a new screening test to predict which employees in a company are at high risk for burnout. To create a confidence interval around an employee's predicted burnout score (the criterion), which of the following psychometric indices must be used?
- The standard error of measurement (SEM) of the screening test.
- The standard deviation of the screening test's normative sample.
- The internal consistency reliability of the screening test.
- The standard error of the estimate (SEE). (correct answer)
Explanation: This is a key distinction. The Standard Error of Measurement (SEM) is used to create a confidence interval around an obtained score on a single test. The Standard Error of the Estimate (SEE) is used in regression and prediction; it quantifies the typical error in predicting a criterion score (burnout) from a predictor score (screening test). To create a confidence interval around the predicted score, the SEE is the correct statistic.
Question 11
When evaluating a new test of executive functioning, a psychologist notes that the test was standardized on a sample of college students from a single university. The psychologist is considering using the test with community-dwelling adults. Which aspect of the test's psychometric properties is most directly compromised when used with this new population?
- Internal consistency.
- Content validity.
- Normative appropriateness. (correct answer)
- Inter-rater reliability.
Explanation: The norms of a test are based on the performance of the standardization sample. Using a test on a population that is significantly different from the normative sample (in this case, community-dwelling adults vs. college students) invalidates the norms. It becomes impossible to accurately interpret how an individual's score compares to a relevant peer group. The other properties are less likely to be directly affected by the change in population.
Question 12
A research team develops a new self-report scale for 'catastrophic thinking.' They administer it along with a well-established measure of depression and a measure of verbal intelligence. They find a strong positive correlation (r = .65) with the depression scale and a near-zero correlation (r = .05) with the verbal intelligence scale. These findings provide initial support for the new scale's:
- test-retest reliability and internal consistency.
- concurrent validity and predictive validity.
- convergent validity and discriminant validity. (correct answer)
- content validity and face validity.
Explanation: This scenario describes the two components of construct validity. Convergent validity is shown by the strong correlation with a related construct (depression). Discriminant (or divergent) validity is shown by the weak correlation with an unrelated construct (verbal intelligence). The other options refer to different types of reliability and validity that are not directly assessed by this specific correlational design.
Question 13
A psychologist is assessing the cognitive functioning of a 15-year-old girl who immigrated from rural Mexico one year ago and is a sequential bilingual. The psychologist must select an appropriate intelligence test. Which of the following options represents the best choice?
- A leading English-language intelligence test administered with a certified interpreter to translate instructions.
- A popular nonverbal intelligence test normed on a large, representative sample of the U.S. population.
- A Spanish-language version of an intelligence test that was normed exclusively on a sample of adolescents in Spain.
- An intelligence test that was developed and normed on a sample of bilingual, Spanish-English speaking adolescents in the U.S. (correct answer)
Explanation: The most appropriate instrument is one normed on a population that is as similar as possible to the client on relevant characteristics. This includes language, culture, and educational background. Using an interpreter (A) invalidates the norms. A nonverbal test (B) reduces linguistic demand but does not eliminate cultural factors. Norms from Spain (C) are not appropriate for a client from Mexico. Option D provides the most appropriate normative comparison group.
Question 14
A test manual indicates its normative sample was stratified to match U.S. Census data on sex, age, parent education level, and geographic region. What is the primary psychometric benefit of this stratification procedure?
- It ensures the test items are free from cultural or demographic bias.
- It allows for the development of separate scoring norms for each demographic subgroup.
- It increases the internal consistency and test-retest reliability of the instrument.
- It enhances the representativeness of the sample, which supports the generalizability of the test scores. (correct answer)
Explanation: Stratification is a sampling technique used to ensure that specific subgroups of a population are represented in the sample in proportion to their presence in the overall population. The primary goal is to create a normative sample that is a miniature, representative version of the population, which in turn allows for greater confidence in generalizing test results from the sample to the broader population.
Question 15
A psychologist is assessing a 75-year-old man with two years of formal education for a suspected neurocognitive disorder. The psychologist is concerned that the client's low educational background may confound the interpretation of his performance on cognitive tests. In selecting an appropriate instrument, which consideration is most critical?
- Choosing a test that can be administered by a technician to reduce assessment time and cost.
- Selecting the briefest available screening tool to minimize client fatigue and frustration.
- Using a test that provides norms stratified by education level or uses regression-based adjustments. (correct answer)
- Prioritizing a test with high test-retest reliability to ensure score stability across administrations.
Explanation: Many cognitive abilities are highly correlated with educational attainment. Using general population norms for an individual with very low education can lead to an underestimation of their premorbid abilities and an over-interpretation of deficits. The best practice is to select an instrument that provides specific normative data for different educational levels, or a method to statistically adjust scores for education, to get a more accurate picture of potential decline.
Question 16
A psychologist in a university counseling center wants to select a self-report measure to administer at each session to track a client's weekly progress in therapy for social anxiety. Which psychometric property is uniquely important for a measure used for this purpose?
- High test-retest reliability over a one-year period.
- Sensitivity to change. (correct answer)
- A large, nationally representative normative sample.
- Strong discriminant validity from measures of academic stress.
Explanation: When a measure is used to track change over short intervals, it must be sensitive to change; that is, its scores must be able to reflect true improvements or deteriorations in the client's state. High long-term test-retest reliability (A) would be undesirable, as it would imply that scores are not expected to change. While norms (C) and discriminant validity (D) are important for diagnostic measures, sensitivity to change is paramount for a progress monitoring tool.
Question 17
The manual for a new, 10-item anxiety screening tool reports a very high internal consistency coefficient (Cronbach's alpha = .97). However, the manual provides no data showing how scores on the tool relate to diagnostic interviews or other established anxiety measures. Based on this information, what is the most accurate conclusion a psychologist can draw?
- The high reliability of the tool ensures that it is also a valid measure of anxiety.
- The tool's items are highly interrelated, but its validity as a measure of anxiety is unknown. (correct answer)
- The high alpha indicates the items are redundant, so the tool lacks construct validity.
- The tool is highly reliable and therefore suitable for tracking changes in anxiety during therapy.
Explanation: Reliability is a necessary but not sufficient condition for validity. A high Cronbach's alpha indicates that the items on the scale are measuring the same underlying construct consistently (internal consistency). However, it does not tell us what that construct is. Without evidence correlating the scores to an external criterion (e.g., a clinical diagnosis), its validity remains unknown. It is a reliable measure, but not necessarily a valid one.
Question 18
A psychologist is choosing between two measures of trait anxiety. Test A has a reliability coefficient of .91 and a standard error of measurement (SEM) of 2.5. Test B has a reliability coefficient of .84 and an SEM of 4.0. A client obtains a T-score of 60 on both tests. Which conclusion is most sound?
- Test A allows for more precision in estimating the client's true score. (correct answer)
- Test B is preferable because its reliability is still within the acceptable range for clinical use.
- The interpretation is identical for both tests because the client's obtained score is the same.
- The test with the larger and more recent normative sample should be preferred, regardless of the SEM.
Explanation: The standard error of measurement (SEM) is used to create a confidence interval around an obtained score. A smaller SEM indicates less error and a narrower confidence interval, allowing for a more precise estimate of the client's true score. Test A has a smaller SEM (2.5 vs. 4.0), making it the more precise instrument. The other options reflect common misconceptions or ignore the direct relevance of the SEM.
Question 19
A validation study for a new college entrance exam is conducted using only students admitted to a highly selective university. The study finds a modest predictive validity coefficient of r = .25 between exam scores and first-year GPA. The test developers had expected a stronger correlation. What is the most likely statistical artifact responsible for this low coefficient?
- A small sample size, which limits statistical power to detect the true effect.
- Low reliability of the criterion variable (first-year GPA), which attenuates the validity coefficient.
- Restriction of range in both the exam scores and GPA within the high-achieving sample. (correct answer)
- Poor content validity of the exam, as it did not match the university's curriculum.
Explanation: Restriction of range occurs when the variability of scores in the sample is smaller than the variability in the full population. In a highly selective university, both the entrance exam scores (predictor) and the GPAs (criterion) will be clustered at the high end. This reduced variability artificially lowers the correlation coefficient, underestimating the test's true predictive validity in a more heterogeneous population.
Question 20
A school psychologist is selecting a new universal behavior screener. The manual for Test A reports that its scores correlate significantly (r=.45, p<.01) with teacher ratings of disruptive behavior. The manual for Test B provides classification accuracy data, reporting a sensitivity of .85 and a specificity of .90 for identifying students who receive a formal diagnosis. For the purpose of universal screening, why is Test B's information more useful?
- The correlation coefficient from Test A is a more precise and continuous measure of the relationship.
- Test B's categorical data is less reliable than the correlational data provided for Test A.
- Test A's p-value indicates a very low probability of error, making it more scientifically rigorous.
- Test B's data directly addresses the test's ability to correctly classify students as at-risk or not at-risk. (correct answer)
Explanation: For a screening instrument, the primary goal is classification: correctly identifying individuals who have a condition (sensitivity) and correctly identifying those who do not (specificity). While a correlation coefficient provides information about the strength of a relationship, sensitivity and specificity provide direct, practical information about the test's diagnostic or screening utility. A significant p-value does not necessarily mean a test is clinically useful.