All questions
Question 1
A researcher is assessing the split-half reliability of a new 40-item survey. She correlates the sum of scores on the 20 odd-numbered items with the sum of scores on the 20 even-numbered items, yielding a correlation coefficient of r = .70. What is the most appropriate next step and conclusion?
- The reliability is .70, which is acceptable, so the scale can be used as is.
- The reliability is too low, indicating that the test has poor internal consistency.
- The .70 correlation must be adjusted using the Spearman-Brown formula to estimate the reliability of the full-length test. (correct answer)
- This procedure measures test-retest reliability, which should be supplemented with a measure of internal consistency like Cronbach's alpha.
Explanation: The correlation between two halves of a test (.70 in this case) underestimates the reliability of the full test, because reliability is partly a function of test length. The Spearman-Brown prophecy formula is used to correct this correlation and estimate what the reliability would be for the full 40-item test. The resulting corrected value will be higher than .70, likely indicating good internal consistency.
Question 2
A high school physics teacher creates a cumulative final exam. The course curriculum was divided equally among four topics: mechanics, thermodynamics, electromagnetism, and optics. However, over 80% of the exam questions are about mechanics. A review by the department head would most likely identify a primary weakness in the exam's:
- predictive validity.
- inter-rater reliability.
- construct validity.
- content validity. (correct answer)
Explanation: Content validity refers to the extent to which a measure or test represents all facets of a given construct or content domain. The final exam is supposed to cover the entire curriculum. By focusing almost exclusively on one of the four topics, the exam fails to adequately sample the content domain, thus giving it poor content validity.
Question 3
A researcher conducts a six-month study to assess the impact of a mindfulness meditation program on employee stress. Three months into the study, the company undergoes a major restructuring, leading to widespread job uncertainty and layoffs. The researcher notes that stress levels, as measured by a self-report questionnaire, are significantly higher at the end of the study than at the beginning. Attributing this increase in stress solely to the failure of the program would ignore which threat to internal validity?
- Testing effect
- Maturation
- History effect (correct answer)
- Instrumentation
Explanation: A history effect is a threat to internal validity that occurs when an external event, unrelated to the treatment, happens during the course of a study and affects the dependent variable. In this case, the company restructuring is a major external event that could have caused the increase in stress, making it impossible to isolate the effect (or lack thereof) of the mindfulness program.
Question 4
A social psychologist publishes a study concluding that people are less likely to help a stranger in need when other bystanders are present. The study was conducted on a public plaza using only American college students as participants. Which of the following critiques most directly addresses the study's external validity?
- The participants may have guessed the study's purpose and altered their behavior accordingly.
- The finding may not generalize to different cultures or to non-student populations. (correct answer)
- The researcher's observations of helping behavior may have been biased by their expectations.
- There may have been a confounding variable, such as the time of day, that influenced helping behavior.
Explanation: External validity concerns the extent to which the results of a study can be generalized to other populations, settings, and times. The critique that the findings may not apply beyond American college students directly questions the generalizability of the results to a broader population, which is a core component of external validity. The other options are threats to internal validity (demand characteristics, observer bias, confounding variables).
Question 5
Two clinical psychologists independently watch a video of a client's therapy session and rate the client's 'expression of affect' on a 1-to-7 scale for 20 distinct one-minute intervals. A comparison of their ratings yields a very low correlation (r = .23). Which of the following is the most direct threat to the validity of the study's conclusions about the client's affect?
- The measure has low inter-rater reliability, which limits its potential validity. (correct answer)
- The measure has low content validity because it only assesses one aspect of the session.
- The rating procedure is subject to a history effect, compromising internal validity.
- The client's behavior may be influenced by demand characteristics.
Explanation: Inter-rater reliability is the degree of agreement between independent observers. A low correlation indicates poor agreement, meaning the ratings are inconsistent and likely contain a large amount of measurement error. A fundamental principle of psychometrics is that a measure cannot be valid if it is not reliable. If the raters cannot agree on what they are seeing, the ratings cannot be an accurate (valid) measure of the client's affect.
Question 6
A cognitive psychologist studies the impact of a new 'brain-training' game on fluid intelligence. An experimental group plays the game for 30 minutes daily for four weeks, while a control group does not. Both groups are given an IQ test before and after the four-week period. The experimental group shows a significantly greater increase in IQ scores. A critic argues that the improvement may be due to the experimental group becoming more comfortable with the testing situation itself. This criticism refers to which potential confound?
- Placebo effect
- Testing effect (correct answer)
- Maturation effect
- Hawthorne effect
Explanation: A testing effect occurs when the act of taking a pre-test influences the scores on a post-test. In this case, the experimental group's improvement might not be due to the brain-training game but to practice or familiarity gained from the first IQ test. While this effect would typically apply to both groups, the criticism implies an interaction where the experimental group benefits more, but the core concept being invoked is the testing effect. It's a plausible alternative explanation for the observed score increase.
Question 7
A school district implements a new anti-bullying program in its middle schools. To identify the schools most in need, they select the 10 schools with the highest reported incidents of bullying in the previous year to receive the program. At the end of the school year, these 10 schools show a marked decrease in bullying incidents. The superintendent claims the program was a success. A methodologist might temper this conclusion by citing which threat to internal validity?
- Instrumentation effect
- Observer-expectancy effect
- Maturation
- Regression to the mean (correct answer)
Explanation: Regression to the mean is a statistical phenomenon where an extreme score on a first measurement tends to be closer to the average on a second measurement. By selecting the schools with the highest (most extreme) rates of bullying, it is statistically likely that their rates would have decreased somewhat in the following year even without any intervention. This makes it difficult to attribute the entire change to the program itself.
Question 8
To establish the construct validity of a new self-report measure of 'emotional intelligence,' a researcher correlates scores on the new measure with scores from three existing scales: a well-validated emotional intelligence test, a measure of social desirability, and a test of general cognitive ability (IQ). Which pattern of correlations would provide the strongest evidence for the new measure's construct validity?
- Strong positive correlation with the validated test, near-zero correlation with social desirability, and moderate positive correlation with IQ. (correct answer)
- Strong positive correlations with all three existing scales, demonstrating that it measures a broad range of related traits.
- Near-zero correlations with all three existing scales, demonstrating that it measures a completely novel construct.
- A strong positive correlation with the validated test and strong negative correlations with social desirability and IQ.
Explanation: Establishing construct validity involves demonstrating both convergent and discriminant validity. Convergent validity is shown by a strong positive correlation with an existing, validated measure of the same construct. Discriminant (or divergent) validity is shown by a weak or near-zero correlation with measures of unrelated or distinct constructs. Social desirability should be unrelated. While some theories link EI and IQ, a moderate correlation is more plausible than a very strong one, showing they are related but distinct. Option A shows the best combination of these.
Question 9
To validate a new leadership aptitude test for corporate managers, a company administers the test to its entire current management team. The researchers then gather the most recent annual performance evaluation scores for each manager and find a strong positive correlation between the test scores and the performance ratings. This method of validation primarily gathers evidence for which specific type of validity?
- Predictive validity
- Concurrent validity (correct answer)
- Divergent validity
- Content validity
Explanation: This is an example of concurrent validity, a type of criterion-related validity. It is called 'concurrent' because the test is administered at roughly the same time as the criterion (performance evaluation) data is collected. If the company had administered the test to new hires and then waited a year to collect performance data to correlate with the initial test scores, it would be an assessment of predictive validity.
Question 10
A psychologist develops a 30-item questionnaire to measure trait conscientiousness. An analysis reveals that the items on the first half of the test are not strongly correlated with the items on the second half. Further analysis shows that several individual items have weak or negative correlations with the total score of the other 29 items. These findings present the most direct challenge to the questionnaire's:
- test-retest reliability.
- internal consistency. (correct answer)
- external validity.
- criterion-related validity.
Explanation: Internal consistency is the extent to which all items on a scale measure the same underlying construct. The finding that different parts of the test (the two halves) are not well-correlated, and that individual items do not correlate with the total score, indicates that the items are not 'hanging together' to measure a single, unified concept. This is a direct measure of poor internal consistency.
Question 11
To avoid socially desirable responding on a new test for hiring police officers, a psychologist includes many subtle, indirect items such as preferences for certain geometric shapes and colors. The test proves to be a very strong predictor of which recruits will receive commendations for bravery two years after being hired. This test most likely has:
- high face validity and high predictive validity.
- low content validity and low predictive validity.
- high internal consistency and high face validity.
- low face validity and high predictive validity. (correct answer)
Explanation: Face validity refers to whether a test appears, on the surface, to measure what it is supposed to measure. A test about shapes and colors does not look like a test for police work, so it has low face validity. Predictive validity is a form of criterion validity that refers to how well a test predicts future outcomes. Since the test strongly predicts future commendations, it has high predictive validity. This is a common trade-off in psychological testing.
Question 12
A high school physics teacher creates a cumulative final exam. The course curriculum was divided equally among four topics: mechanics, thermodynamics, electromagnetism, and optics. However, over 80% of the exam questions are about mechanics. A review by the department head would most likely identify a primary weakness in the exam's:
- predictive validity.
- inter-rater reliability.
- construct validity.
- content validity. (correct answer)
Explanation: Content validity refers to the extent to which a measure or test represents all facets of a given construct or content domain. The final exam is supposed to cover the entire curriculum. By focusing almost exclusively on one of the four topics, the exam fails to adequately sample the content domain, thus giving it poor content validity.
Question 13
A psychologist develops a 30-item questionnaire to measure trait conscientiousness. An analysis reveals that the items on the first half of the test are not strongly correlated with the items on the second half. Further analysis shows that several individual items have weak or negative correlations with the total score of the other 29 items. These findings present the most direct challenge to the questionnaire's:
- test-retest reliability.
- internal consistency. (correct answer)
- external validity.
- criterion-related validity.
Explanation: Internal consistency is the extent to which all items on a scale measure the same underlying construct. The finding that different parts of the test (the two halves) are not well-correlated, and that individual items do not correlate with the total score, indicates that the items are not 'hanging together' to measure a single, unified concept. This is a direct measure of poor internal consistency.
Question 14
A research team develops a new scale to measure 'scholarly grit.' The scale is administered to a group of graduate students twice, one month apart, and the scores show a high correlation (r = .89). However, scores on the scale do not correlate with existing measures of perseverance, conscientiousness, or long-term goal attainment. Which statement best describes the psychometric properties of this new scale?
- High content validity, but low test-retest reliability.
- High test-retest reliability, but low construct validity. (correct answer)
- Low internal consistency, but high criterion validity.
- Low face validity, but high internal validity.
Explanation: The high correlation between scores from two different administrations indicates high test-retest reliability (consistency over time). However, the scale's failure to correlate with theoretically related constructs (perseverance, conscientiousness) and relevant outcomes (goal attainment) suggests it lacks construct validity; it does not appear to be measuring the intended concept of 'scholarly grit.'
Question 15
A school district implements a new anti-bullying program in its middle schools. To identify the schools most in need, they select the 10 schools with the highest reported incidents of bullying in the previous year to receive the program. At the end of the school year, these 10 schools show a marked decrease in bullying incidents. The superintendent claims the program was a success. A methodologist might temper this conclusion by citing which threat to internal validity?
- Instrumentation effect
- Observer-expectancy effect
- Maturation
- Regression to the mean (correct answer)
Explanation: Regression to the mean is a statistical phenomenon where an extreme score on a first measurement tends to be closer to the average on a second measurement. By selecting the schools with the highest (most extreme) rates of bullying, it is statistically likely that their rates would have decreased somewhat in the following year even without any intervention. This makes it difficult to attribute the entire change to the program itself.
Question 16
A researcher conducts a six-month study to assess the impact of a mindfulness meditation program on employee stress. Three months into the study, the company undergoes a major restructuring, leading to widespread job uncertainty and layoffs. The researcher notes that stress levels, as measured by a self-report questionnaire, are significantly higher at the end of the study than at the beginning. Attributing this increase in stress solely to the failure of the program would ignore which threat to internal validity?
- Testing effect
- Maturation
- History effect (correct answer)
- Instrumentation
Explanation: A history effect is a threat to internal validity that occurs when an external event, unrelated to the treatment, happens during the course of a study and affects the dependent variable. In this case, the company restructuring is a major external event that could have caused the increase in stress, making it impossible to isolate the effect (or lack thereof) of the mindfulness program.
Question 17
A researcher is assessing the split-half reliability of a new 40-item survey. She correlates the sum of scores on the 20 odd-numbered items with the sum of scores on the 20 even-numbered items, yielding a correlation coefficient of r = .70. What is the most appropriate next step and conclusion?
- The reliability is .70, which is acceptable, so the scale can be used as is.
- The reliability is too low, indicating that the test has poor internal consistency.
- The .70 correlation must be adjusted using the Spearman-Brown formula to estimate the reliability of the full-length test. (correct answer)
- This procedure measures test-retest reliability, which should be supplemented with a measure of internal consistency like Cronbach's alpha.
Explanation: The correlation between two halves of a test (.70 in this case) underestimates the reliability of the full test, because reliability is partly a function of test length. The Spearman-Brown prophecy formula is used to correct this correlation and estimate what the reliability would be for the full 40-item test. The resulting corrected value will be higher than .70, likely indicating good internal consistency.
Question 18
A researcher investigating the effect of caffeine on motor skills tells the experimental group, 'You are receiving a high dose of caffeine that should improve your performance.' He tells the control group, 'You are receiving a placebo.' The researcher is not blinded to the group assignments. He observes that the experimental group performs better. The outcome is potentially compromised by which two confounds?
- Selection bias and regression to the mean
- History effect and maturation
- Demand characteristics and experimenter bias (correct answer)
- Attrition and instrumentation effects
Explanation: Demand characteristics occur when participants form an interpretation of the experiment's purpose and subconsciously change their behavior to fit that interpretation. Telling the experimental group to expect improvement creates such a demand. Experimenter bias (or observer-expectancy effect) occurs when a researcher's expectations influence the outcome. Since the researcher is not blinded, his knowledge of who got the caffeine could unconsciously affect how he observes or interacts with the participants, thereby influencing their performance.
Question 19
Two clinical psychologists independently watch a video of a client's therapy session and rate the client's 'expression of affect' on a 1-to-7 scale for 20 distinct one-minute intervals. A comparison of their ratings yields a very low correlation (r = .23). Which of the following is the most direct threat to the validity of the study's conclusions about the client's affect?
- The measure has low inter-rater reliability, which limits its potential validity. (correct answer)
- The measure has low content validity because it only assesses one aspect of the session.
- The rating procedure is subject to a history effect, compromising internal validity.
- The client's behavior may be influenced by demand characteristics.
Explanation: Inter-rater reliability is the degree of agreement between independent observers. A low correlation indicates poor agreement, meaning the ratings are inconsistent and likely contain a large amount of measurement error. A fundamental principle of psychometrics is that a measure cannot be valid if it is not reliable. If the raters cannot agree on what they are seeing, the ratings cannot be an accurate (valid) measure of the client's affect.
Question 20
A cognitive psychologist studies the impact of a new 'brain-training' game on fluid intelligence. An experimental group plays the game for 30 minutes daily for four weeks, while a control group does not. Both groups are given an IQ test before and after the four-week period. The experimental group shows a significantly greater increase in IQ scores. A critic argues that the improvement may be due to the experimental group becoming more comfortable with the testing situation itself. This criticism refers to which potential confound?
- Placebo effect
- Testing effect (correct answer)
- Maturation effect
- Hawthorne effect
Explanation: A testing effect occurs when the act of taking a pre-test influences the scores on a post-test. In this case, the experimental group's improvement might not be due to the brain-training game but to practice or familiarity gained from the first IQ test. While this effect would typically apply to both groups, the criticism implies an interaction where the experimental group benefits more, but the core concept being invoked is the testing effect. It's a plausible alternative explanation for the observed score increase.