All questions
Question 1
A study uses an automated text analysis program to code the sentiment of tweets about a political candidate. To check the quality of this measure, the researchers manually code a random subsample of 1,000 tweets and find that the computer's sentiment score for a given tweet is highly correlated with the human-coded score. This procedure is primarily designed to establish the measure's:
- test-retest reliability.
- internal validity.
- criterion validity. (correct answer)
- external validity.
Explanation: Criterion validity assesses how well a measure relates to an established, external criterion. In this case, the 'gold standard' or criterion is the sentiment assigned by trained human coders. By showing that the automated program's scores correlate highly with the human scores (the criterion), the researchers are providing strong evidence for the criterion validity (specifically, concurrent validity) of their automated measure. Test-retest reliability (A) would involve running the program on the same tweets at two different times. Internal (B) and external (D) validity are properties of study designs, not measures.
Question 2
A political science study finds that a get-out-the-vote (GOTV) intervention conducted by highly charismatic and trained student volunteers significantly increases voter turnout. Several campaigns attempt to replicate the intervention using their regular, less-trained volunteers but see no effect. The failure to replicate the original finding most likely stems from an issue with the original study's:
- construct validity, as 'voter turnout' was measured incorrectly.
- internal validity, as the original study failed to establish a true causal link.
- reliability, as the intervention was not administered consistently in the original study.
- external validity, as the effect was dependent on specific conditions not present elsewhere. (correct answer)
Explanation: External validity concerns whether the results of a study can be generalized to other people, settings, and conditions. In this case, the effect found in the original study appears to be dependent on a specific condition: the use of highly charismatic and trained volunteers. When this condition is changed (by using regular volunteers), the effect disappears. This suggests the original finding is not generalizable, which is a problem of external validity. There is no information to suggest issues with construct validity (A) or reliability (C) in the original study. The original study may have had high internal validity (B), but the causal relationship it identified was specific to its unique context.
Question 3
A researcher is studying the impact of campaign spending on election outcomes using observational data from 50 states over 20 years. The researcher uses a sophisticated statistical model to control for dozens of potentially confounding variables. A reviewer raises a concern that, despite the controls, the states that spend more on campaigns may differ from those that spend less in unmeasured ways, such as the underlying political culture. This concern fundamentally questions the study's:
- reliability of the spending data.
- external validity to other countries.
- internal validity of the causal claim. (correct answer)
- content validity of the 'election outcome' measure.
Explanation: Internal validity is about the soundness of the causal inference. The reviewer's concern is that an unmeasured confounding variable (political culture) might be responsible for both higher campaign spending and certain election outcomes. This is a problem of omitted variable bias, which is a direct threat to internal validity. Even with many controls, observational studies can never be certain they have accounted for all potential confounds. The criticism is not about the consistency of the data (A), its generalizability to other countries (B), or how an outcome is measured (D), but about whether the claimed causal relationship between spending and outcomes is real or spurious.
Question 4
To measure 'support for democratic norms,' a professor creates a survey consisting of ten questions about economic policy preferences, such as views on taxation and regulation. The survey produces consistent results when administered multiple times. The primary weakness of this measure is its lack of:
- test-retest reliability, because economic views can fluctuate.
- internal validity, because it cannot establish causality between norms and policies.
- external validity, because the sample may not be representative.
- content validity, because it fails to capture the full domain of the concept. (correct answer)
Explanation: Content validity refers to how well a measure covers the full range of meanings and dimensions of the concept it is intended to measure. 'Support for democratic norms' includes concepts like free speech, minority rights, and rule of law, not just economic policy. By focusing exclusively on economic questions, the survey omits crucial aspects of the concept, thus suffering from poor content validity. The stem states the results are consistent, so reliability (A) is not the primary weakness. Internal (B) and external (C) validity are properties of a study's design and inferences, not a measure itself.
Question 5
To measure 'support for democratic norms,' a professor creates a survey consisting of ten questions about economic policy preferences, such as views on taxation and regulation. The survey produces consistent results when administered multiple times. The primary weakness of this measure is its lack of:
- test-retest reliability, because economic views can fluctuate.
- internal validity, because it cannot establish causality between norms and policies.
- external validity, because the sample may not be representative.
- content validity, because it fails to capture the full domain of the concept. (correct answer)
Explanation: Content validity refers to how well a measure covers the full range of meanings and dimensions of the concept it is intended to measure. 'Support for democratic norms' includes concepts like free speech, minority rights, and rule of law, not just economic policy. By focusing exclusively on economic questions, the survey omits crucial aspects of the concept, thus suffering from poor content validity. The stem states the results are consistent, so reliability (A) is not the primary weakness. Internal (B) and external (C) validity are properties of a study's design and inferences, not a measure itself.
Question 6
A researcher develops a new scale to measure populist attitudes. To assess its validity, the researcher finds that scores on the scale are strongly correlated with support for parties widely considered populist, and negatively correlated with measures of support for global institutions. This process is an example of establishing:
- test-retest reliability.
- content validity.
- construct validity. (correct answer)
- inter-rater reliability.
Explanation: Construct validity is the extent to which a measure behaves in a way consistent with theoretical expectations. It involves examining the relationships between the measure and other related concepts. By showing that the new scale correlates positively with support for populist parties (convergent validity) and negatively with support for global institutions (discriminant validity), the researcher is accumulating evidence for its construct validity. The measure is relating to other constructs in the theoretically expected ways. This is not about reliability (A, D) or simply covering the concept's domain (B).
Question 7
A researcher conducts a longitudinal study on the effects of a civics education program on high school students' political engagement. A pre-test is administered in September, the program is implemented throughout the school year, and a post-test is administered in May. Between the pre-test and post-test, a highly contentious and mobilizing presidential election occurs. The researcher finds a large increase in political engagement. The election represents a significant threat to the study's internal validity known as:
- Maturation
- Testing
- History (correct answer)
- Instrumentation
Explanation: The 'history' threat to internal validity occurs when an external event, unrelated to the treatment (the civics program), happens between the pre-test and post-test and affects the outcome variable. The presidential election is a major historical event that could increase political engagement, making it impossible to attribute the change solely to the education program. Maturation (A) refers to natural changes in the subjects over time (e.g., growing older). Testing (B) refers to the effect of the pre-test on the post-test scores. Instrumentation (D) refers to changes in the measurement instrument itself.
Question 8
A researcher hypothesizes that a new online deliberation platform increases participants' political efficacy. To test this, the researcher needs to ensure that the study has high internal validity. Which of the following design choices is most critical for achieving this goal?
- Recruiting a large and nationally representative sample of participants.
- Using a survey with high test-retest reliability to measure political efficacy.
- Randomly assigning participants to either use the platform or be in a control group. (correct answer)
- Ensuring the experimental conditions closely mirror real-world online discussions.
Explanation: Internal validity is concerned with establishing a cause-and-effect relationship and ruling out alternative explanations (confounds). The most critical element for this is random assignment to treatment and control groups. Randomization helps ensure that, on average, the two groups are comparable on all potential confounding variables before the treatment is administered, thus isolating the effect of the treatment itself. A representative sample (A) and realistic conditions (D) are crucial for external validity. A reliable measure (B) is necessary for a good study, but it does not by itself ensure internal validity; one could reliably measure an effect that is actually caused by a confounding variable.
Question 9
A study uses an automated text analysis program to code the sentiment of tweets about a political candidate. To check the quality of this measure, the researchers manually code a random subsample of 1,000 tweets and find that the computer's sentiment score for a given tweet is highly correlated with the human-coded score. This procedure is primarily designed to establish the measure's:
- test-retest reliability.
- internal validity.
- criterion validity. (correct answer)
- external validity.
Explanation: Criterion validity assesses how well a measure relates to an established, external criterion. In this case, the 'gold standard' or criterion is the sentiment assigned by trained human coders. By showing that the automated program's scores correlate highly with the human scores (the criterion), the researchers are providing strong evidence for the criterion validity (specifically, concurrent validity) of their automated measure. Test-retest reliability (A) would involve running the program on the same tweets at two different times. Internal (B) and external (D) validity are properties of study designs, not measures.
Question 10
A political scientist wants to generalize the findings of their experiment on campaign messaging to the entire voting population of the United States. Which of the following is most essential for ensuring high external validity of the findings?
- Randomly assigning participants to the different messaging groups.
- Using a measurement of voting intention that is highly reliable.
- Conducting the experiment in a tightly controlled laboratory setting.
- Recruiting a sample of participants that is representative of the U.S. voting population. (correct answer)
Explanation: External validity is the extent to which study findings can be generalized to other populations, settings, or times. To generalize findings to the entire U.S. voting population, it is most critical to have a sample that is representative of that population in terms of key demographics like age, gender, race, income, and political affiliation. Random assignment (A) is key for internal validity. A reliable measure (B) is important for any good study but doesn't guarantee generalizability. A controlled lab setting (C) often weakens, rather than strengthens, external validity by creating an artificial environment.
Question 11
A researcher is concerned that in her experiment on political persuasion, participants in the control group might become aware of the treatment being received by the other group and feel resentful, potentially altering their behavior on the outcome measure. This phenomenon, where the control group's behavior is altered due to their awareness of the treatment group, is a threat to internal validity known as:
- compensatory rivalry.
- resentful demoralization. (correct answer)
- the Hawthorne effect.
- diffusion of treatment.
Explanation: Resentful demoralization is a specific threat to internal validity where participants in the control group, who are not receiving a desirable treatment, become discouraged or resentful. This can lead them to perform worse on the outcome measure than they otherwise would have, artificially inflating the apparent effect of the treatment. Compensatory rivalry (A) is the opposite, where the control group works harder to overcome the perceived disadvantage. The Hawthorne effect (C) is when subjects' behavior changes simply because they are being observed. Diffusion of treatment (D) is when the treatment itself accidentally spreads to the control group.
Question 12
An experiment is designed to test whether watching a negative political advertisement increases voter apathy. The researchers measure apathy before and after the ad is shown. They are concerned that the pre-test measurement itself might make participants more aware of their own feelings of apathy, causing them to respond differently to the post-test, regardless of the ad's content. This potential issue is a threat to internal validity known as:
- selection bias.
- maturation.
- history.
- testing. (correct answer)
Explanation: The 'testing' effect occurs when the act of taking a pre-test influences the scores on a post-test. The concern that the initial apathy measurement makes participants more sensitive or aware, thereby changing their subsequent responses, is a classic example of this threat to internal validity. The effect is not due to the treatment (the ad) but to the measurement process itself. Selection bias (A) relates to non-random assignment. Maturation (B) refers to internal changes in participants over time. History (C) refers to an external event affecting the outcome.
Question 13
A researcher conducts a longitudinal study on the effects of a civics education program on high school students' political engagement. A pre-test is administered in September, the program is implemented throughout the school year, and a post-test is administered in May. Between the pre-test and post-test, a highly contentious and mobilizing presidential election occurs. The researcher finds a large increase in political engagement. The election represents a significant threat to the study's internal validity known as:
- Maturation
- Testing
- History (correct answer)
- Instrumentation
Explanation: The 'history' threat to internal validity occurs when an external event, unrelated to the treatment (the civics program), happens between the pre-test and post-test and affects the outcome variable. The presidential election is a major historical event that could increase political engagement, making it impossible to attribute the change solely to the education program. Maturation (A) refers to natural changes in the subjects over time (e.g., growing older). Testing (B) refers to the effect of the pre-test on the post-test scores. Instrumentation (D) refers to changes in the measurement instrument itself.
Question 14
Two graduate students are tasked with coding the tone of newspaper editorials regarding a new environmental policy. Student 1 codes 75% of the editorials as 'negative,' while Student 2, using the same coding scheme, codes only 45% as 'negative.' This discrepancy indicates a significant problem with which of the following?
- Inter-rater reliability (correct answer)
- Test-retest reliability
- Predictive validity
- Content validity
Explanation: Inter-rater reliability refers to the degree of agreement between different observers or coders who are measuring the same phenomenon. The large difference in the percentage of editorials coded as 'negative' by the two students shows a lack of consistency between them, which is a clear failure of inter-rater reliability. Test-retest reliability (B) involves consistency over time, not between coders. Predictive validity (C) concerns how well a measure predicts a future outcome. Content validity (D) concerns whether a measure covers all facets of a concept.
Question 15
A study compared voting rates between two groups of citizens. Group A received a series of mailers encouraging them to vote, while Group B did not. The researchers did not randomly assign citizens to the groups; instead, Group A consisted of individuals who had opted in to receive political communications. The study found that Group A had a higher turnout rate. Why is the study's internal validity compromised?
- The sample was not representative of all voters, threatening external validity.
- The mailers may not have been a reliable method for encouraging voting.
- A selection bias exists because the groups may have differed in political engagement from the start. (correct answer)
- The Hawthorne effect may have caused participants to change their behavior because they were being studied.
Explanation: Internal validity is the degree to which a study can establish a causal relationship. The lack of random assignment creates a selection bias. The individuals who opted in to receive communications (Group A) are likely more politically engaged to begin with than those in Group B. Therefore, their higher turnout may be due to this pre-existing difference rather than the mailers, confounding the causal claim. This is a classic example of selection bias threatening internal validity. Choice A describes a threat to external validity. Choice B questions the treatment's effectiveness, but the core methodological flaw is the bias. Choice D (Hawthorne effect) is possible but selection bias is the more direct and certain flaw given the design.
Question 16
A research firm is hired to evaluate public support for a proposed tax increase. They conduct a phone survey using random digit dialing. However, the survey is conducted only between 9 a.m. and 5 p.m. on weekdays. The results show that 60% of the population opposes the tax. Critics argue the results are flawed.
The methodology described in the passage introduces a significant threat to the survey's external validity because:
- the survey questions may have been worded in a leading manner.
- the sample is likely to be unrepresentative of the entire population. (correct answer)
- respondents might not give truthful answers over the phone.
- the causal relationship between demographics and tax opinion is unclear.
Explanation: External validity is the extent to which results can be generalized. By surveying only during standard working hours (9 a.m. to 5 p.m. on weekdays), the methodology systematically excludes people who work during those hours and cannot answer the phone. This creates a non-representative sample that is likely skewed towards retirees, stay-at-home parents, and the unemployed, whose views on a tax increase might differ from the general working population. This sampling bias directly threatens the ability to generalize the findings to the entire population. Leading questions (A) and untruthful answers (C) are measurement validity issues. Causal relationships (D) relate to internal validity, which is not the primary goal of this descriptive survey.
Question 17
A team of researchers is studying the effect of exposure to partisan news on political polarization. They recruit participants from a national, representative sample. Participants are randomly assigned to watch either a partisan news broadcast or a neutral news broadcast for 30 minutes. The researchers find that the group watching partisan news reports significantly higher levels of polarization. However, a critic notes that the study took place in a university laboratory, a highly controlled and artificial environment. This criticism most directly challenges the study's:
- internal validity, because the causal mechanism cannot be isolated from the setting.
- external validity, because the findings may not be generalizable to real-world media consumption contexts. (correct answer)
- reliability, because the measurement of polarization might be inconsistent across different settings.
- construct validity, because the concept of 'polarization' is not being accurately measured by the survey instrument.
Explanation: The criticism focuses on the artificiality of the laboratory setting. This is a classic threat to external validity, which concerns the extent to which the results of a study can be generalized to other settings, populations, or times. While the random assignment and controlled environment give the study high internal validity (confidence in the causal link within the study), the lab setting may not reflect how people consume media in their daily lives, thus threatening the generalizability (external validity) of the findings. Internal validity (A) is actually strengthened by the controlled setting. Reliability (C) refers to the consistency of the measure, not its generalizability. Construct validity (D) refers to whether the measure accurately captures the concept, which is a different issue from the study's setting.
Question 18
Two graduate students are tasked with coding the tone of newspaper editorials regarding a new environmental policy. Student 1 codes 75% of the editorials as 'negative,' while Student 2, using the same coding scheme, codes only 45% as 'negative.' This discrepancy indicates a significant problem with which of the following?
- Inter-rater reliability (correct answer)
- Test-retest reliability
- Predictive validity
- Content validity
Explanation: Inter-rater reliability refers to the degree of agreement between different observers or coders who are measuring the same phenomenon. The large difference in the percentage of editorials coded as 'negative' by the two students shows a lack of consistency between them, which is a clear failure of inter-rater reliability. Test-retest reliability (B) involves consistency over time, not between coders. Predictive validity (C) concerns how well a measure predicts a future outcome. Content validity (D) concerns whether a measure covers all facets of a concept.
Question 19
A political science study finds that a get-out-the-vote (GOTV) intervention conducted by highly charismatic and trained student volunteers significantly increases voter turnout. Several campaigns attempt to replicate the intervention using their regular, less-trained volunteers but see no effect. The failure to replicate the original finding most likely stems from an issue with the original study's:
- construct validity, as 'voter turnout' was measured incorrectly.
- internal validity, as the original study failed to establish a true causal link.
- reliability, as the intervention was not administered consistently in the original study.
- external validity, as the effect was dependent on specific conditions not present elsewhere. (correct answer)
Explanation: External validity concerns whether the results of a study can be generalized to other people, settings, and conditions. In this case, the effect found in the original study appears to be dependent on a specific condition: the use of highly charismatic and trained volunteers. When this condition is changed (by using regular volunteers), the effect disappears. This suggests the original finding is not generalizable, which is a problem of external validity. There is no information to suggest issues with construct validity (A) or reliability (C) in the original study. The original study may have had high internal validity (B), but the causal relationship it identified was specific to its unique context.
Question 20
Which of the following statements accurately describes the logical relationship between reliability and validity in social science measurement?
- A measure that is valid must also be reliable. (correct answer)
- A measure that is reliable must also be valid.
- Reliability and validity are independent concepts with no logical relationship.
- A measure can be valid for one population but unreliable for another.
Explanation: Reliability is a necessary but not sufficient condition for validity. A measure cannot be valid (i.e., accurately measuring the intended concept) if it is not reliable (i.e., producing consistent results). If a measure gives wildly different results each time it is used, it cannot possibly be measuring the true underlying concept accurately. Therefore, if a measure is determined to be valid, it must, by definition, also be reliable. The converse is not true; a measure can be highly reliable—consistently measuring the same thing—but that 'thing' may not be the concept of interest (B). They are not independent (C). While validity can be population-dependent, reliability is a property of the measure itself, though it is tested on populations (D).