All questions
Question 1
To estimate the mean household income in a town, a researcher first divides the town into 20 distinct neighborhoods (clusters). They then draw a simple random sample of 50 households from the entire town's residential address list. A critic argues that because the number of selected households varies unpredictably from neighborhood to neighborhood, the sample is biased. How should this criticism be evaluated?
- The critic is correct; the sample should have been a stratified random sample with households selected from every neighborhood.
- The critic is incorrect; the procedure described is a valid simple random sample (SRS) of households, which is an unbiased method. (correct answer)
- The critic is correct; the procedure is a haphazard sample and does not qualify as a probability sample.
- The critic is incorrect; while the procedure is a cluster sample, the randomness of the selection ensures it is unbiased.
Explanation: The question is designed to create confusion between different sampling methods. The researcher drew a simple random sample from the entire town's list. The fact that the town is also divisible into neighborhoods is extraneous information designed to mislead. The described method is a pure SRS. In an SRS, it is expected that the number of units selected from any particular subgroup (like a neighborhood) will vary randomly. This does not introduce bias. The critic is incorrectly conflating this random variation with a flawed design. The method is not stratified (A), haphazard (C), or cluster sampling (D).
Question 2
To estimate the mean household income in a town, a researcher first divides the town into 20 distinct neighborhoods (clusters). They then draw a simple random sample of 50 households from the entire town's residential address list. A critic argues that because the number of selected households varies unpredictably from neighborhood to neighborhood, the sample is biased. How should this criticism be evaluated?
- The critic is correct; the sample should have been a stratified random sample with households selected from every neighborhood.
- The critic is incorrect; the procedure described is a valid simple random sample (SRS) of households, which is an unbiased method. (correct answer)
- The critic is correct; the procedure is a haphazard sample and does not qualify as a probability sample.
- The critic is incorrect; while the procedure is a cluster sample, the randomness of the selection ensures it is unbiased.
Explanation: The question is designed to create confusion between different sampling methods. The researcher drew a simple random sample from the entire town's list. The fact that the town is also divisible into neighborhoods is extraneous information designed to mislead. The described method is a pure SRS. In an SRS, it is expected that the number of units selected from any particular subgroup (like a neighborhood) will vary randomly. This does not introduce bias. The critic is incorrectly conflating this random variation with a flawed design. The method is not stratified (A), haphazard (C), or cluster sampling (D).
Question 3
A large corporation wants to assess employee satisfaction. The company has 50 office locations across the country. Management randomly selects 5 of these 50 office locations and surveys every single employee who works at those 5 locations.
Which of the following is a primary statistical concern with this sampling methodology?
- The sample size is too small because surveying only 10% of the office locations is insufficient for inference.
- The method does not constitute a probability sample, making it impossible to quantify the sampling error.
- The employees within any given office location may have experiences that are not representative of the entire company. (correct answer)
- The study will suffer from severe nonresponse bias because employees will be hesitant to criticize their employer.
Explanation: This is a cluster sampling design. A major concern with cluster sampling is that the clusters themselves (the office locations) may not be representative of the population of all clusters. For example, the 5 selected offices might be in a particular region or have similar management styles, leading to satisfaction levels that are systematically different from the company as a whole. This intra-cluster correlation can lead to a biased estimate. The sample size (A) is not the core issue. The method (B) is a probability sample because every cluster has a known chance of being selected. Nonresponse bias (D) is a potential problem in any survey but not a flaw inherent to the cluster sampling design itself.
Question 4
A psychology researcher needs to collect data on test anxiety among college students. The researcher offers extra credit to students in their large 'Introduction to Psychology' course in exchange for completing a survey. Why might the findings of this study not be generalizable to the entire college student population?
- The use of extra credit as an incentive is an unethical research practice that invalidates the results.
- The sample is a convenience sample, and students in an introductory psychology course may differ systematically from the general student population. (correct answer)
- The sample size is not sufficiently large to allow for generalization to a larger population.
- Self-reporting on anxiety is inherently unreliable, introducing a high degree of response bias.
Explanation: The primary issue is the use of a convenience sample. The students in one specific course do not represent all students. They may be younger on average, concentrated in certain majors, or have different academic motivations than the overall student body. These systematic differences mean the results on test anxiety may not generalize. While ethics (A) are important, offering minor extra credit is often permissible. Sample size (C) affects precision, not generalizability in this fundamental way. Response bias (D) is a concern about measurement, but the core problem for generalizability is the sampling method itself.
Question 5
A state's department of education wants to estimate the average number of hours high school students spend on homework per week. They first randomly select 10 of the state's 150 school districts. Within each of these 10 districts, they randomly select 3 high schools. Finally, from each selected high school, they obtain a list of all honors students and randomly select 15 of them to survey. Which of the following describes the most significant source of bias in this study design?
- Nonresponse bias, as some honors students may not complete or return the survey.
- Response bias, because honors students may feel pressured to over-report their study hours.
- Undercoverage bias, because the sampling frame is restricted to honors students. (correct answer)
- Bias from using a cluster sample, because students within a school may have similar homework loads.
Explanation: The most significant flaw in the design is a systematic error in defining the sampling frame. By exclusively sampling from lists of honors students, the study systematically excludes all non-honors students. This is a severe form of undercoverage bias, as the sample cannot possibly be representative of all high school students. While nonresponse (A) and response bias (B) could occur, they are potential issues in execution, not a fundamental flaw in the sampling frame design. The use of cluster/multi-stage sampling (D) introduces complexity in analysis but is a valid technique; the main issue is the biased selection within the final stage.
Question 6
A news website places a poll on its homepage asking, 'Do you support the proposed city-wide ban on plastic bags?' The poll runs for 24 hours, and thousands of visitors to the site cast a vote. The website reports the results as an accurate reflection of public opinion. What is the most likely consequence of this polling method?
- The estimate of support for the ban is likely biased because the sample size is excessively large.
- The estimate of support for the ban is likely biased due to the strong opinions of those who choose to participate. (correct answer)
- The results are likely unbiased because the poll was open to all visitors, ensuring a diverse range of opinions.
- The results are likely unbiased because the anonymity of the web encourages truthful responses, reducing response bias.
Explanation: This is a classic example of a voluntary response sample. People who feel strongly about an issue are more likely to participate in such polls. This self-selection creates a sample that is not representative of the general population, leading to voluntary response bias. The results will likely reflect the opinions of a passionate subgroup rather than the entire city. A large sample size (A) does not correct for bias. The poll being open to all visitors (C) is the very reason for the self-selection bias. Anonymity (D) might reduce response bias on sensitive topics, but it does not address the fundamental selection bias of this design.
Question 7
In the 1990s, many national political polls were conducted using random-digit dialing (RDD) of landline telephones. If this same methodology were used today, what would be the most significant source of bias for estimating the political opinions of the entire adult population?
- Nonresponse bias, as people are less willing to answer polls on the phone today than in the past.
- Response bias, as people are more likely to lie about their political opinions over the phone.
- Undercoverage bias, as a large and growing portion of the population uses only cell phones and has no landline. (correct answer)
- Wording bias, as political questions have become more complex and difficult to frame neutrally.
Explanation: The most significant change in the polling landscape since the 1990s is the rise of cell phones and the decline of landlines. A large segment of the population, particularly younger people, is 'cell phone only.' A sampling frame that includes only landlines would systematically exclude these individuals, leading to severe undercoverage bias. The demographics and political opinions of the 'cell phone only' population are known to differ from the landline population, making this a critical flaw. While nonresponse (A), response (B), and wording (D) are always concerns in polling, the structural issue of the outdated landline sampling frame is the most significant new source of bias.
Question 8
A research institute conducted a mail survey using a simple random sample of 2,000 households, but only 800 responded (a 40% response rate). To address potential nonresponse bias, the researchers then drew a simple random sample of 100 of the 1,200 non-responding households and used intensive follow-up (phone calls, home visits) to achieve a 95% response rate within that subsample. How can the researchers best use the data from this follow-up sample?
- They should replace 95 of the original non-responses with the new responses to increase the overall sample size.
- They should combine the original 800 responses with the 95 follow-up responses and treat it as a single random sample of 895 households.
- They should analyze the follow-up sample separately and report two different sets of findings for the two groups.
- They should compare the responses of the follow-up sample to the original respondents and use weighting to adjust the initial results. (correct answer)
Explanation: The purpose of sampling non-respondents is to understand if and how they differ from the original respondents. If the follow-up group's characteristics or answers differ systematically, it indicates the presence of nonresponse bias. The correct statistical procedure is to use this information to develop weights for the original 800 responses to make the final, weighted sample more representative of the entire population of 2,000. Simply adding the responses (A, B) is incorrect because the follow-up group was sampled at a different rate and is not representative of the whole sample. Analyzing them separately (C) is a preliminary step, but the ultimate goal is to produce a single, corrected estimate for the target population.
Question 9
To estimate the proportion of students who use the campus recreation center, a university researcher obtains a complete list of all students living in university dormitories and draws a simple random sample of 200 students from this list. The university has a large population of students who live off-campus. The primary source of bias in this study is:
- Voluntary response bias, because students can choose whether or not to use the recreation center.
- Response bias, because students might not accurately recall how often they use the center.
- Nonresponse bias, because it is likely that not all 200 selected students will complete the survey.
- Undercoverage, because students living off-campus have no chance of being included in the sample. (correct answer)
Explanation: The sampling frame (the list from which the sample is drawn) consists only of students living in dormitories. This systematically excludes the entire population of students living off-campus. This is the definition of undercoverage bias. It is very likely that the recreation center usage patterns of off-campus students differ from those of on-campus students, so their exclusion will almost certainly bias the results. Voluntary response bias (A) applies to self-selected samples, which this is not. Response bias (B) and nonresponse bias (C) are potential issues, but the most fundamental flaw in the study's design is the incomplete sampling frame leading to undercoverage.
Question 10
A researcher conducts a survey using a random sample of registered voters. The final sample contains 60% women and 40% men. However, the population of registered voters is known to be 52% women and 48% men. The researcher adjusts for this discrepancy by giving each man's response a weight of 1.2 (since 48/40=1.2) and each woman's response a weight of approximately 0.87 (since 52/60≈0.87). What is the primary purpose of this weighting procedure?
- To increase the effective sample size and thus increase the precision of the estimates.
- To correct for potential bias introduced by the disproportionate representation of genders in the sample. (correct answer)
- To ethically ensure that the opinions of men and women are counted equally in the final analysis.
- To transform the random sample into a stratified sample after data collection has been completed.
Explanation: This procedure is called post-stratification weighting. It is used to correct for situations where the final sample, due to random chance or nonresponse, is not representative of the population on key demographic variables. By down-weighting the overrepresented group (women) and up-weighting the underrepresented group (men), the researcher aims to make the sample results better reflect the true population structure, thus reducing a potential source of bias. This does not increase the sample size or precision (A). While it has an ethical dimension (C), its primary purpose is statistical correction of bias. It doesn't retroactively create a stratified sample (D), but rather adjusts a simple random sample to align with population strata.
Question 11
A political campaign wants to gauge voter opinion in a district with 35% urban, 45% suburban, and 20% rural voters. They plan to survey 1,000 voters using proportional stratified sampling. They acquire voter lists for each of the three regions. Which of the following procedures would violate the principles of this sampling design and introduce selection bias?
- Randomly selecting 350 voters from the urban list, 450 from the suburban list, and 200 from the rural list.
- Using a random number generator to select voters from each list, but stopping when the target number for one stratum is reached.
- Numbering the voters on each list and selecting every 100th voter after a random start for each list.
- Contacting the first 350 listed urban voters, 450 suburban voters, and 200 rural voters who are reachable by phone. (correct answer)
Explanation: Proportional stratified sampling requires two key steps: (1) dividing the population into strata and (2) taking a simple random sample (or other probability sample) from within each stratum. Option (D) violates the second step by selecting voters based on convenience (being first on the list and reachable by phone) rather than random chance. This introduces selection bias. Option (A) correctly describes the goal of the design. Option (B) describes a valid way of executing the random sampling. Option (C) describes systematic sampling within each stratum, which is a valid probability sampling method.
Question 12
A university is considering a campus-wide smoking ban. To gauge opinion, they plan to ask students the following question: 'Given the well-documented health risks associated with secondhand smoke, do you agree that the university has an obligation to protect its community by implementing a campus-wide smoking ban?' Which of the following is the primary concern with this survey question?
- It may create undercoverage bias by excluding faculty and staff from the survey.
- It is a leading question that is likely to produce biased responses in favor of the ban. (correct answer)
- The question is too long, which could lead to a high nonresponse rate.
- The use of the term 'obligation' is too ambiguous for a formal survey question.
Explanation: This question suffers from wording bias. By prefacing the question with phrases like 'well-documented health risks' and 'obligation to protect its community,' it frames the issue in a way that encourages respondents to agree with the ban. This is a leading question that pressures the respondent toward a particular answer, biasing the results. Undercoverage (A) is a sampling issue, not an issue with the question itself. While the question length (C) or ambiguity (D) could be minor issues, the most significant flaw is the biased and leading wording.
Question 13
A public health organization conducted a survey on personal health habits. One question asked, 'In the past year, how many times, if any, have you engaged in binge drinking (consuming 5 or more drinks on one occasion)?' The survey was conducted with a well-designed random sample and had a high response rate. The resulting estimate of binge drinking frequency is likely to be...
- Biased downward, because respondents may be hesitant to admit to a socially undesirable behavior. (correct answer)
- Biased upward, because respondents may exaggerate their behavior to appear more social.
- Unbiased, because the random sample and high response rate eliminate major sources of error.
- Unreliable, because the definition of 'binge drinking' is not universally understood by all participants.
Explanation: When evaluating survey data about sensitive behaviors, you need to consider how social desirability bias affects responses. Even with excellent survey methodology, certain types of questions create systematic response patterns that can skew results.
Binge drinking carries negative social connotations and potential legal concerns. When asked about such behaviors, respondents often underreport the true frequency, even in anonymous surveys. This creates a consistent downward bias - the survey estimate will likely be lower than the actual population parameter. The well-designed random sample and high response rate don't eliminate this psychological barrier to truthful reporting.
Let's examine why the other options miss the mark: Option B suggests upward bias from exaggeration, but research consistently shows people underreport rather than overreport stigmatized behaviors like excessive drinking. Option C incorrectly assumes that good sampling methodology eliminates all bias - while random sampling addresses selection bias and high response rates reduce nonresponse bias, neither addresses response bias from sensitive questions. Option D focuses on definitional confusion, but the question provides a clear, specific definition of binge drinking (5+ drinks per occasion), making comprehension issues unlikely.
The key insight is that methodological excellence can't overcome human psychology. When people feel judged about their behavior, they tend to present themselves in a more favorable light, leading to systematic underreporting.
Study tip: For statistics questions about sensitive topics (drinking, drugs, income, sexual behavior), always consider social desirability bias as a source of systematic error that sampling design alone cannot fix.
Question 14
A researcher wants to study the relationship between weekly exercise hours and stress levels among employees at a large tech company. The employees are categorized into three job levels: entry-level, mid-career, and senior management. To ensure the sample reflects the company's structure, the researcher decides to use stratified random sampling. Which of the following statements provides the best statistical justification for using job level as the stratification variable?
- It is best to have strata of approximately equal size to simplify the statistical analysis.
- Employees within each job level are likely to be more homogeneous with respect to exercise and stress than the employee population as a whole. (correct answer)
- Using job level as a stratification variable guarantees that the sample will be geographically diverse.
- This method ensures that every employee has an equal probability of being selected for the survey.
Explanation: The primary goal of stratification is to increase the precision of estimates by dividing the population into subgroups (strata) that are internally homogeneous but externally heterogeneous. In this context, it is reasonable to assume that job level is related to both stress and available time for exercise. Therefore, employees within a single job level are likely more similar to each other on these variables than they are to employees at different levels. This homogeneity within strata makes stratified sampling more efficient than a simple random sample. Strata do not need to be of equal size (A). Stratifying by job level has no bearing on geographic diversity (C). In proportional stratified sampling, individual selection probabilities are equal, but not in disproportional sampling; more importantly, this is a result, not the primary justification for the method (D).
Question 15
A door-to-door survey is being conducted in a politically diverse neighborhood to ask about opinions on a local school bond. The research firm discovers that interviewers who are older and formally dressed are getting significantly more conservative responses than younger, casually dressed interviewers, even when asking the questions identically. This phenomenon, where characteristics of the interviewer affect the answers given by respondents, is known as what type of bias?
- Interviewer effect (correct answer)
- Nonresponse bias
- Selection bias
- Undercoverage bias
Explanation: When you encounter survey research scenarios, focus on identifying what specific aspect of the data collection process is creating problems. Different types of bias arise from different stages and mechanisms in the research process.
In this scenario, the key issue is that the interviewer's personal characteristics (age and dress style) are systematically influencing how respondents answer the same questions. This is a classic example of interviewer effect (choice A) - a type of bias where the interviewer's attributes, behavior, or presence affects respondents' answers, even when the questions are asked identically.
Let's examine why the other options don't fit: Nonresponse bias (B) occurs when certain types of people systematically refuse to participate in the survey, skewing results toward those who do respond. Here, people are responding - they're just responding differently based on who's asking. Selection bias (C) happens when the sample isn't representative of the target population due to how participants were chosen. The neighborhood appears properly selected; the issue is data collection, not sampling. Undercoverage bias (D) occurs when some groups in the target population have little or no chance of being included in the sample. Again, the coverage seems adequate - the problem is how responses vary by interviewer.
Study tip: Remember that interviewer effect specifically involves the interviewer's characteristics or behavior influencing responses. Look for scenarios where identical questions yield different responses based on who's asking, rather than who's being selected or who's refusing to participate.
Question 16
To estimate the proportion of students who use the campus recreation center, a university researcher obtains a complete list of all students living in university dormitories and draws a simple random sample of 200 students from this list. The university has a large population of students who live off-campus. The primary source of bias in this study is:
- Voluntary response bias, because students can choose whether or not to use the recreation center.
- Response bias, because students might not accurately recall how often they use the center.
- Nonresponse bias, because it is likely that not all 200 selected students will complete the survey.
- Undercoverage, because students living off-campus have no chance of being included in the sample. (correct answer)
Explanation: The sampling frame (the list from which the sample is drawn) consists only of students living in dormitories. This systematically excludes the entire population of students living off-campus. This is the definition of undercoverage bias. It is very likely that the recreation center usage patterns of off-campus students differ from those of on-campus students, so their exclusion will almost certainly bias the results. Voluntary response bias (A) applies to self-selected samples, which this is not. Response bias (B) and nonresponse bias (C) are potential issues, but the most fundamental flaw in the study's design is the incomplete sampling frame leading to undercoverage.
Question 17
A news website places a poll on its homepage asking, 'Do you support the proposed city-wide ban on plastic bags?' The poll runs for 24 hours, and thousands of visitors to the site cast a vote. The website reports the results as an accurate reflection of public opinion. What is the most likely consequence of this polling method?
- The estimate of support for the ban is likely biased because the sample size is excessively large.
- The estimate of support for the ban is likely biased due to the strong opinions of those who choose to participate. (correct answer)
- The results are likely unbiased because the poll was open to all visitors, ensuring a diverse range of opinions.
- The results are likely unbiased because the anonymity of the web encourages truthful responses, reducing response bias.
Explanation: This is a classic example of a voluntary response sample. People who feel strongly about an issue are more likely to participate in such polls. This self-selection creates a sample that is not representative of the general population, leading to voluntary response bias. The results will likely reflect the opinions of a passionate subgroup rather than the entire city. A large sample size (A) does not correct for bias. The poll being open to all visitors (C) is the very reason for the self-selection bias. Anonymity (D) might reduce response bias on sensitive topics, but it does not address the fundamental selection bias of this design.
Question 18
A university is considering a campus-wide smoking ban. To gauge opinion, they plan to ask students the following question: 'Given the well-documented health risks associated with secondhand smoke, do you agree that the university has an obligation to protect its community by implementing a campus-wide smoking ban?' Which of the following is the primary concern with this survey question?
- It may create undercoverage bias by excluding faculty and staff from the survey.
- It is a leading question that is likely to produce biased responses in favor of the ban. (correct answer)
- The question is too long, which could lead to a high nonresponse rate.
- The use of the term 'obligation' is too ambiguous for a formal survey question.
Explanation: This question suffers from wording bias. By prefacing the question with phrases like 'well-documented health risks' and 'obligation to protect its community,' it frames the issue in a way that encourages respondents to agree with the ban. This is a leading question that pressures the respondent toward a particular answer, biasing the results. Undercoverage (A) is a sampling issue, not an issue with the question itself. While the question length (C) or ambiguity (D) could be minor issues, the most significant flaw is the biased and leading wording.
Question 19
A polling organization uses a perfectly executed simple random sampling (SRS) design to select 1,200 voters from a large population. They find that 54% of the sample supports Candidate A. However, an election held the next day reveals that 51% of the entire population of voters supports Candidate A. Assuming no voters changed their minds overnight, which is the most likely explanation for the discrepancy between the sample statistic (54%) and the population parameter (51%)?
- The poll must have suffered from nonresponse bias, as some voters did not respond.
- The discrepancy is most likely due to random sampling error or variability. (correct answer)
- The wording of the poll question must have been biased in favor of Candidate A.
- The sampling frame must have been incomplete, leading to undercoverage bias.
Explanation: The question specifies that a 'perfectly executed simple random sampling' design was used. This is meant to direct the student away from sources of bias (non-sampling errors). Even with a perfect random sample, the sample statistic (e.g., sample proportion) will almost never be exactly equal to the population parameter due to random chance. This natural, chance variation is called sampling error or sampling variability. The 3% difference is a plausible amount of sampling error for a sample of this size. The other options (A, C, D) all describe types of bias (non-sampling errors), which the problem setup implies are not the issue.
Question 20
A public health organization conducted a survey on personal health habits. One question asked, 'In the past year, how many times, if any, have you engaged in binge drinking (consuming 5 or more drinks on one occasion)?' The survey was conducted with a well-designed random sample and had a high response rate. The resulting estimate of binge drinking frequency is likely to be...
- Biased downward, because respondents may be hesitant to admit to a socially undesirable behavior. (correct answer)
- Biased upward, because respondents may exaggerate their behavior to appear more social.
- Unbiased, because the random sample and high response rate eliminate major sources of error.
- Unreliable, because the definition of 'binge drinking' is not universally understood by all participants.
Explanation: When evaluating survey data about sensitive behaviors, you need to consider how social desirability bias affects responses. Even with excellent survey methodology, certain types of questions create systematic response patterns that can skew results.
Binge drinking carries negative social connotations and potential legal concerns. When asked about such behaviors, respondents often underreport the true frequency, even in anonymous surveys. This creates a consistent downward bias - the survey estimate will likely be lower than the actual population parameter. The well-designed random sample and high response rate don't eliminate this psychological barrier to truthful reporting.
Let's examine why the other options miss the mark: Option B suggests upward bias from exaggeration, but research consistently shows people underreport rather than overreport stigmatized behaviors like excessive drinking. Option C incorrectly assumes that good sampling methodology eliminates all bias - while random sampling addresses selection bias and high response rates reduce nonresponse bias, neither addresses response bias from sensitive questions. Option D focuses on definitional confusion, but the question provides a clear, specific definition of binge drinking (5+ drinks per occasion), making comprehension issues unlikely.
The key insight is that methodological excellence can't overcome human psychology. When people feel judged about their behavior, they tend to present themselves in a more favorable light, leading to systematic underreporting.
Study tip: For statistics questions about sensitive topics (drinking, drugs, income, sexual behavior), always consider social desirability bias as a source of systematic error that sampling design alone cannot fix.