All questions
Question 1
In a double-blind, placebo-controlled trial of a new vaccine, 10,000 participants were randomly assigned to either the vaccine group or the placebo group. After the assignment, the researchers compiled a table comparing baseline characteristics (e.g., age, sex, pre-existing conditions) of the two groups. What is the primary purpose of creating and examining this table?
- To identify which baseline characteristics are the most significant predictors of the disease outcome.
- To ensure the results of the study can be generalized to the entire population.
- To confirm that the double-blinding procedure was effectively implemented.
- To assess whether the randomization process was successful in creating comparable groups. (correct answer)
Explanation: This table, often called 'Table 1' in clinical trial reports, is used to check the success of randomization. While randomization is expected to balance all confounders (known and unknown) in the long run, it's important to check in any given study whether the groups appear comparable on key baseline variables. Large imbalances could occur by chance and might need to be adjusted for in the analysis. This process checks internal validity, not generalizability (B). It does not check blinding (C) or identify predictors of the outcome (A) as its primary goal.
Question 2
A hospital analyzes its survival rates for two surgical procedures, Procedure A and Procedure B. The overall data show that Procedure A has a higher survival rate. However, when the data are stratified by the patient's pre-operative condition (rated as 'good' or 'poor'), the data show that Procedure B has a higher survival rate for patients in good condition, and also a higher survival rate for patients in poor condition.
Based on the passage, what is the most likely explanation for this statistical reversal?
- There must be a calculation error, as it is mathematically impossible for the overall trend to be the reverse of the trend in all subgroups.
- The sample sizes within each patient condition group were too small, leading to random fluctuations that caused the apparent reversal.
- Patient condition is a confounding variable, and Procedure A was performed disproportionately on patients who were in good condition to begin with. (correct answer)
- There is a strong interaction effect between the procedure type and patient condition, but this does not involve any confounding.
Explanation: This is a classic example of Simpson's Paradox. The reversal occurs because a confounding variable (patient condition) is associated with both the explanatory variable (procedure type) and the response (survival). If the 'easier' patients (good condition) are disproportionately given Procedure A, it can make Procedure A look better overall, even if it is inferior for every type of patient. This is not a calculation error (A) but a real phenomenon. While small samples (B) can be misleading, this systematic reversal points to confounding. Interaction (D) is not sufficient to explain the complete reversal of the association.
Question 3
Researchers are studying the relationship between the number of hours a student works at a part-time job (explanatory variable) and their final exam score (response variable). They consider the student's field of study (e.g., STEM, Humanities) as a potential third variable. Under which of the following conditions would 'field of study' fail to meet the criteria of a confounding variable?
- If the relationship between work hours and exam score is different for STEM students compared to Humanities students.
- If students in STEM fields, on average, work the same number of hours as students in Humanities fields. (correct answer)
- If STEM courses are generally more demanding, leading to lower average exam scores regardless of work hours.
- If students in more demanding fields tend to work fewer hours to have more study time.
Explanation: A variable is a confounder only if it is associated with both the explanatory and response variables. If students in different fields of study work the same number of hours on average, then the link between 'field of study' and the explanatory variable (work hours) is broken. Therefore, it cannot be a confounder, even if it is related to the exam score. Choice A describes interaction. Choice C establishes a link to the response variable. Choice D establishes links to both explanatory and response variables, which would make it a confounder.
Question 4
A medical researcher wants to test the effectiveness of a new blood pressure medication. They are concerned that the effect of the medication might differ significantly between men and women, and that any imbalance of sex between the treatment and control groups could bias the results. To control for the potential confounding effect of sex, which experimental design is most appropriate?
- Randomly assign all subjects to the medication or a placebo, then use statistical adjustment for sex in the final analysis.
- Group subjects by sex, and then within each group, randomly assign subjects to either the medication or a placebo. (correct answer)
- Match each male subject with a female subject of a similar age, creating pairs that are then randomly assigned to the two groups.
- Assign all male subjects to receive the medication and all female subjects to receive the placebo to isolate the effect within each sex.
Explanation: This strategy is known as blocking. By separating subjects into blocks based on a known potential confounder (sex) and then randomizing within each block, the researcher ensures that the confounding variable is balanced between the treatment and control groups. This is a powerful design-based method of control. Choice A describes an analysis-based control, which is valid but generally less preferred than a design-based control like blocking when the confounder is known beforehand. Choice C describes an incorrect application of matching. Choice D is not a randomized experiment and would completely confound the effect of the drug with the effect of sex.
Question 5
An ecologist observes that in a certain forest, trees with a higher prevalence of a specific fungus also have a lower growth rate. The ecologist considers that 'sunlight exposure' could be a confounding variable. Which of the following scenarios would best support the claim that sunlight exposure is a confounder?
- The fungus grows best in shady conditions, and trees that receive less sunlight have lower growth rates. (correct answer)
- The fungus drains nutrients from the tree, which directly causes the lower growth rate, regardless of sunlight.
- The effect of the fungus on tree growth is much more severe for trees in low-sunlight conditions.
- Trees with low growth rates are more susceptible to fungal infections, but sunlight does not affect growth.
Explanation: For sunlight exposure to be a confounder, it must be associated with both the explanatory variable (fungus) and the response variable (growth rate). Choice A establishes both links: the fungus is more prevalent in low sunlight (association with explanatory), and low sunlight leads to lower growth (association with response). This provides an alternative explanation for the fungus-growth correlation. Choice B argues for a direct causal link, ignoring confounding. Choice C describes interaction (effect modification). Choice D describes reverse causality and breaks the link between sunlight and growth rate.
Question 6
Researchers conduct a cross-sectional study and find a positive association between the number of fitness apps on a person's smartphone and their body mass index (BMI). They cannot conclude that having more fitness apps causes a higher BMI. Which of the following is the most plausible confounding variable that could explain this counterintuitive result?
- The brand of smartphone used by the individual.
- The subscription cost of the fitness apps.
- An individual's concern about their current weight. (correct answer)
- The amount of time spent engaging with the fitness apps.
Explanation: An individual's pre-existing concern about their weight is a plausible confounder. A person who is concerned about being overweight (and thus may have a higher BMI) might be more likely to download multiple fitness apps in an attempt to address the issue. This concern is associated with both the explanatory variable (number of apps) and the response variable (BMI), creating a spurious association. Smartphone brand (A) and app cost (B) are less likely to be strongly associated with both variables. Time spent using the apps (D) is related to the explanatory variable but is not the underlying common cause.
Question 7
An educational software company wants to test if its new math game improves student test scores more than its old game. They will conduct an experiment in a large school district. They know that math achievement levels vary substantially between elementary, middle, and high school students. Which design would be most effective for controlling for the pre-existing differences among school levels?
- Separate students into three blocks (elementary, middle, high school) and then randomly assign students within each block to either the new or old game. (correct answer)
- Conduct the study only with high school students to hold the school level constant, then generalize the results.
- Allow teachers to choose which game they think is best for their class, ensuring equal numbers of students use each game.
- Randomly assign half of all students in the district to the new game and half to the old game, and then use regression to adjust for school level.
Explanation: When you encounter experimental design questions involving known confounding variables, think about how to control for factors that could mask or distort your treatment effect. Here, school level creates systematic differences in math ability that could overwhelm any game effects.
Option A is correct because it uses a randomized block design. By creating blocks (elementary, middle, high school) and randomly assigning students within each block to treatments, you ensure fair comparisons at each level. Any difference between games within a block can't be attributed to developmental differences, since both treatment groups have the same school level. This design also lets you detect whether the game effect varies by school level.
Option B limits the study to only high school students, which does control for school level but severely restricts generalizability. The company needs to know if their game works across all levels, not just high schoolers.
Option C eliminates randomization entirely by letting teachers self-select treatments. This introduces selection bias—teachers might choose games based on their class's ability level or their own preferences, confounding the results.
Option D relies on statistical adjustment after the fact. While regression can help control for known variables, it's less reliable than experimental control through blocking. Plus, with substantial differences between school levels, the randomization might create unbalanced groups that are hard to adjust for properly.
Study tip: When you see experiments with known confounding variables, look for randomized block designs first. They're more powerful than post-hoc statistical adjustments and more generalizable than restriction strategies.
Question 8
In an observational study, researchers find that patients prescribed a new, powerful anti-inflammatory drug have worse long-term outcomes than patients prescribed an older, standard drug. The researchers are hesitant to conclude the new drug is harmful due to 'confounding by indication.' What does this specific type of confounding imply in this context?
- Doctors' indications of which drug to prescribe were not randomly assigned, violating experimental principles.
- Patients with more severe disease and a worse prognosis are preferentially prescribed the new, more powerful drug. (correct answer)
- The new drug has side effects that indicate the presence of other health problems which are the true cause of poor outcomes.
- Patients who were indicated for the new drug were also more likely to have poor health behaviors, like smoking.
Explanation: 'Confounding by indication' (or 'confounding by severity') is a specific and common problem in observational studies of treatments. It occurs when the clinical reasons (the indication) for prescribing a particular treatment are also risk factors for the outcome of interest. In this case, doctors are likely giving the stronger new drug to the sickest patients, and their poor outcomes are a result of their underlying disease severity, not the drug itself. While A and D describe general issues in observational studies, B is the specific definition of confounding by indication.
Question 9
A public health study finds that individuals who drink more coffee have a higher incidence of coronary heart disease. A researcher suspects that smoking status is a confounding variable in this relationship. For smoking to be a true confounder, which of the following set of conditions must be met?
- The effect of coffee consumption on heart disease risk must be significantly different for smokers than for non-smokers.
- Coffee consumption must be a direct cause of smoking, which in turn is a direct cause of heart disease.
- Smokers must be more likely to drink coffee than non-smokers, and smoking must be an independent risk factor for heart disease. (correct answer)
- Smoking must be a risk factor for heart disease, but there must be no association between coffee consumption and smoking status.
Explanation: For a variable to be a confounder, it must be associated with both the explanatory variable (coffee drinking) and the response variable (heart disease). Choice C correctly states these two conditions: smokers are more likely to drink coffee (association with explanatory) and smoking is a risk factor for heart disease (association with response). This creates a potential alternative explanation for the observed association between coffee and heart disease. Choice A describes interaction (effect modification). Choice B describes mediation. Choice D breaks one of the necessary conditions for confounding (the association with the explanatory variable).
Question 10
A city's department of public health observes that neighborhoods with more public parks per capita have lower rates of obesity. They acknowledge that neighborhood wealth is a potential confounder. To address this, they propose a study where they identify pairs of neighborhoods across the city, such that both neighborhoods in a pair have very similar median household incomes, but one has a high number of parks and the other has a low number. They will then compare obesity rates within these pairs. This study design primarily relies on what control strategy?
- Random assignment
- Matching (correct answer)
- Blocking
- Statistical adjustment
Explanation: This is a classic example of a matched-pair design for an observational study. By selecting pairs of neighborhoods that are similar on the key confounding variable (wealth), the researchers are attempting to control for its effect. This is distinct from blocking (B), which is used in experiments before random assignment. Since this is an observational study, random assignment (A) is not used. Statistical adjustment (D) is an analytical technique, whereas matching is a study design strategy.
Question 11
A pharmaceutical company conducts a clinical trial for a new antidepressant. To ensure the results are widely applicable, they recruit participants from diverse geographic and demographic backgrounds. To establish a causal link between the drug and symptom improvement, they use a computer to assign each participant to either the new drug or a placebo. Which statement best distinguishes the primary roles of these two procedures?
- The diverse recruitment controls for confounding variables, while random assignment ensures the sample is representative.
- Both procedures are primarily designed to reduce sampling error and increase the statistical power of the study.
- The diverse recruitment addresses external validity, while random assignment addresses internal validity. (correct answer)
- The diverse recruitment minimizes measurement bias, while random assignment minimizes selection bias.
Explanation: These two procedures address two different aspects of validity. The diverse recruitment aims to create a sample that reflects the broader population, which is a matter of generalizability, or external validity. The random assignment to treatment groups is designed to create comparable groups and control for confounding, which is essential for making a valid causal inference about the drug's effect within the study sample, i.e., internal validity. Choice A reverses the roles. Choice B is too general. Choice D uses terminology incorrectly; random assignment deals with allocation bias, not selection bias in the sampling sense.
Question 12
An epidemiologist is designing a case-control study to investigate the association between exposure to an industrial solvent and a rare form of kidney disease. To control for confounding by age and smoking status, the researcher intends to select a control group of individuals without the disease. Which method of selecting the control group would be the most rigorous design-based strategy for controlling for these specific confounders?
- Select a control group with the same overall proportion of smokers and the same average age as the case group.
- Select a completely random sample of the general population to serve as the control group.
- Select a control group and then use a multiple regression model to adjust for differences in age and smoking during the analysis phase.
- For each individual with the disease (a case), select one or more individuals without the disease who are in the same age bracket and have the same smoking history. (correct answer)
Explanation: When you encounter case-control study design questions, focus on how the study controls for confounding variables. Confounders are variables that affect both the exposure and outcome, potentially creating false associations. The most rigorous design-based approach controls for confounders during subject selection rather than relying solely on statistical adjustment afterward.
Answer D represents matching, the gold standard for confounder control in case-control studies. By selecting controls who share the same age bracket and smoking history as each case, you ensure these potential confounders are distributed equally between groups. This eliminates their ability to distort the relationship between solvent exposure and kidney disease, since matched variables cannot explain differences in disease rates between cases and controls.
Answer A is problematic because matching only on overall proportions allows for imbalanced distributions within subgroups. You might have the same average age overall, but cases could be clustered in high-risk age groups while controls aren't.
Answer B fails to control for confounders at all. A random population sample will likely have different age and smoking distributions than your disease cases, making it impossible to isolate the effect of solvent exposure.
Answer C relies on analytical control rather than design control. While regression can adjust for confounders, it's less rigorous than preventing the confounding through careful subject selection and assumes you've measured and modeled all relevant confounders correctly.
Study tip: Remember that matching in case-control studies is about selecting controls who are similar to cases on potential confounders, not on the exposure or outcome of interest.
Question 13
A researcher notes that cities with a higher number of hospitals per capita also have a higher mortality rate. The association is statistically significant. A sociologist suggests that a city's population age structure is a key unmeasured variable that could explain this finding. In this context, the age structure of the population is best described as what?
- A mediating variable, because hospitals cause the population to age, which in turn leads to higher mortality.
- An interacting variable, because the effect of hospitals on mortality might be different in younger vs. older cities.
- A response variable, since the number of hospitals is designed to respond to the population's age structure.
- A confounding variable, because a higher proportion of older residents likely leads to both more hospitals and a higher mortality rate. (correct answer)
Explanation: When you encounter a research scenario where two variables are correlated but a third variable might explain both, you're dealing with questions about confounding variables. This is a fundamental concept in understanding causation versus correlation in statistics.
The key insight here is recognizing the causal pathway. Cities with older populations naturally need more hospitals to serve their residents' greater healthcare needs. Simultaneously, older populations have higher mortality rates regardless of hospital availability. This creates a spurious correlation between hospitals and mortality that disappears once you account for age structure.
Answer D correctly identifies age structure as a confounding variable because it influences both the predictor (hospitals per capita) and the outcome (mortality rate), creating a non-causal association between them.
Answer A incorrectly suggests hospitals cause population aging, which reverses the actual causal direction. Hospitals don't make populations older; older populations require more hospitals.
Answer B mischaracterizes this as an interaction effect. While hospital effectiveness might vary by population age, the sociologist is suggesting age structure explains the entire association, not that it modifies the hospital-mortality relationship.
Answer C incorrectly labels age structure as the response variable. The response variable is what you're trying to explain (mortality rate). Age structure is an explanatory factor, not the outcome being measured.
Remember: confounding variables are "lurking" factors that influence both your predictor and outcome variables. Always ask yourself whether an unmeasured variable could be driving the relationship you observe, rather than assuming direct causation.
Question 14
A key advantage of a well-designed randomized controlled experiment over an observational study is that random assignment of a treatment is expected to:
- ensure that the study participants are a representative sample of the population, allowing for generalizability.
- eliminate all sources of bias, ensuring the observed treatment effect is exactly the true effect.
- hold all variables constant between the treatment and control groups except for the explanatory variable.
- create approximate balance on all potential confounding variables, both measured and unmeasured, between groups. (correct answer)
Explanation: This question tests your understanding of what randomization achieves in experimental design versus observational studies. The key insight is recognizing what random assignment can and cannot accomplish.
Random assignment's primary strength is creating balance between treatment groups. When you randomly assign participants to treatment and control groups, you're essentially letting chance distribute all variables—both those you've measured and those you haven't even thought of—approximately equally between groups. This includes potential confounders like socioeconomic status, genetic factors, lifestyle habits, and countless unmeasured variables that could influence your outcome. This balance is what makes randomized experiments so powerful for establishing causation.
Option A confuses random assignment with random sampling. Random assignment divides your existing participants into groups; random sampling determines who participates in your study. You need random sampling (not assignment) for generalizability.
Option B is too absolute. Randomization doesn't eliminate ALL bias—measurement bias, selection bias in recruitment, and other sources can still exist. It specifically addresses confounding bias by balancing confounders between groups.
Option C suggests randomization "holds variables constant," but that's incorrect. Randomization doesn't control or hold variables constant; it distributes them randomly so they vary similarly across groups. If you actually held all other variables constant, you'd have perfect experimental control, not randomization.
Remember: randomization's magic is in creating comparable groups by balancing confounders through chance, not by eliminating or controlling variables directly. This is why randomized controlled trials are considered the gold standard for causal inference.
Question 15
An urban planning group finds a negative correlation between a neighborhood's average commute time to work and residents' self-reported life satisfaction. They suspect that neighborhood income is a major confounder. They propose an analysis where they first group neighborhoods into three categories: low-income, middle-income, and high-income. They then calculate the correlation between commute time and life satisfaction separately within each of these three income categories. This analytical strategy is known as:
- Stratification (correct answer)
- Blocking
- Matching
- Instrumental variable analysis
Explanation: When researchers suspect that a third variable might be confounding the relationship between two variables of interest, they need strategies to control for that confounder. In this case, the planners worry that neighborhood income affects both commute times and life satisfaction, potentially creating a spurious correlation.
The strategy described here is stratification (A). This involves dividing the sample into homogeneous subgroups based on the confounding variable (income level), then analyzing the relationship of interest separately within each stratum. By examining the commute time-satisfaction correlation within low-income, middle-income, and high-income neighborhoods separately, researchers can see if the relationship holds when income is held relatively constant.
Blocking (B) is used in experimental design, where you group similar experimental units together before randomly assigning treatments. This is observational data, not an experiment with assigned treatments.
Matching (C) involves pairing individual observations that are similar on confounding variables, then comparing outcomes within these matched pairs. The planners aren't matching specific neighborhoods; they're creating broad income categories.
Instrumental variable analysis (D) uses a third variable that affects the exposure variable but only affects the outcome through the exposure. This sophisticated technique requires finding an appropriate instrument, which isn't described here.
Remember: stratification creates subgroups and analyzes relationships separately within each group. When you see researchers dividing their sample by a potential confounder and repeating their analysis in each subset, think stratification.
Question 16
Researchers conducted a large, well-designed randomized, double-blind, placebo-controlled trial to test a new drug for migraines. The results showed no statistically significant difference in headache frequency between the drug group and the placebo group. Assuming the study had adequate statistical power, what is the most appropriate conclusion regarding the drug's efficacy?
- The study provides strong evidence that the drug lacks efficacy, as the design controlled for known and unknown confounders. (correct answer)
- The study's results are likely biased because the participants were not a random sample of the general population.
- The study is inconclusive because an unmeasured variable may have confounded the relationship.
- The lack of effect is likely due to the placebo effect, which was not adequately controlled for by the study design.
Explanation: When evaluating clinical trial results, the study design determines how confidently you can interpret the findings. A randomized, double-blind, placebo-controlled trial represents the gold standard because it systematically eliminates bias and confounding variables.
The correct answer is A because this study design provides the strongest possible evidence about drug efficacy. Randomization ensures that known and unknown confounding variables are equally distributed between groups, while double-blinding prevents both participant and researcher bias from influencing results. When such a well-designed study with adequate statistical power shows no significant difference, it provides strong evidence that the drug truly lacks efficacy.
Answer B incorrectly focuses on external validity (generalizability) rather than internal validity. While participants weren't randomly sampled from the general population, this doesn't bias the comparison between drug and placebo groups within the study.
Answer C misunderstands how randomization works. The beauty of randomized controlled trials is that they control for unmeasured confounders by distributing them equally between groups, making this concern irrelevant.
Answer D reveals a fundamental misunderstanding of study design. The placebo group serves specifically to control for placebo effects - that's the entire point of having it. Both groups experience any placebo effect equally, so it cannot explain the lack of difference between groups.
Remember: When a well-designed RCT with adequate power shows no effect, trust the result. The rigorous design methodology is specifically built to give you confidence in null findings, not just positive ones.
Question 17
A public health study finds that individuals who drink more coffee have a higher incidence of coronary heart disease. A researcher suspects that smoking status is a confounding variable in this relationship. For smoking to be a true confounder, which of the following set of conditions must be met?
- The effect of coffee consumption on heart disease risk must be significantly different for smokers than for non-smokers.
- Coffee consumption must be a direct cause of smoking, which in turn is a direct cause of heart disease.
- Smokers must be more likely to drink coffee than non-smokers, and smoking must be an independent risk factor for heart disease. (correct answer)
- Smoking must be a risk factor for heart disease, but there must be no association between coffee consumption and smoking status.
Explanation: For a variable to be a confounder, it must be associated with both the explanatory variable (coffee drinking) and the response variable (heart disease). Choice C correctly states these two conditions: smokers are more likely to drink coffee (association with explanatory) and smoking is a risk factor for heart disease (association with response). This creates a potential alternative explanation for the observed association between coffee and heart disease. Choice A describes interaction (effect modification). Choice B describes mediation. Choice D breaks one of the necessary conditions for confounding (the association with the explanatory variable).
Question 18
Researchers conduct a cross-sectional study and find a positive association between the number of fitness apps on a person's smartphone and their body mass index (BMI). They cannot conclude that having more fitness apps causes a higher BMI. Which of the following is the most plausible confounding variable that could explain this counterintuitive result?
- The brand of smartphone used by the individual.
- The subscription cost of the fitness apps.
- An individual's concern about their current weight. (correct answer)
- The amount of time spent engaging with the fitness apps.
Explanation: An individual's pre-existing concern about their weight is a plausible confounder. A person who is concerned about being overweight (and thus may have a higher BMI) might be more likely to download multiple fitness apps in an attempt to address the issue. This concern is associated with both the explanatory variable (number of apps) and the response variable (BMI), creating a spurious association. Smartphone brand (A) and app cost (B) are less likely to be strongly associated with both variables. Time spent using the apps (D) is related to the explanatory variable but is not the underlying common cause.
Question 19
A pharmaceutical company conducts a clinical trial for a new antidepressant. To ensure the results are widely applicable, they recruit participants from diverse geographic and demographic backgrounds. To establish a causal link between the drug and symptom improvement, they use a computer to assign each participant to either the new drug or a placebo. Which statement best distinguishes the primary roles of these two procedures?
- The diverse recruitment controls for confounding variables, while random assignment ensures the sample is representative.
- Both procedures are primarily designed to reduce sampling error and increase the statistical power of the study.
- The diverse recruitment addresses external validity, while random assignment addresses internal validity. (correct answer)
- The diverse recruitment minimizes measurement bias, while random assignment minimizes selection bias.
Explanation: These two procedures address two different aspects of validity. The diverse recruitment aims to create a sample that reflects the broader population, which is a matter of generalizability, or external validity. The random assignment to treatment groups is designed to create comparable groups and control for confounding, which is essential for making a valid causal inference about the drug's effect within the study sample, i.e., internal validity. Choice A reverses the roles. Choice B is too general. Choice D uses terminology incorrectly; random assignment deals with allocation bias, not selection bias in the sampling sense.
Question 20
An epidemiologist is designing a case-control study to investigate the association between exposure to an industrial solvent and a rare form of kidney disease. To control for confounding by age and smoking status, the researcher intends to select a control group of individuals without the disease. Which method of selecting the control group would be the most rigorous design-based strategy for controlling for these specific confounders?
- Select a control group with the same overall proportion of smokers and the same average age as the case group.
- Select a completely random sample of the general population to serve as the control group.
- Select a control group and then use a multiple regression model to adjust for differences in age and smoking during the analysis phase.
- For each individual with the disease (a case), select one or more individuals without the disease who are in the same age bracket and have the same smoking history. (correct answer)
Explanation: When you encounter case-control study design questions, focus on how the study controls for confounding variables. Confounders are variables that affect both the exposure and outcome, potentially creating false associations. The most rigorous design-based approach controls for confounders during subject selection rather than relying solely on statistical adjustment afterward.
Answer D represents matching, the gold standard for confounder control in case-control studies. By selecting controls who share the same age bracket and smoking history as each case, you ensure these potential confounders are distributed equally between groups. This eliminates their ability to distort the relationship between solvent exposure and kidney disease, since matched variables cannot explain differences in disease rates between cases and controls.
Answer A is problematic because matching only on overall proportions allows for imbalanced distributions within subgroups. You might have the same average age overall, but cases could be clustered in high-risk age groups while controls aren't.
Answer B fails to control for confounders at all. A random population sample will likely have different age and smoking distributions than your disease cases, making it impossible to isolate the effect of solvent exposure.
Answer C relies on analytical control rather than design control. While regression can adjust for confounders, it's less rigorous than preventing the confounding through careful subject selection and assumes you've measured and modeled all relevant confounders correctly.
Study tip: Remember that matching in case-control studies is about selecting controls who are similar to cases on potential confounders, not on the exposure or outcome of interest.