Math 2 Quiz: Interpreting Correlation
20 questions · exam conditions
0:00
Interpreting CorrelationQuestion 1 of 20

A longitudinal study tracks the same individuals over 10 years and finds r=0.52r = 0.52 between exercise frequency and cardiovascular health scores. Compared to a cross-sectional study with the same correlation, what additional insight does the longitudinal design provide regarding causation?

The longitudinal design provides stronger evidence for causation by establishing temporal order, though confounding variables may still exist.
The longitudinal design proves causation because it eliminates all potential confounding variables through repeated measurements over time.
The longitudinal design produces more reliable correlations because it uses more data points from the same subjects over time.
The longitudinal design eliminates the need to consider causation since correlations from repeated measures automatically indicate causal relationships.
← Back to quizzes

Math 2 Quiz

Math 2 Quiz: Interpreting Correlation

Practice Interpreting Correlation in Math 2 with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Interpreting Correlation, giving you a quick way to practice the rules, question types, and explanations that matter most for Math 2.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A longitudinal study tracks the same individuals over 10 years and finds r=0.52r = 0.52 between exercise frequency and cardiovascular health scores. Compared to a cross-sectional study with the same correlation, what additional insight does the longitudinal design provide regarding causation?

  1. The longitudinal design provides stronger evidence for causation by establishing temporal order, though confounding variables may still exist. (correct answer)
  2. The longitudinal design proves causation because it eliminates all potential confounding variables through repeated measurements over time.
  3. The longitudinal design produces more reliable correlations because it uses more data points from the same subjects over time.
  4. The longitudinal design eliminates the need to consider causation since correlations from repeated measures automatically indicate causal relationships.
Explanation: When you encounter questions about research design and causation, remember that establishing causality requires more than just correlation - you need to consider timing, confounding variables, and the strength of evidence each design provides. Longitudinal studies track the same individuals over time, which gives them a crucial advantage over cross-sectional studies: they can establish temporal order. Since this study measured exercise frequency and cardiovascular health repeatedly over 10 years, researchers can see whether changes in exercise precede changes in cardiovascular health. This temporal sequence is essential for supporting causal claims, since causes must come before effects. However, even with this timing advantage, confounding variables (like genetics, diet, or socioeconomic factors) could still influence both exercise habits and health outcomes. Let's examine why the other options miss the mark. Option B incorrectly claims longitudinal designs eliminate all confounding variables - this is false because repeated measurements don't control for external factors that might influence both variables. Option C focuses on data reliability rather than causation, missing the key distinction between correlation strength and causal inference. Option D makes the dangerous error of suggesting that repeated correlations automatically prove causation, which violates the fundamental principle that correlation never equals causation, regardless of study design. Study tip for research design questions: Remember the hierarchy of causal evidence. Longitudinal studies provide stronger causal evidence than cross-sectional studies because they establish temporal order, but only randomized controlled experiments can truly isolate causal relationships by controlling for confounding variables.

Question 2

A study finds that cities with more bookstores have higher literacy rates (r=0.58r = 0.58). A policy maker concludes that building more bookstores will improve literacy. Which alternative explanation should be considered before accepting this causal claim?

  1. The correlation may be spurious because bookstores and literacy rates both tend to be higher in wealthier, more educated communities. (correct answer)
  2. The correlation is too moderate to support policy decisions since urban planning requires correlations above 0.75 for cost-effectiveness.
  3. The relationship direction may be reversed, with bookstores closing in areas where literacy rates decline over time periods.
  4. The correlation calculation may be incorrect because geographic data typically produces inflated correlation coefficients due to spatial clustering.
Explanation: Economic factors likely influence both bookstore density and literacy rates, creating a confounded relationship. Wealthier areas can support more bookstores and have better educational resources. Choice B arbitrarily sets correlation thresholds. Choice C suggests reverse causation but less plausibly. Choice D incorrectly assumes geographic data automatically inflates correlations.

Question 3

Researchers find r=0.33r = 0.33 between class size and student achievement. They also note that schools with smaller classes tend to have more experienced teachers and better resources. What is the most appropriate interpretation?

  1. The correlation demonstrates that reducing class size will improve achievement, with the additional resources providing extra benefits beyond the main effect.
  2. The correlation proves that class size is less important than teacher quality since the relationship is only moderate in strength.
  3. The correlation is artificially inflated because multiple beneficial factors occur together, requiring statistical adjustment to find the true relationship.
  4. The correlation is confounded by school resources and teacher quality, making it impossible to isolate the effect of class size alone. (correct answer)
Explanation: When interpreting correlations in research, you must always consider confounding variables—factors that are related to both variables being studied and could explain the observed relationship. Here, the correlation of r=0.33r = 0.33 between class size and achievement exists alongside the fact that smaller classes tend to occur with better teachers and resources. This creates a confounding situation where multiple beneficial factors cluster together, making it impossible to determine which factor(s) actually cause improved achievement. The observed correlation could be due to class size, teacher experience, better resources, or any combination of these factors. Answer D correctly identifies this confounding problem—you cannot isolate the pure effect of class size when it's bundled with other advantageous conditions. Answer A incorrectly assumes causation from correlation and treats the additional factors as simply "extra benefits" rather than potential confounders. Answer B misinterprets the correlation strength (0.33 is moderate, not necessarily indicating class size is "less important") and makes an unjustified comparison about relative importance. Answer C suggests the correlation is "artificially inflated," but confounding doesn't necessarily inflate correlations—it makes them uninterpretable regarding causation. Remember: correlation plus confounding equals uncertainty about causation. When you see research describing relationships between variables that tend to occur together, immediately think about confounding variables. The key warning phrase here is "schools with smaller classes tend to have..." which signals that multiple factors are clustered together, preventing clean causal interpretation.

Question 4

A meta-analysis combines 15 studies and finds an overall correlation of r=0.28r = 0.28 between social media use and depression scores, with individual study correlations ranging from 0.15-0.15 to 0.710.71. What does this variability suggest about interpreting the overall correlation?

  1. The wide range invalidates the meta-analysis results because consistent correlations are required for meaningful synthesis across studies.
  2. The overall correlation represents the true population effect since meta-analyses eliminate individual study limitations through statistical averaging.
  3. The variability suggests important moderating factors exist that influence the relationship between social media use and depression. (correct answer)
  4. The positive overall correlation proves causation despite individual study variation since the meta-analytic approach controls for confounding variables.
Explanation: High variability in correlations across studies suggests moderating variables (age, type of social media, depression measurement, etc.) that influence the relationship. Choice A incorrectly dismisses valuable meta-analytic information. Choice B oversimplifies what meta-analyses can conclude. Choice D incorrectly assumes meta-analyses establish causation.

Question 5

A researcher finds r=0.45r = -0.45 between hours of television watching and academic performance. She controls for socioeconomic status and finds the correlation becomes r=0.22r = -0.22. What does this suggest about the original correlation?

  1. The original correlation was more accurate because controlling for variables artificially reduces the true strength of relationships between primary variables.
  2. The difference between correlations is too small to be meaningful since both values indicate the same negative relationship direction.
  3. The dramatic reduction proves that television watching has no real effect on academic performance independent of family income levels.
  4. Socioeconomic status partially confounded the original relationship, and the controlled correlation better isolates the TV-performance association. (correct answer)
Explanation: When you encounter correlation problems involving control variables, focus on how third variables can inflate or mask true relationships between your primary variables of interest. The dramatic drop from r=0.45r = -0.45 to r=0.22r = -0.22 when controlling for socioeconomic status reveals that SES was acting as a confounding variable. Both TV watching and academic performance are related to family income—lower-income families might watch more TV and have fewer educational resources, while higher-income families often watch less TV and have better academic support. This creates an artificially stronger correlation between TV and performance in the original analysis. Option D correctly identifies this pattern: the controlled correlation of r=0.22r = -0.22 better isolates the direct relationship between TV watching and academic performance, separate from socioeconomic influences. Option A misunderstands confounding—controlling for relevant variables actually improves accuracy by removing spurious associations, not reducing "true" relationships. Option B dismisses a meaningful change; dropping from -0.45 to -0.22 represents a substantial reduction in effect size, even though direction remains the same. Option C overstates the conclusion—a correlation of -0.22 still indicates a meaningful negative relationship, just weaker than originally appeared. Remember this pattern: when controlling for a third variable dramatically changes a correlation, that third variable was likely confounding the original relationship. The controlled correlation typically provides a cleaner estimate of the true association between your primary variables.

Question 6

A researcher reports: "The correlation between study time and grades is r=0.40r = 0.40 (p=0.03p = 0.03). Therefore, students should study more to improve their grades." Which limitation of this reasoning is most critical?

  1. The statistical significance only indicates the relationship is unlikely due to chance, not that studying more will cause grade improvements. (correct answer)
  2. The correlation is not strong enough to justify the recommendation since values below 0.5 indicate weak relationships with limited practical value.
  3. The p-value is too high to support the conclusion since correlational studies require significance levels below 0.01 for causal claims.
  4. The correlation could be negative in different populations, making the recommendation potentially harmful without additional demographic analysis.
Explanation: When you encounter research claims that jump from correlation to causation, you need to immediately ask: "Does this statistical relationship actually prove one thing causes another?" The fundamental issue here is that correlation never implies causation, regardless of statistical significance. The p=0.03p = 0.03 tells you there's only a 3% chance this correlation occurred by random chance alone—but it says nothing about whether studying more will actually cause grade improvements. The relationship could exist because better students naturally choose to study more, because both variables are influenced by a third factor (like motivation or socioeconomic status), or because grades influence study habits rather than vice versa. Choice A correctly identifies this core limitation: statistical significance only indicates the relationship is unlikely due to chance, not that increasing study time will cause better grades. Choice B misunderstands correlation strength. While r=0.40r = 0.40 represents a moderate correlation, the threshold of 0.5 for "practical value" is arbitrary—correlations of 0.40 can be meaningful in educational research. Choice C incorrectly suggests correlational studies need p<0.01p < 0.01 for causal claims. The p-value threshold doesn't matter because correlation can never establish causation, regardless of significance level. Choice D focuses on population differences, but this misses the main point. Even if the correlation were consistent across populations, it still wouldn't justify the causal recommendation. Remember: whenever you see researchers making policy recommendations based solely on correlational data, the primary concern should always be the correlation-causation fallacy, not the statistical details.

Question 7

A researcher finds that the correlation coefficient between hours of sleep and test scores is r=0.65r = 0.65. She concludes that getting more sleep causes higher test scores. Which statement best describes this conclusion?

  1. The conclusion is valid because the correlation is positive and moderately strong, indicating a causal relationship.
  2. The conclusion is invalid because correlation does not establish causation, and confounding variables may explain the relationship. (correct answer)
  3. The conclusion is valid because correlations above 0.6 are considered statistically significant for causal inference.
  4. The conclusion is invalid because the correlation is not strong enough to indicate causation, which requires r>0.8r > 0.8.
Explanation: Correlation never establishes causation, regardless of strength. The relationship could be due to confounding variables (like stress levels affecting both sleep and performance) or reverse causation. Choice A incorrectly assumes correlation implies causation. Choice C misunderstands statistical significance. Choice D incorrectly suggests a threshold for causal inference based on correlation strength.

Question 8

Two researchers study the same data and report different correlations: Researcher A finds r=0.73r = 0.73 between income and happiness, while Researcher B finds r=0.31r = 0.31 for the same variables. Investigation reveals Researcher A included outliers while Researcher B removed them. How should these results be interpreted?

  1. Researcher A's result is more valid because removing data points artificially manipulates correlations to support predetermined conclusions.
  2. Researcher B's result is more reliable because outliers always distort correlations and should be systematically removed from analyses.
  3. Both correlations may be valid depending on the research question, but outlier influence suggests the relationship may not be robust. (correct answer)
  4. The dramatic difference proves measurement error occurred, so neither correlation should be trusted without data verification.
Explanation: When outliers dramatically affect correlation, it suggests the relationship may not be consistent across the population. Both approaches can be valid for different research questions. Choice A oversimplifies outlier treatment. Choice B incorrectly states outliers should always be removed. Choice D assumes measurement error without justification.

Question 9

A study finds r=0.15r = 0.15 between daily coffee consumption and productivity scores among office workers. The sample size is 2,000 and the result is statistically significant at p<0.01p < 0.01. What is the most appropriate interpretation?

  1. Coffee consumption significantly improves productivity, so companies should provide free coffee to increase worker output.
  2. There is a weak positive association that is statistically reliable, but practical significance and causation remain unclear. (correct answer)
  3. The correlation is too weak to be meaningful, so the statistical significance must be due to calculation error.
  4. Strong evidence exists for a causal relationship because the large sample size eliminates confounding variables.
Explanation: With a large sample (n=2,000), even weak correlations can be statistically significant. The correlation is weak but reliable, though not necessarily practically meaningful or causal. Choice A incorrectly infers causation and practical significance. Choice C misunderstands how large samples can detect small but real effects. Choice D incorrectly assumes sample size eliminates confounding.

Question 10

A correlation study shows r=0.45r = 0.45 between years of education and annual income. However, when the data is separated by geographic region, the correlations are: Urban areas: r=0.21r = 0.21, Rural areas: r=0.18r = 0.18. What does this suggest?

  1. The original correlation is more reliable because it uses the complete dataset without arbitrary subdivisions that reduce statistical power.
  2. Geographic location may be a confounding variable that inflates the overall correlation between education and income levels. (correct answer)
  3. The weaker correlations within regions prove that education has no real effect on income in specific locations.
  4. Calculation errors occurred because correlations within subgroups should sum to equal the overall correlation coefficient.
Explanation: When overall correlation is stronger than correlations within subgroups, it suggests the grouping variable (geography) may be confounding the relationship. This is Simpson's paradox in reverse. Choice A misses the confounding issue. Choice C overinterprets weaker correlations as no effect. Choice D misunderstands how correlations work with subgroups.

Question 11

A pharmaceutical company reports r=0.42r = 0.42 between drug dosage and symptom improvement in a clinical trial. However, patients were not randomly assigned to dosage levels; doctors chose dosages based on symptom severity. How does this affect interpretation of the correlation?

  1. The correlation is strengthened because doctor expertise ensures optimal dosage-patient matching, reducing random variation in the relationship.
  2. The correlation is unaffected because statistical relationships remain valid regardless of how participants were assigned to treatment conditions.
  3. The correlation becomes more generalizable because real-world prescribing practices are reflected rather than artificial random assignment conditions.
  4. The correlation may be confounded because sicker patients likely received higher doses, potentially masking the true drug effect. (correct answer)
Explanation: When interpreting correlations from observational studies, you must always consider how participants were assigned to different conditions and whether confounding variables might influence the relationship. In this scenario, doctors chose dosages based on symptom severity, creating a systematic bias in treatment assignment. This means sicker patients (with more severe symptoms) likely received higher doses, while patients with milder symptoms received lower doses. This confounding relationship can mask or distort the true effect of the drug. Even if higher doses are more effective, the correlation might appear weaker because the sickest patients receiving high doses may still show less improvement than healthier patients receiving low doses. Answer D correctly identifies this confounding issue. Answer A is wrong because doctor expertise doesn't eliminate confounding bias—it actually creates it by systematically linking dosage to initial symptom severity. Answer B incorrectly assumes that correlation coefficients are interpretation-free; the validity of statistical relationships absolutely depends on study design and potential confounders. Answer C misunderstands generalizability—while this design might reflect real-world practices, it compromises our ability to isolate the drug's true causal effect, making the correlation less meaningful for understanding drug efficacy. Remember this key principle: correlation coefficients are just numbers until you understand the study design behind them. Always ask yourself whether systematic differences between groups (confounding variables) might explain the observed relationship instead of—or in addition to—the variables you're actually measuring.

Question 12

A health researcher reports that cities with more ice cream sales also have higher rates of drowning incidents (r=0.67r = 0.67). A journalist writes that "ice cream consumption increases drowning risk." What is the most likely explanation for this correlation?

  1. The correlation is spurious due to measurement error in either ice cream sales data or drowning incident reporting systems.
  2. A confounding variable, such as temperature or season, influences both ice cream sales and swimming activity levels simultaneously. (correct answer)
  3. The sample size was too small to detect the true relationship, leading to an inflated correlation coefficient estimate.
  4. Reverse causation is occurring, where communities with higher drowning rates actually drive increased ice cream sales through tourism.
Explanation: Choice B correctly identifies a confounding variable (temperature/season) that affects both variables simultaneously. Hot weather increases both ice cream sales and swimming activities (thus drowning risk). Choice A incorrectly attributes the correlation to measurement error. Choice C misunderstands how sample size affects correlation estimates. Choice D proposes an implausible reverse causation mechanism.

Question 13

A longitudinal study tracking the same individuals over 10 years finds a correlation of r=0.41r = -0.41 between years of education completed and number of cigarettes smoked per day. Compared to a cross-sectional study showing the same correlation, what advantage does the longitudinal design provide for interpreting this relationship?

  1. Longitudinal studies automatically establish causation because they follow the same subjects over time, unlike cross-sectional studies.
  2. The extended time frame allows for detection of stronger correlations that would be missed in single-timepoint cross-sectional measurements.
  3. Longitudinal studies produce more accurate correlation coefficients because they eliminate measurement error that occurs in cross-sectional designs.
  4. The temporal sequence in longitudinal data helps rule out reverse causation, though confounding variables may still explain the relationship. (correct answer)
Explanation: When comparing study designs in statistics, you need to understand how the structure of data collection affects what conclusions you can draw from correlational relationships. The key advantage of longitudinal studies is their ability to track changes over time in the same individuals. This temporal dimension helps address one major limitation of correlation: the directionality problem. While correlation doesn't prove causation, longitudinal data can at least help determine which variable came first. If education levels are measured early and smoking behaviors are tracked later, you can rule out the possibility that smoking habits caused lower education levels (reverse causation). However, longitudinal designs still can't eliminate the possibility that some third variable (like socioeconomic status or family background) influences both education and smoking habits. Option A is incorrect because longitudinal studies do not automatically establish causation - they only help with temporal sequence, not with eliminating confounding variables. Option B misunderstands correlation strength; the time frame doesn't inherently make correlations stronger or weaker, and the same correlation coefficient (r=0.41r = -0.41) was found in both designs. Option C wrongly suggests that longitudinal studies eliminate measurement error - both study types can have measurement issues, and following the same people over time can actually introduce new sources of error like participant dropout. Remember: longitudinal designs help with establishing temporal order (which came first), but only experimental designs with random assignment can truly establish causation. Always distinguish between correlation, temporal sequence, and actual causal relationships.

Question 14

A study finds that students who sit in the front rows of classrooms have higher GPAs (r=0.54r = 0.54). The administration decides to assign all struggling students to front-row seats. What assumption is the administration making that may not be justified?

  1. They assume the correlation is strong enough to guarantee individual students will benefit from the seating change intervention.
  2. They assume that front-row seating directly causes higher academic performance, rather than motivated students choosing to sit in front. (correct answer)
  3. They assume the correlation will remain stable over time as more students are moved to front-row positions.
  4. They assume that correlation coefficients above 0.5 indicate relationships that are suitable for policy-making decisions in educational settings.
Explanation: Choice B correctly identifies the causal assumption: that seating location causes performance rather than student motivation affecting both seating choice and performance. Choice A incorrectly focuses on correlation strength rather than causation. Choice C addresses stability but misses the main causation issue. Choice D incorrectly suggests there's a threshold correlation for policy decisions.

Question 15

A company finds that employee satisfaction ratings correlate positively with productivity metrics (r=0.61r = 0.61). Before implementing satisfaction-boosting initiatives, what would be the most important consideration regarding this correlation?

  1. Whether the correlation remains significant when controlling for potential confounding variables like salary, job tenure, and department type. (correct answer)
  2. Whether the sample included enough employees from each department to ensure the correlation is representative of the entire organization.
  3. Whether productivity was measured using objective metrics rather than subjective supervisor ratings that might bias the correlation.
  4. Whether the correlation coefficient is large enough to justify the cost of implementing organization-wide satisfaction improvement programs.
Explanation: Choice A correctly identifies the need to control for confounding variables to better understand the relationship. Higher-paid, longer-tenured employees might be both more satisfied and more productive for reasons unrelated to satisfaction causing productivity. Choice B addresses sampling but not causation. Choice C addresses measurement validity but not the causal interpretation issue. Choice D incorrectly suggests correlation size determines intervention justification.

Question 16

A researcher finds a correlation coefficient of r=0.78r = -0.78 between hours of social media use per day and academic performance scores among high school students. The researcher concludes that reducing social media use will improve academic performance. Which statement best identifies a limitation of this conclusion?

  1. The negative correlation indicates that the relationship is not statistically significant enough to support causal claims.
  2. Correlation does not establish causation; confounding variables like study habits or sleep patterns could explain both low academic performance and high social media use. (correct answer)
  3. The correlation coefficient is too weak to draw any meaningful conclusions about the relationship between the variables.
  4. The sample size was likely too small since correlation coefficients above 0.70.7 in absolute value are only possible with very large datasets.
Explanation: Choice B correctly identifies that correlation does not imply causation. A third variable (like poor time management skills) could cause both high social media use and low academic performance. Choice A incorrectly suggests that negative correlations cannot support statistical significance. Choice C is wrong because r=0.78r = -0.78 represents a strong correlation. Choice D incorrectly claims that strong correlations require large sample sizes.

Question 17

Researchers find correlations between various lifestyle factors and life expectancy in different countries. Which correlation would be most difficult to interpret causally due to potential confounding?

  1. Average daily caloric intake and life expectancy (r=0.32r = -0.32), because the relationship between diet and longevity is well-established in medical literature.
  2. Physical activity levels and life expectancy (r=0.51r = 0.51), because exercise has direct biological mechanisms that are known to affect human health and mortality.
  3. Average hours of sleep per night and life expectancy (r=0.28r = 0.28), because sleep duration has a relatively weak correlation with longevity outcomes.
  4. Healthcare spending per capita and life expectancy (r=0.67r = 0.67), because economic development affects both healthcare investment and numerous other factors influencing population health. (correct answer)
Explanation: When you encounter questions about correlation and causation, the key challenge is identifying which relationships are most vulnerable to confounding variables—hidden factors that influence both variables being studied and create misleading apparent connections. Healthcare spending per capita (option D) presents the most challenging causal interpretation because economic development simultaneously drives both higher healthcare investment and countless other health-promoting factors. Wealthier nations can afford better medical care, but they also have superior sanitation systems, education, nutrition, infrastructure, and environmental regulations. The strong correlation (r=0.67r = 0.67) likely reflects this complex web of interconnected prosperity-related factors rather than healthcare spending alone extending life. Option A is wrong because while caloric intake does have established health effects, the relationship is relatively straightforward with fewer major confounding variables than economic factors. Option B incorrectly suggests that having known biological mechanisms makes confounding less problematic—but confounding depends on hidden variables, not whether direct effects exist. Option C misses the point entirely by focusing on correlation strength (r=0.28r = 0.28) rather than confounding potential; weak correlations can still have clear causal interpretations. The trap here is confusing "established mechanisms" or "correlation strength" with "causal clarity." Strong correlations with known biological pathways can still be heavily confounded, while weak correlations might have cleaner causal stories. Study tip: When evaluating causal interpretation difficulty, look for scenarios involving broad socioeconomic factors like wealth, education, or development—these create the most complex webs of confounding variables that make isolating true causal relationships nearly impossible.

Question 18

Researchers studying ice cream sales and drowning incidents find r=0.78r = 0.78. A news report claims "Ice cream consumption leads to increased drowning risk." Which response best addresses this claim?

  1. The claim is reasonable because the correlation is strong enough to suggest a direct causal mechanism between consumption and drowning.
  2. The claim ignores that both variables likely increase during summer months, making temperature a probable confounding variable. (correct answer)
  3. The claim is invalid because positive correlations between unrelated variables always indicate measurement error or sampling bias.
  4. The claim requires verification through calculating the coefficient of determination to assess the true strength of the relationship.
Explanation: This is a classic example of confounding variables. Both ice cream sales and drowning incidents increase in summer due to hot weather and increased swimming activity. Choice A ignores confounding. Choice C incorrectly assumes all unexpected correlations indicate error. Choice D misses that the issue is confounding, not correlation strength.

Question 19

Two variables have a correlation of r=0.82r = -0.82. Based on the scatter plot, the relationship appears linear with no obvious outliers. What can be reasonably concluded?

  1. Strong evidence exists that one variable directly causes changes in the other variable in the opposite direction.
  2. A strong negative linear association exists, but additional information is needed to determine any causal relationship. (correct answer)
  3. The negative correlation indicates that the data collection method was flawed or biased toward negative outcomes.
  4. Since the correlation is strong, random chance can be ruled out as an explanation for the relationship.
Explanation: A correlation of -0.82 indicates a strong negative linear association, but correlation alone cannot establish causation. Additional experimental or longitudinal data would be needed. Choice A incorrectly infers causation. Choice C misinterprets negative correlation as indicating flawed methodology. Choice D confuses correlation strength with statistical significance testing.

Question 20

Two variables show a correlation of r=0.85r = 0.85. A student claims this proves a strong causal relationship exists. To best evaluate this claim, which additional information would be most critical?

  1. The p-value associated with the correlation coefficient to determine if the relationship is statistically significant at the 0.05 level.
  2. Whether the data collection involved random sampling from the target population to ensure external validity of results.
  3. Whether the study was observational or experimental, and what potential confounding variables might explain the association. (correct answer)
  4. The sample size used to calculate the correlation, since correlations above 0.8 require at least 100 observations to be reliable.
Explanation: Choice C correctly identifies that determining causation requires knowing the study design and considering confounding variables. Even strong correlations in observational studies cannot establish causation. Choice A focuses on statistical significance, which doesn't address causation. Choice B addresses generalizability but not causation. Choice D incorrectly states a rule about sample size and correlation strength.