Math 1 Quiz: Scatterplots And Association
10 questions · exam conditions
0:00
Scatterplots And AssociationQuestion 1 of 10

A scatterplot of study hours versus exam scores shows most points clustered along a positive linear trend, but three students who studied 8-10 hours scored much lower than expected. How should these points be characterized in terms of their impact on association measures?

They are outliers that should be removed because they represent data collection errors or anomalies
They are influential points that strengthen the positive association by increasing the range of data
They represent natural variation and have minimal impact on the overall association pattern
They are outliers that weaken the association and increase variability around the trend line
← Back to quizzes

Math 1 Quiz

Math 1 Quiz: Scatterplots And Association

Practice Scatterplots And Association in Math 1 with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Scatterplots And Association, giving you a quick way to practice the rules, question types, and explanations that matter most for Math 1.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A scatterplot of study hours versus exam scores shows most points clustered along a positive linear trend, but three students who studied 8-10 hours scored much lower than expected. How should these points be characterized in terms of their impact on association measures?

  1. They are outliers that should be removed because they represent data collection errors or anomalies
  2. They are influential points that strengthen the positive association by increasing the range of data
  3. They represent natural variation and have minimal impact on the overall association pattern
  4. They are outliers that weaken the association and increase variability around the trend line (correct answer)
Explanation: When analyzing scatterplots, you need to understand how outliers affect measures of association like correlation coefficients and regression lines. Outliers are data points that fall far from the general pattern, and their impact depends on where they deviate from the trend. In this scenario, three students studied 8-10 hours (high study time) but scored much lower than the positive trend would predict. These points pull the trend line downward, reducing the slope and weakening the correlation coefficient. They also increase the scatter of points around the line, making the relationship appear less predictable. This perfectly describes answer D - these outliers weaken the association and increase variability. Answer A is wrong because you shouldn't automatically assume outliers represent errors. These could be legitimate data points - perhaps these students had test anxiety or other factors affecting performance. Answer B incorrectly suggests these points strengthen the association. While they do increase the range of study hours, they actually weaken the linear relationship by deviating from the expected pattern. Answer C underestimates their impact - points that fall far below a strong linear trend are not "minimal impact" variations but genuine outliers that meaningfully affect correlation measures. Remember this pattern: outliers that fall off the main trend line (whether above or below) generally weaken correlations and increase variability. The key is identifying where points deviate from the expected pattern, not just whether they're at the extremes of your data range.

Question 2

A researcher creates a scatterplot of temperature versus ice cream sales and observes a strong positive linear association (r = 0.87). However, when examining the residuals plot, the researcher notices the residuals show a clear curved pattern. What does this suggest about the original interpretation?

  1. The correlation coefficient is incorrect and should be recalculated using proper statistical methods
  2. The relationship is actually nonlinear, and a curved model would better fit the data pattern (correct answer)
  3. The strong correlation is valid, but outliers are present that need to be removed from analysis
  4. The linear model is appropriate, but heteroscedasticity is affecting the reliability of predictions
Explanation: A curved pattern in residuals indicates that the linear model is not capturing the true relationship adequately - the relationship is nonlinear. The correlation coefficient r = 0.87 measures linear association, but when residuals show curvature, it suggests a curved model would fit better. Choice A is wrong because the correlation calculation isn't necessarily incorrect. Choice C misinterprets what curved residuals indicate. Choice D confuses heteroscedasticity (changing variability) with nonlinearity.

Question 3

Two researchers create scatterplots for the same dataset but reach different conclusions about the strength of association. Researcher A reports a strong positive association, while Researcher B reports a weak positive association. Assuming both are competent, what most likely explains this discrepancy?

  1. One researcher incorrectly calculated the correlation coefficient due to computational errors in the analysis
  2. The researchers used different scales on their axes, making the same data appear different visually
  3. One researcher included outliers in the analysis while the other excluded them from consideration (correct answer)
  4. The researchers focused on different subsets of the data or used different criteria for grouping
Explanation: Outliers can dramatically affect the apparent strength of association. Including or excluding them can change a relationship from strong to weak or vice versa. Choice A assumes incompetence, which contradicts the problem. Choice B is incorrect because axis scaling affects visual appearance but shouldn't change professional assessment of association strength. Choice D assumes they analyzed different data, but the problem states it's the same dataset.

Question 4

Two variables show a correlation coefficient of r = -0.23. A student examines the scatterplot and concludes there is "no meaningful relationship" between the variables. Under what circumstances would this conclusion be most problematic?

  1. When the sample size is very large, making even small correlations statistically significant and meaningful
  2. When the relationship is actually nonlinear, making the correlation coefficient an inappropriate measure (correct answer)
  3. When outliers are present that are masking a stronger underlying linear relationship in the data
  4. When the variables are measured on different scales, affecting the magnitude of correlation
Explanation: A correlation coefficient near zero doesn't necessarily mean no relationship exists - it means no linear relationship. If the true relationship is curved (like U-shaped or S-shaped), r could be near zero even with a strong nonlinear association. Choice A is partially true but less problematic since weak relationships remain weak regardless of significance. Choice C could affect correlation but outliers typically don't completely mask relationships. Choice D is incorrect because correlation is scale-independent.

Question 5

A scatterplot of height versus shoe size shows a strong positive linear association for a sample of adults. A student concludes that increasing someone's shoe size will cause them to become taller. What is the primary flaw in this reasoning?

  1. The student confused correlation with causation and ignored potential confounding variables like genetics (correct answer)
  2. The student failed to consider that the association might not be linear across all ranges of data
  3. The student incorrectly assumed that strong associations always indicate perfect predictive relationships
  4. The student overlooked the possibility that outliers might be artificially inflating the correlation coefficient
Explanation: The fundamental error is inferring causation from correlation. Height and shoe size are both influenced by underlying factors like genetics and age, but neither directly causes the other. Choice B is incorrect because the linearity isn't the main issue with the causal claim. Choice C is wrong because the student's error isn't about prediction accuracy. Choice D is irrelevant since the problem states the association is genuinely strong.

Question 6

A scatterplot shows the relationship between years of education and annual income. The pattern shows a strong positive association for incomes up to $80,000, but above this threshold, the relationship becomes much weaker with high variability. What type of association does this describe?

  1. Linear association with heteroscedasticity affecting the higher income ranges specifically
  2. Nonlinear association that might be better described as logarithmic or exponential in nature
  3. Piecewise linear association with different slopes in different ranges of the explanatory variable (correct answer)
  4. Weak overall association that appears strong only due to clustering in lower income ranges
Explanation: The description indicates different relationship patterns in different ranges of income - strong linear up to $80,000, then weak above that threshold. This suggests a piecewise linear relationship with a break point. Choice A only addresses variability, not the change in association strength. Choice B suggests a smooth curve, but the description implies a distinct change at a threshold. Choice D incorrectly characterizes the overall pattern.

Question 7

A marketing analyst creates a scatterplot showing the relationship between advertising expenditure (in thousands of dollars) and sales revenue (in thousands of dollars) for 25 products. The plot shows a strong positive linear association, but when examining the data more closely, the analyst notices that the relationship only holds true for advertising expenditures below $50,000. Above this threshold, increased advertising shows no clear relationship with sales. What type of association pattern is this?

  1. Nonlinear association that requires a curved model to describe the complete relationship
  2. Linear association with outliers that should be removed to improve the model fit
  3. Piecewise linear association with different relationships in different ranges of the data (correct answer)
  4. Weak linear association that appears strong due to the presence of lurking variables
Explanation: The correct answer is C. This describes a piecewise linear relationship where the pattern changes at a threshold ($50,000). Below this point, there's a strong positive linear relationship, but above it, there's no clear relationship. A is incorrect because each piece is linear, not curved. B is incorrect because the high-expenditure points aren't outliers but represent a different relationship regime. D is incorrect because the issue isn't about lurking variables but about different relationship patterns in different ranges.

Question 8

A researcher studying the relationship between temperature and plant growth observes that as temperature increases from 10°C to 25°C, plant height increases steadily. However, as temperature continues to increase from 25°C to 40°C, plant height begins to decrease. When creating a scatterplot of this data, what would be the most accurate description of the association?

  1. Strong positive linear association across the entire temperature range with some random variation
  2. No association since the positive and negative portions cancel each other out completely
  3. Strong nonlinear association that cannot be described as simply positive or negative (correct answer)
  4. Two separate linear associations: positive below 25°C and negative above 25°C with moderate strength
Explanation: The correct answer is C. The relationship described follows a curved (likely quadratic) pattern that peaks around 25°C. This is a strong nonlinear association that cannot be characterized as simply positive or negative since it changes direction. A is incorrect because the relationship is not linear. B is incorrect because there is a clear, strong relationship, just not linear. D is incorrect because while it correctly identifies the changing relationship, the overall pattern is better described as nonlinear rather than piecewise linear.

Question 9

A sports analyst examines the relationship between a basketball player's height and their free-throw percentage for 30 professional players. The scatterplot reveals virtually no pattern, with a correlation coefficient near zero. However, when the analyst separates the data by position (guards vs. centers), each group shows a moderate negative association. What phenomenon does this illustrate?

  1. Simpson's paradox, where subgroup patterns differ from the overall pattern due to confounding variables (correct answer)
  2. Regression toward the mean, where extreme values in height produce average free-throw percentages
  3. Sampling bias, where the data collection method favored certain types of players over others
  4. Measurement error, where imprecise height measurements obscured the true underlying relationship
Explanation: The correct answer is A. This is a classic example of Simpson's paradox, where the overall relationship (no correlation) is different from the relationships within subgroups (negative correlations for both guards and centers). Position acts as a confounding variable that affects both height and shooting ability. B is incorrect because this isn't about extreme values regressing. C is incorrect because there's no indication of sampling bias. D is incorrect because the issue isn't measurement precision but rather the need to account for player position.

Question 10

A researcher finds that hours spent on social media and GPA show a correlation of r = -0.31 in a sample of 200 students. When the data is plotted, the scatterplot appears to show two distinct groups with different patterns. What does this suggest about interpreting the correlation?

  1. The correlation accurately represents the relationship, but the two groups indicate different populations
  2. The correlation should be recalculated separately to determine which group shows stronger association
  3. The correlation is too weak to be meaningful, and the grouping confirms no association exists
  4. The overall correlation may be misleading because subgroups might have different relationships (correct answer)
Explanation: When you encounter correlation problems involving scatterplots with distinct groupings, you're dealing with a classic statistical phenomenon called Simpson's Paradox or subgroup effects. The key insight is that an overall correlation can mask or misrepresent what's actually happening within different subgroups of your data. Here, the correlation of r = -0.31 suggests a moderate negative relationship between social media use and GPA across all 200 students. However, the presence of two distinct groups in the scatterplot is a red flag that this overall correlation might not tell the complete story. The correct answer is D because when data clusters into distinct groups, the overall correlation can be misleading. For example, one group might be graduate students (high GPA, low social media use) and another might be freshmen (lower GPA, high social media use). Within each group, the actual relationship between social media and GPA might be different—or even nonexistent. Answer A incorrectly assumes the overall correlation remains valid despite the grouping. Answer B misses the point by focusing on which group has stronger association rather than recognizing that the overall correlation is problematic. Answer C wrongly dismisses a correlation of -0.31 as meaningless—this is actually a moderate correlation that would typically be considered significant with 200 subjects. Remember: whenever you see distinct clustering or grouping in correlation problems, immediately question whether the overall correlation accurately represents the relationships within subgroups. Always examine scatterplots for patterns that might invalidate summary statistics.