Biostatistics Quiz: Interpreting Scatterplots
15 questions · exam conditions
0:00
Interpreting ScatterplotsQuestion 1 of 15

A researcher examining the relationship between study hours and exam scores notices that the scatterplot shows a generally positive trend, but the correlation coefficient is only 0.45. What is the most likely explanation for this moderate correlation despite the apparent positive relationship?

The sample size is too small to detect the true strong correlation that exists
There is substantial variability around the regression line, reducing the correlation strength
The relationship is actually curvilinear rather than linear, which correlation doesn't capture well
Measurement error in recording study hours has artificially inflated the correlation coefficient
The presence of outliers is artificially deflating what would otherwise be a strong correlation
← Back to quizzes

Biostatistics Quiz

Biostatistics Quiz: Interpreting Scatterplots

Practice Interpreting Scatterplots in Biostatistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Interpreting Scatterplots, giving you a quick way to practice the rules, question types, and explanations that matter most for Biostatistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A researcher examining the relationship between study hours and exam scores notices that the scatterplot shows a generally positive trend, but the correlation coefficient is only 0.45. What is the most likely explanation for this moderate correlation despite the apparent positive relationship?

  1. The sample size is too small to detect the true strong correlation that exists
  2. There is substantial variability around the regression line, reducing the correlation strength (correct answer)
  3. The relationship is actually curvilinear rather than linear, which correlation doesn't capture well
  4. Measurement error in recording study hours has artificially inflated the correlation coefficient
  5. The presence of outliers is artificially deflating what would otherwise be a strong correlation
Explanation: When you encounter correlation questions in biostatistics, remember that the correlation coefficient measures how tightly data points cluster around a straight line, not just whether a relationship exists. A correlation of 0.45 indicates a moderate positive relationship, which perfectly aligns with seeing a positive trend in the scatterplot. The key insight is that correlation strength depends on how much the data points deviate from the best-fit line. Even with a clear positive relationship, substantial scatter around the regression line will reduce the correlation coefficient. Think of correlation as measuring the "tightness" of the linear relationship - the more spread out the points are vertically from the line, the weaker the correlation becomes. Option A is incorrect because sample size affects our ability to detect correlation, but doesn't explain why a visible positive trend shows only moderate correlation. A small sample would more likely result in an unreliable estimate rather than consistently moderate values. Option C misses the mark because while curvilinear relationships can reduce linear correlation, the question states there's a positive trend visible, suggesting the linear model is reasonable. Option D gets the direction wrong - measurement error typically reduces correlation strength rather than inflating it, and wouldn't explain moderate correlation when a relationship is clearly visible. The correct answer is B because substantial variability around the regression line directly explains how you can have both a visible positive trend and a moderate correlation coefficient. Study tip: Remember that correlation quantifies scatter around the line - perfect correlation means all points lie exactly on the line, while moderate correlation means considerable vertical spread despite a clear trend.

Question 2

A scatterplot displays the relationship between age (X-axis) and blood pressure (Y-axis) for 50 adults. The plot shows an overall positive trend, but the relationship appears to strengthen after age 40. What pattern would you expect to see in the data points?

  1. Points scattered randomly across all age groups with no discernible pattern change
  2. Points showing weak positive association before age 40, then stronger positive association after age 40 (correct answer)
  3. Points showing strong positive association before age 40, then weaker association after age 40
  4. Points showing negative association before age 40, then positive association after age 40
  5. Points clustered tightly around the mean with little variation across age groups
Explanation: When interpreting scatterplots in biostatistics, you need to carefully analyze how the strength and direction of relationships can vary across different ranges of your independent variable. This question tests your ability to visualize how correlation patterns change within subgroups of data. The key information tells you there's an "overall positive trend" that "strengthens after age 40." This means that while blood pressure generally increases with age throughout the entire age range, this relationship becomes more pronounced in older adults. Visually, you'd expect to see data points that show a modest upward trend in younger adults (before 40), followed by points that cluster more tightly around a steeper upward slope after age 40. Answer B correctly describes this pattern: a weak positive association that transitions to a stronger positive association at the age 40 threshold. Answer A is wrong because it describes random scatter with no pattern change, contradicting the described strengthening relationship. Answer C reverses the pattern, suggesting the relationship weakens after 40 rather than strengthens. Answer D describes a direction change from negative to positive, but the question states there's an overall positive trend throughout. Study tip: When analyzing scatterplots with changing patterns, always pay attention to phrases like "strengthens," "weakens," or "changes direction" at specific thresholds. These indicate piecewise relationships where the correlation coefficient differs between subgroups. Practice sketching what these descriptions would look like visually—it helps you quickly eliminate incorrect interpretations on exams.

Question 3

A scatterplot shows the relationship between hours of sleep (X) and cognitive test scores (Y). The plot reveals that both very low sleep (< 4 hours) and very high sleep (> 10 hours) are associated with lower test scores, while moderate sleep (6-8 hours) shows the highest scores. What type of association pattern does this describe?

  1. Strong positive linear association between sleep hours and cognitive performance
  2. Strong negative linear association between sleep hours and cognitive performance
  3. Curvilinear association with an inverted U-shape pattern (correct answer)
  4. No association since the correlation coefficient would be approximately zero
  5. Weak positive association due to high variability in the cognitive test scores
Explanation: When analyzing relationships between variables, you need to distinguish between linear and non-linear patterns. This question tests your ability to recognize curvilinear associations, which are common in biological and psychological research where "too little" and "too much" of something can both be harmful. The described pattern shows cognitive scores peaking at moderate sleep levels (6-8 hours) and declining at both extremes (< 4 hours and > 10 hours). This creates an inverted U-shaped curve, which is a classic curvilinear relationship. Answer C correctly identifies this pattern. Answer A is wrong because a positive linear association would show test scores consistently increasing as sleep hours increase - but here, scores actually decrease after 8 hours. Answer B is incorrect because a negative linear association would show scores consistently decreasing as sleep increases, which doesn't match the pattern where moderate sleep produces the highest scores. Answer D represents a common trap: while the overall correlation coefficient might indeed be near zero due to the U-shape canceling out positive and negative portions, this doesn't mean there's "no association." There's clearly a strong relationship - it's just not linear. Study tip: When you see biological relationships involving optimal ranges or "sweet spots," think curvilinear. Many physiological processes follow inverted U-curves (like the relationship between arousal and performance, or drug dose and effectiveness). Don't let a low correlation coefficient fool you into thinking there's no relationship - always visualize the data pattern first.

Question 4

Two researchers examine the same scatterplot but reach different conclusions. Researcher A claims there is a moderate positive correlation (r = 0.6), while Researcher B claims there is no meaningful linear relationship (r = 0.1). What characteristic of the data would most likely explain this disagreement?

  1. The researchers used different sample sizes for their correlation calculations
  2. One researcher included influential outliers while the other excluded them (correct answer)
  3. The researchers measured the variables using different units of measurement
  4. One researcher incorrectly calculated the correlation using the wrong formula
  5. The data shows a strong curvilinear relationship that one researcher linearized
Explanation: When you encounter questions about correlation discrepancies, focus on what can dramatically alter correlation coefficients while keeping the same basic dataset. The Pearson correlation coefficient measures the strength and direction of linear relationships, but it's highly sensitive to extreme values. The key insight here is that outliers can drastically inflate or deflate correlation values. If the scatterplot contains influential outliers that fall in line with a positive trend, including them would strengthen the correlation (pushing it toward r = 0.6). However, excluding these same outliers might reveal that the remaining data points show little to no linear relationship (r = 0.1). This explains how two researchers examining the "same" scatterplot could reach such different conclusions - they made different decisions about outlier treatment. Let's examine why the other options don't work: (A) is incorrect because if researchers are looking at the same scatterplot, they're working with identical sample sizes - the visual display represents a fixed dataset. (C) is wrong because changing units of measurement doesn't affect correlation coefficients; correlation is unitless and remains constant regardless of whether you measure height in inches or centimeters, for example. (D) is implausible because while calculation errors can occur, the dramatic difference between 0.6 and 0.1 suggests a systematic difference in approach rather than a simple computational mistake. Study tip: Always consider outlier sensitivity when interpreting correlation studies. In biostatistics, deciding whether to include or exclude outliers significantly impacts your conclusions, so document and justify these decisions clearly.

Question 5

A scatterplot displays data with the following pattern: for X < 5, there is a strong positive relationship; for X > 5, there is a strong negative relationship. What would you expect the overall correlation coefficient for the entire dataset to be?

  1. Close to +1.0 since the positive relationship dominates
  2. Close to -1.0 since negative relationships are typically stronger
  3. Close to 0 since the positive and negative portions cancel each other out (correct answer)
  4. Impossible to determine without knowing the exact number of points in each region
  5. Close to +0.5 representing the average of the positive and negative correlations
Explanation: When you encounter correlation questions involving datasets with mixed patterns, remember that correlation coefficients measure the overall linear relationship across the entire dataset, not just portions of it. In this scenario, the dataset contains two distinct regions with opposing relationships. For X < 5, points follow a strong positive trend (as X increases, Y increases). For X > 5, points follow a strong negative trend (as X increases, Y decreases). When you calculate a single correlation coefficient for the entire dataset, these opposing patterns effectively counteract each other. The positive correlation in the first region and the negative correlation in the second region cancel out, resulting in an overall correlation close to zero. Answer A is incorrect because positive relationships don't automatically "dominate" - the correlation coefficient treats all data points equally regardless of the direction of local relationships. Answer B contains a false premise; negative relationships aren't inherently stronger than positive ones, and strength depends on how tightly points cluster around a trend line, not the direction. Answer D is wrong because while the exact number of points in each region affects the precise value, the fundamental principle remains: opposing strong relationships in different regions will produce a correlation near zero regardless of the specific distribution of points. Study tip: Remember that correlation coefficients can be misleading when datasets contain multiple distinct patterns. Always examine scatterplots visually before interpreting correlation values - a correlation near zero doesn't always mean "no relationship," but could indicate opposing relationships that cancel out.

Question 6

Two variables show a scatterplot pattern where most points cluster in two distinct groups: one group with low X and low Y values, and another group with high X and high Y values, with very few points in between. What characteristic would this pattern most likely exhibit?

  1. A correlation coefficient close to zero due to the gap in the middle range
  2. A high positive correlation coefficient despite the unusual distribution pattern (correct answer)
  3. A negative correlation coefficient because the groups are separated
  4. An undefined correlation coefficient because the data is not continuous
  5. A correlation coefficient that varies depending on which group is analyzed separately
Explanation: When you encounter scatterplot patterns in biostatistics, focus on the overall trend between variables rather than getting distracted by unusual distributions or gaps in the data. The described pattern shows two distinct clusters where both X and Y values move together - low X with low Y, and high X with high Y. This indicates a strong positive relationship between the variables. The correlation coefficient measures the strength and direction of linear association, and it will detect this positive trend regardless of whether the data points are evenly distributed across the range or clustered in groups. Option B correctly identifies that this pattern would produce a high positive correlation coefficient. The clustering doesn't eliminate the strong positive relationship - it actually reinforces it by showing that when X increases, Y consistently increases as well. Option A incorrectly assumes that gaps in data automatically weaken correlation. The correlation coefficient isn't calculated based on whether every possible X-value has a corresponding data point, but rather on how consistently the existing points follow a linear trend. Option C misunderstands what creates negative correlation. Separated groups don't automatically mean negative correlation - it's the direction of the relationship that matters. Since both variables increase together, the relationship remains positive. Option D confuses correlation requirements with data types. Correlation can be calculated for any numeric variables, and gaps in the middle range don't make data "non-continuous" in the statistical sense. Study tip: Remember that correlation measures consistency of relationship direction, not how evenly distributed your data points are across the range.

Question 7

A researcher creates a scatterplot of study time versus test performance and observes that students who study either very little (< 2 hours) or excessively (> 12 hours) tend to perform poorly, while those with moderate study time (4-8 hours) perform best. If the researcher calculates a linear correlation coefficient for this data, what would be the most appropriate interpretation?

  1. The correlation coefficient accurately captures the strength of the relationship between study time and performance
  2. The correlation coefficient may underestimate the true relationship strength because the association is curvilinear (correct answer)
  3. The correlation coefficient will be artificially inflated due to the presence of extreme values
  4. The correlation coefficient is the most appropriate measure since all educational relationships are linear
  5. The correlation coefficient cannot be calculated for this type of behavioral data
Explanation: When analyzing relationships between variables, you need to consider whether the association follows a linear pattern before interpreting correlation coefficients. The Pearson correlation coefficient specifically measures the strength of linear relationships, which becomes problematic when dealing with curvilinear (non-linear) patterns. The scenario describes a curvilinear relationship where moderate study time yields the best performance, while both very low and very high study times result in poor performance. This creates an inverted U-shaped or quadratic pattern. When you calculate a linear correlation coefficient for this type of relationship, it will likely produce a value close to zero because the linear trend line tries to average out the upward and downward portions of the curve, missing the true underlying association. Option B correctly identifies that the correlation coefficient may underestimate the relationship strength because it's designed for linear associations, not curvilinear ones. Option A is wrong because linear correlation cannot accurately capture non-linear relationships. Option C incorrectly suggests the correlation will be inflated—actually, the curvilinear pattern will likely produce a weak linear correlation despite a strong overall relationship. Option D makes the false assumption that all educational relationships are linear, which contradicts the evidence presented in the scatterplot. Study tip: Always examine scatterplots visually before interpreting correlation coefficients. If you see curved patterns, U-shapes, or inverted U-shapes, remember that Pearson correlation may not accurately represent the relationship strength. Consider alternative measures or transformations for non-linear associations.

Question 8

A researcher notices that their scatterplot shows a clear positive trend, but when they calculate the correlation coefficient, it is only 0.3. Which of the following interpretations is most appropriate?

  1. The correlation calculation must contain an error since visual trends should match correlation strength
  2. The relationship is positive but weak, with substantial unexplained variability around the trend (correct answer)
  3. The sample size is too small to accurately estimate the true population correlation
  4. The relationship is actually nonlinear, making the correlation coefficient inappropriate
  5. The low correlation indicates that the positive trend is likely due to random chance
Explanation: When interpreting correlation coefficients, you need to understand that correlation measures both the direction and strength of a linear relationship. A correlation can show a clear directional pattern while still being considered weak in magnitude. A correlation coefficient of 0.3 indicates a positive but weak relationship. This means that as one variable increases, the other tends to increase as well (positive direction), but there's considerable scatter around the trend line. The strength is determined by how close the coefficient is to ±1, not by how obvious the visual pattern appears. Even with substantial variability, you can still observe a clear upward trend in a scatterplot. Option A is incorrect because there's no contradiction between seeing a visual trend and having a low correlation coefficient. The correlation accurately reflects both the positive direction and weak strength of the relationship. Option C is wrong because sample size affects the precision of correlation estimates, not the fundamental interpretation of the coefficient itself. A correlation of 0.3 means the same thing regardless of whether it's from 50 or 500 observations. Option D is incorrect because nothing in the question suggests nonlinearity - a correlation of 0.3 is perfectly valid for linear relationships that simply have high variability. Remember that correlation strength is about how tightly clustered the points are around the trend line, not about whether you can see the trend. Visual clarity doesn't equal statistical strength. When you see low correlations with apparent trends, think "weak but real relationship" rather than assuming an error.

Question 9

A scatterplot shows the relationship between temperature (°F) and ice cream sales ($). The plot displays a strong positive linear relationship for temperatures between 60°F and 90°F, but shows no clear pattern for temperatures below 60°F. What does this suggest about the relationship between these variables?

  1. Temperature is always a strong predictor of ice cream sales regardless of the range
  2. The relationship between temperature and sales is conditional on temperature being above a threshold (correct answer)
  3. There is measurement error in the temperature data for readings below 60°F
  4. Ice cream sales are independent of temperature across the entire temperature range
  5. The correlation coefficient should be calculated separately for each temperature range
Explanation: When you encounter scatterplots in biostatistics, you're looking for patterns that reveal how variables relate to each other across different ranges of values. This question tests your ability to interpret conditional relationships—where the strength or nature of an association changes depending on the values of the variables involved. The key insight here is recognizing that relationships between variables don't always remain constant across their entire range. The scatterplot shows a strong positive linear relationship between temperature and ice cream sales, but only when temperatures are above 60°F. Below this threshold, there's no clear pattern, suggesting the relationship fundamentally changes. This indicates that temperature's predictive power for ice cream sales is conditional—it depends on crossing a certain temperature threshold where people actually want cold treats. Option A is incorrect because it claims temperature is always a strong predictor regardless of range, which contradicts the observed pattern below 60°F. Option C assumes measurement error, but there's no evidence provided for this—the lack of pattern could simply reflect real-world behavior where people don't buy ice cream when it's too cold. Option D wrongly suggests complete independence across all temperatures, ignoring the clear strong relationship that exists above 60°F. When analyzing scatterplots, always examine whether relationships hold consistently across the entire data range. Look for threshold effects, changing slopes, or regions where patterns break down—these often reveal important conditional relationships that simple correlation coefficients might miss.

Question 10

Examine the scatterplot. If this data were used to predict Y from X, which range of X values would likely produce the most reliable predictions?

  1. X values below 3, where the data points are most densely packed
  2. X values between 6 and 9, where the linear relationship appears strongest (correct answer)
  3. X values above 10, where there is less competition from other data points
  4. X values near 5, where the data shows the least variability in Y
  5. All X ranges would produce equally reliable predictions since correlation is constant
Explanation: The range X = 6 to 9 shows the strongest linear relationship with points closely following the trend line and minimal scatter. This would produce the most reliable predictions because there is less unexplained variability around the trend. Option A may have dense packing but not necessarily strong linear relationship. Option C suggests fewer data points would improve prediction reliability. Option D may have low variability but not necessarily good linear fit. Option E incorrectly assumes uniform reliability across all ranges.

Question 11

In the scatterplot provided, if you were to draw a best-fit line through the data, which statement about the residuals (vertical distances from points to the line) would be most accurate?

  1. Most residuals would be positive since the majority of points fall above any potential line
  2. Most residuals would be negative since the majority of points fall below any potential line
  3. Residuals would be approximately balanced between positive and negative values (correct answer)
  4. All residuals would be exactly zero if the line is drawn correctly
  5. Residuals would show a systematic pattern indicating the linear model is inappropriate
Explanation: A best-fit regression line is positioned so that the sum of residuals equals zero, meaning roughly equal numbers of points fall above and below the line, creating a balance between positive and negative residuals. This is a fundamental property of least squares regression. Options A and B incorrectly suggest systematic bias in residual signs. Option D is impossible unless all points fall exactly on the line. Option E assumes systematic patterns without evidence from the scatterplot description.

Question 12

Referring to the scatterplot, which data point would be considered the most influential in determining the correlation coefficient?

  1. The point at approximately (2, 18) because it has the highest Y-value
  2. The point at approximately (15, 8) because it deviates most from the main trend (correct answer)
  3. The point at approximately (8, 12) because it is closest to the center of the data
  4. The point at approximately (1, 15) because it has the lowest X-value
  5. All points are equally influential since correlation treats each observation the same
Explanation: The point at (15,8) appears to be an outlier that deviates substantially from the main positive trend shown by other points. Points that deviate from the overall pattern, especially those at the extremes of the X distribution, have the most influence on correlation coefficients. Option A focuses on extreme Y-value but ignores trend deviation. Option C incorrectly suggests central points are most influential (they typically have less influence). Option D focuses on extreme X-value but ignores trend conformity. Option E is incorrect because outliers and extreme points have disproportionate influence.

Question 13

What would happen to a correlation coefficient if all Y-values in a dataset were increased by exactly 10 units?

  1. The correlation would increase because all Y-values are larger
  2. The correlation would decrease because the variance in Y increases
  3. The correlation would remain exactly the same as before the transformation (correct answer)
  4. The correlation would become undefined because the transformation changes the scale
  5. The correlation would approach 1.0 because the linear relationship is strengthened
Explanation: Adding a constant to all Y-values is a linear transformation that shifts the entire distribution but does not change the correlation coefficient. Correlation measures the strength of linear relationship and is invariant under linear transformations (adding/subtracting constants, multiplying by positive constants). The relative positions and spread of points remain the same. Options A, B, D, and E all incorrectly suggest that correlation changes with this transformation.

Question 14

Based on the scatterplot shown, which statement about the strength of association is most accurate when comparing the left half (X < 6) versus the right half (X ≥ 6) of the data?

  1. The association is stronger in the left half due to higher density of data points
  2. The association is stronger in the right half due to clearer linear trend with less scatter
  3. The association strength is approximately equal in both halves of the data range
  4. The association is stronger in the left half due to steeper slope of the relationship
Explanation: B

Question 15

Based on the scatterplot shown, what would be the most reasonable expectation for the Y-value when X = 7?

  1. Approximately 14, based on extrapolating the linear trend
  2. Approximately 18, based on the highest Y-value in the nearby region
  3. Approximately 10, based on the lowest Y-value in the nearby region
  4. Approximately 16, based on interpolating between nearby data points (correct answer)
  5. Cannot be determined because X = 7 is outside the range of the data
Explanation: Looking at the scatterplot, X = 7 falls within the data range, and interpolating between the nearby points around X = 6 (Y ≈ 15) and X = 8 (Y ≈ 17) suggests a Y-value of approximately 16. This represents reasonable interpolation within the data range. Option A gives a value too low based on visible trend. Options B and C use extreme values rather than central tendency. Option E incorrectly states that X = 7 is outside the data range when it clearly falls within it.