For a sample of data, the covariance between variables X and Y is Cov(X,Y)=−40. The variance of X is sx2=25, and the standard deviation of Y is sy=10. What is the Pearson correlation coefficient r?
Practice Correlation in College Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.
What this quiz covers
This quiz focuses on Correlation, giving you a quick way to practice the rules, question types, and explanations that matter most for College Statistics.
How to use this quiz
Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.
All questions
Question 1
For a sample of data, the covariance between variables X and Y is Cov(X,Y)=−40. The variance of X is sx2=25, and the standard deviation of Y is sy=10. What is the Pearson correlation coefficient r?
-0.08
-0.16
-0.40
-0.80 (correct answer)
Explanation: The formula for the correlation coefficient using covariance is r=sxsyCov(X,Y). We are given Cov(X,Y)=−40, sx2=25, and sy=10. First, we need the standard deviation of X, which is sx=sx2=25=5. Now, substitute the values into the formula: r=5×10−40=50−40=−0.80. Distractors are based on common errors like using variance instead of standard deviation for X (−40/(25∗10)=−0.16) or other miscalculations.
Question 2
A marketing team is studying the relationship between advertising spending and monthly sales for 50 different retail stores. Initially, the data show a correlation of r=0.15. Upon investigation, they discover one store that is an extreme outlier: it had almost zero advertising spending but exceptionally high sales due to a celebrity endorsement. If this one store is removed from the analysis, what is the most likely effect on the correlation coefficient?
r will increase significantly towards a positive value. (correct answer)
r will decrease significantly, likely becoming negative.
r will remain close to 0.15, as removing one point from 50 is a minor change.
r will become exactly 0.
Explanation: The outlier store (low X, high Y) contradicts the expected positive relationship between ad spending and sales. This type of outlier, which goes against the general trend of the other data, suppresses the correlation coefficient, pulling it down towards zero or even making it negative. The initial low correlation of 0.15 suggests there is a weak positive trend being obscured. By removing the outlier that violates this trend, the underlying positive relationship among the other 49 stores will become more apparent, causing the correlation coefficient to increase significantly.
Question 3
A researcher finds a correlation of r=0.7 between the average number of hours of sleep per night and the average GPA for a sample of college students. The researcher concludes that individual students who sleep more tend to get higher grades. What is the primary flaw in this conclusion?
The correlation is not strong enough to draw any conclusions.
The sample of college students may not be representative of the entire population.
The conclusion implies causation, but the relationship could be reversed: higher grades may lead to more sleep.
The conclusion makes an assertion about individuals based on aggregated data, which may be an ecological fallacy. (correct answer)
Explanation: When you encounter correlation studies in statistics, always pay careful attention to the level of analysis in both the data and the conclusions. This question tests your understanding of ecological fallacy—a critical concept in interpreting research findings.The researcher collected data on average sleep hours and average GPA, meaning the correlation of r=0.7 describes the relationship between group-level (aggregated) measures. However, the conclusion jumps to making claims about individual students—that individual students who sleep more tend to get higher grades. This shift from group-level data to individual-level conclusions represents an ecological fallacy, making answer D correct.Looking at the wrong answers: A is incorrect because r=0.7 actually represents a strong positive correlation—certainly strong enough to warrant analysis. B identifies a real concern about external validity, but it's not the primary flaw in this specific conclusion; the researcher could still validly conclude about individuals in their sample if the ecological fallacy weren't present. C addresses causation versus correlation, which is important, but the researcher's conclusion doesn't necessarily claim causation—just that individuals "tend to" have higher grades with more sleep, which could be interpreted as correlation.Study tip: Watch for this pattern on statistics exams: when research uses aggregated/group-level data but conclusions are drawn about individuals, consider ecological fallacy. Always match the level of analysis in the data with the level of analysis in the conclusion.
Question 4
In a study of housing prices, the correlation between the size of a house (in square feet) and its selling price is found to be r=0.7. Which of the following is a correct interpretation of the coefficient of determination, r2?
49% of the data points on the scatterplot fall on the least-squares regression line.
49% of the variation in house selling prices can be explained by the linear relationship with house size. (correct answer)
For a 1% increase in house size, the selling price is expected to increase by 49%.
The selling price of a house is, on average, 49% of its size in square feet.
Explanation: The coefficient of determination, r2, is calculated as (0.7)2=0.49, or 49%. It represents the proportion of the total variance in the dependent variable (selling price) that can be explained by its linear relationship with the independent variable (house size). Option A confuses r2 with the fit of the points. Option C misinterprets r2 as a slope or elasticity. Option D is a nonsensical interpretation of the value.
Question 5
For a dataset of 30 pairs of (X, Y) values, the Pearson correlation coefficient is r=0.85. A new data point is added. This point is an outlier in the Y-direction but has an X-value equal to the mean of the original X-values, xˉ. How will the addition of this point most likely affect the correlation coefficient?
It will decrease r. (correct answer)
It will increase r towards 1.
It will have a negligible effect on r because its X-value is the mean.
It will change the sign of r from positive to negative.
Explanation: The original data has a strong positive correlation. The new point (xˉ, youtlier) lies vertically away from the center of the data cloud. This point does not follow the linear trend of the other data points and will 'pull' the best-fit line towards it, weakening the observed linear association. Therefore, the correlation coefficient r will decrease.
Question 6
A city-wide study finds a strong positive correlation (r>0.8) between the number of ice cream shops per capita in a neighborhood and the rate of reported crime in that neighborhood. Which of the following is the most valid conclusion based on this finding?
The presence of ice cream shops in a neighborhood leads to an increase in crime.
Higher crime rates in a neighborhood cause more ice cream shops to open.
A lurking variable, such as higher population density, is likely associated with both more ice cream shops and higher crime rates. (correct answer)
The positive correlation must be a statistical error, as there is no logical connection between ice cream and crime.
Explanation: Correlation does not imply causation. A strong correlation between two variables can be explained by a third, lurking variable that affects both. In this case, neighborhoods with higher population density or more commercial activity are likely to have both more retail outlets (like ice cream shops) and higher reported crime rates. Options A and B infer a causal relationship, which is not justified. Option D incorrectly dismisses the statistical finding without considering confounding factors.
Question 7
A researcher investigating the link between caffeine consumption and sleep quality finds a correlation of r=−0.60 between daily milligrams of caffeine consumed and hours of sleep per night. The researcher then converts the caffeine measurements from milligrams to grams (1000 mg = 1 g). What will be the new correlation coefficient after this conversion?
-0.60 (correct answer)
-0.0006
0.60
The correlation cannot be determined without the raw data.
Explanation: The Pearson correlation coefficient (r) is a unitless measure of linear association. It is invariant to linear transformations of the variables. Converting caffeine from milligrams to grams is a linear transformation (dividing by 1000). Therefore, the correlation coefficient remains unchanged at -0.60.
Question 8
A biologist studies the relationship between the dosage of a certain chemical and the growth rate of a plant. The data show that the growth rate increases with the dosage up to a certain point, and then decreases as the dosage becomes toxic. The biologist computes the linear correlation coefficient for the entire range of dosages and finds r=0.05. What is the most appropriate interpretation?
There is no association of any kind between the chemical dosage and plant growth.
There is a very weak positive linear association between the chemical dosage and plant growth.
A strong non-linear association exists, which the linear correlation coefficient fails to capture. (correct answer)
The calculation must be incorrect, as an inverted-U shape implies a negative correlation.
Explanation: The correlation coefficient r measures the strength and direction of linear association. The described relationship is quadratic (an inverted U-shape). For such a relationship, the initial positive trend and the later negative trend can cancel each other out, resulting in a linear correlation coefficient close to zero. This does not mean there is no relationship, but rather that the relationship is not linear. Option A is too strong. Option B incorrectly interprets r without considering the context. Option D is incorrect because the overall linear trend is close to flat.
Question 9
A researcher calculates two correlation coefficients from different datasets. Dataset A yields rA=0.65 for variables X and Y. Dataset B yields rB=−0.75 for variables P and Q. Based on these coefficients, which statement accurately compares the strength of the linear relationships?
The relationship in Dataset A is stronger because the correlation is positive.
The relationship in Dataset B is stronger because ∣−0.75∣>∣0.65∣. (correct answer)
The strengths are not comparable because the variables are different in each dataset.
The relationship in Dataset A is stronger because 0.65 is a larger number than -0.75.
Explanation: The strength of a linear relationship is determined by the absolute value of the correlation coefficient. The sign (positive or negative) only indicates the direction of the relationship. Since ∣−0.75∣=0.75 and ∣0.65∣=0.65, the correlation in Dataset B indicates a stronger linear association than the one in Dataset A. The specific variables being measured do not prevent comparison of the correlation strengths.
Question 10
A sociologist wants to investigate the relationship between a person's highest level of education and their annual income. Education level is recorded as an ordinal variable (High School, Bachelor's, Master's, PhD), and income is recorded in dollars. The sociologist calculates a Pearson correlation coefficient between these two variables. What is the primary issue with this approach?
The Pearson correlation coefficient cannot be calculated if one of the variables is not normally distributed.
The Pearson correlation coefficient requires two quantitative variables, and the education level is categorical. (correct answer)
The relationship is likely to be non-linear, which Pearson's r does not measure effectively.
The units of the two variables (education level vs. dollars) are not compatible for calculating a correlation.
Explanation: The Pearson correlation coefficient (r) is designed to measure the linear relationship between two quantitative variables. While the education levels have a natural order, they are categorical (ordinal) and do not have a consistent, measurable interval between them. Therefore, calculating Pearson's r is inappropriate. A different measure, like Spearman's rank correlation, would be more suitable. While normality (A) is an assumption for inference, not calculation, and non-linearity (C) is a limitation, the fundamental problem is the variable type. The units (D) are irrelevant as r is unitless.
Question 11
A researcher examines the relationship between the daily average temperature and the number of visitors at a public park. For a sample of 100 days, the data are restricted to only include days where the temperature was between 20°C and 25°C, yielding a correlation of r=0.25. If the researcher had instead used data from the entire year, including very cold and very hot days, how would the correlation coefficient most likely have changed?
It would likely be closer to 0.
It would likely become negative.
It would likely remain approximately 0.25.
It would likely be closer to 1. (correct answer)
Explanation: When you encounter correlation questions involving restricted ranges, think about how limiting your data affects the relationship you can observe between variables. The key concept here is restriction of range - a phenomenon that typically weakens correlations by cutting off parts of the data where the relationship might be strongest.The correct answer is D because expanding from the narrow 20-25°C range to include the full year would likely reveal a stronger relationship. Think about it logically: on very cold days (say, -10°C), almost no one visits the park, while on pleasant warm days (around 22°C), many people visit. On extremely hot days (35°C+), visits probably drop again due to discomfort. This creates a clearer, stronger pattern across the full temperature spectrum than what you can see in just the narrow 20-25°C window, where the relationship appears modest at r=0.25.Option A is wrong because restriction of range typically weakens correlations - the full dataset would likely show a stronger, not weaker, relationship. Option B misunderstands the logical relationship between temperature and park visits, which wouldn't suddenly become negative with more data. Option C incorrectly assumes that correlation coefficients remain stable regardless of the range of data included, ignoring how restricted ranges can mask stronger underlying relationships.Study tip: Remember that restricted ranges usually weaken observed correlations. When you see a modest correlation in a limited dataset, expanding the range often reveals stronger relationships by including the extreme values where differences are most pronounced.
Question 12
The correlation between two variables X and Y is r=−0.5. What percentage of the variance in Y is not explained by the linear relationship with X?
25%
50%
75% (correct answer)
It cannot be determined from the information given.
Explanation: First, calculate the coefficient of determination, r2. Here, r2=(−0.5)2=0.25. This value, 0.25 or 25%, represents the proportion of the variance in Y that is explained by the linear relationship with X. The question asks for the percentage of variance that is not explained. This is calculated as 1−r2. So, 1−0.25=0.75, or 75%.
Question 13
A researcher finds a correlation of r=0.50 between a self-reported measure of weekly exercise hours and a physiological measure of cardiovascular health. It is later discovered that the self-reported hours are unreliable and contain significant random measurement error. If the researcher were able to obtain the true, error-free exercise hours for each participant, how would the correlation with cardiovascular health most likely change?
The magnitude of the correlation would increase. (correct answer)
The magnitude of the correlation would decrease.
The correlation would remain unchanged.
The sign of the correlation would reverse.
Explanation: Random measurement error in one of the variables (in this case, self-reported exercise) tends to 'attenuate' the correlation, meaning it biases the observed correlation coefficient toward zero. The random errors add noise to the data, obscuring the true strength of the linear relationship. If the true, error-free data were used, this noise would be removed, and the underlying correlation would be revealed to be stronger. Therefore, the magnitude of r would increase from 0.50 towards a higher value.
Question 14
Four researchers are studying the same relationship between variables X and Y. Each researcher works with a different dataset. Which researcher's results indicate the strongest linear relationship between X and Y?
Researcher A computes a correlation of r=0.70.
Researcher B finds that the coefficient of determination is r2=0.50.
Researcher C reports a correlation of r=−0.80. (correct answer)
Researcher D finds that the covariance is Cov(X,Y)=100.
Explanation: To compare the strength of linear relationships, we must compare the absolute values of the correlation coefficients. For Researcher A, ∣r∣=0.70. For Researcher B, r2=0.50 implies ∣r∣=0.50≈0.707. For Researcher C, ∣r∣=∣−0.80∣=0.80. For Researcher D, covariance is not a standardized measure of association; its magnitude depends on the units and variability of X and Y, so it cannot be directly compared to r. Comparing the absolute values, 0.80 is the largest, indicating the strongest linear relationship.
Question 15
A researcher investigating the link between caffeine consumption and sleep quality finds a correlation of r=−0.60 between daily milligrams of caffeine consumed and hours of sleep per night. The researcher then converts the caffeine measurements from milligrams to grams (1000 mg = 1 g). What will be the new correlation coefficient after this conversion?
-0.60 (correct answer)
-0.0006
0.60
The correlation cannot be determined without the raw data.
Explanation: The Pearson correlation coefficient (r) is a unitless measure of linear association. It is invariant to linear transformations of the variables. Converting caffeine from milligrams to grams is a linear transformation (dividing by 1000). Therefore, the correlation coefficient remains unchanged at -0.60.
Question 16
For a dataset of 30 pairs of (X, Y) values, the Pearson correlation coefficient is r=0.85. A new data point is added. This point is an outlier in the Y-direction but has an X-value equal to the mean of the original X-values, xˉ. How will the addition of this point most likely affect the correlation coefficient?
It will decrease r. (correct answer)
It will increase r towards 1.
It will have a negligible effect on r because its X-value is the mean.
It will change the sign of r from positive to negative.
Explanation: The original data has a strong positive correlation. The new point (xˉ, youtlier) lies vertically away from the center of the data cloud. This point does not follow the linear trend of the other data points and will 'pull' the best-fit line towards it, weakening the observed linear association. Therefore, the correlation coefficient r will decrease.
Question 17
In a study of housing prices, the correlation between the size of a house (in square feet) and its selling price is found to be r=0.7. Which of the following is a correct interpretation of the coefficient of determination, r2?
49% of the data points on the scatterplot fall on the least-squares regression line.
49% of the variation in house selling prices can be explained by the linear relationship with house size. (correct answer)
For a 1% increase in house size, the selling price is expected to increase by 49%.
The selling price of a house is, on average, 49% of its size in square feet.
Explanation: The coefficient of determination, r2, is calculated as (0.7)2=0.49, or 49%. It represents the proportion of the total variance in the dependent variable (selling price) that can be explained by its linear relationship with the independent variable (house size). Option A confuses r2 with the fit of the points. Option C misinterprets r2 as a slope or elasticity. Option D is a nonsensical interpretation of the value.
Question 18
A researcher calculates two correlation coefficients from different datasets. Dataset A yields rA=0.65 for variables X and Y. Dataset B yields rB=−0.75 for variables P and Q. Based on these coefficients, which statement accurately compares the strength of the linear relationships?
The relationship in Dataset A is stronger because the correlation is positive.
The relationship in Dataset B is stronger because ∣−0.75∣>∣0.65∣. (correct answer)
The strengths are not comparable because the variables are different in each dataset.
The relationship in Dataset A is stronger because 0.65 is a larger number than -0.75.
Explanation: The strength of a linear relationship is determined by the absolute value of the correlation coefficient. The sign (positive or negative) only indicates the direction of the relationship. Since ∣−0.75∣=0.75 and ∣0.65∣=0.65, the correlation in Dataset B indicates a stronger linear association than the one in Dataset A. The specific variables being measured do not prevent comparison of the correlation strengths.
Question 19
An analyst is studying the relationship between employees' scores on a standardized logic test and their annual performance rating. For sales staff, the correlation is r=0.5. For technical staff, the correlation is r=0.6. However, when the two groups are combined, the overall correlation is r=−0.2. Which of the following provides the most plausible explanation for this change?
There must be a calculation error, as combining groups with positive correlations cannot result in a negative one.
The technical staff generally scores much higher on the logic test but receives lower performance ratings than the sales staff. (correct answer)
The relationship between test scores and performance ratings is non-linear for the combined dataset.
The sample sizes of the two groups are too different, which invalidates the combined correlation.
Explanation: This is an example of Simpson's Paradox. The overall correlation can be different from, and even opposite to, the correlations within subgroups. This occurs when there is a confounding variable (in this case, employee type) that is associated with both variables of interest. If the technical staff forms a cluster of points with high logic scores and low performance ratings, and the sales staff forms a cluster with low logic scores and high performance ratings, combining them can create an overall negative trend, even if the trend within each cluster is positive.
Question 20
A sociologist wants to investigate the relationship between a person's highest level of education and their annual income. Education level is recorded as an ordinal variable (High School, Bachelor's, Master's, PhD), and income is recorded in dollars. The sociologist calculates a Pearson correlation coefficient between these two variables. What is the primary issue with this approach?
The Pearson correlation coefficient cannot be calculated if one of the variables is not normally distributed.
The Pearson correlation coefficient requires two quantitative variables, and the education level is categorical. (correct answer)
The relationship is likely to be non-linear, which Pearson's r does not measure effectively.
The units of the two variables (education level vs. dollars) are not compatible for calculating a correlation.
Explanation: The Pearson correlation coefficient (r) is designed to measure the linear relationship between two quantitative variables. While the education levels have a natural order, they are categorical (ordinal) and do not have a consistent, measurable interval between them. Therefore, calculating Pearson's r is inappropriate. A different measure, like Spearman's rank correlation, would be more suitable. While normality (A) is an assumption for inference, not calculation, and non-linearity (C) is a limitation, the fundamental problem is the variable type. The units (D) are irrelevant as r is unitless.