AP Statistics Quiz: Linear Regression Models
20 questions · exam conditions
0:00
Linear Regression ModelsQuestion 1 of 20

A streaming service sampled 20 users and recorded age (xx, years, from 13 to 62) and average hours streamed per week (yy). A scatterplot with the least-squares regression line is shown, with fitted equation y^=18.00.15x\hat{y}=18.0-0.15x. The purpose of the linear model is to describe the linear association and predict typical weekly streaming time from age within the observed range. Which interpretation of the model is correct?

For each additional year of age, the predicted weekly streaming time decreases by about 0.15 hours, on average, for users like those sampled.
If a user gets 1 year older, that causes their weekly streaming time to drop by exactly 0.15 hours.
At age 0, the model predicts 18 hours per week, so newborns would typically stream about 18 hours weekly.
A 120-year-old is predicted to stream 18.00.15(120)18.0-0.15(120) hours per week, and this prediction is as reliable as those for ages 13–62.
The intercept 18.0 means that 18% of users stream each week.
← Back to quizzes

AP Statistics Quiz

AP Statistics Quiz: Linear Regression Models

Practice Linear Regression Models in AP Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Linear Regression Models, giving you a quick way to practice the rules, question types, and explanations that matter most for AP Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A streaming service sampled 20 users and recorded age (xx, years, from 13 to 62) and average hours streamed per week (yy). A scatterplot with the least-squares regression line is shown, with fitted equation y^=18.00.15x\hat{y}=18.0-0.15x. The purpose of the linear model is to describe the linear association and predict typical weekly streaming time from age within the observed range. Which interpretation of the model is correct?

  1. For each additional year of age, the predicted weekly streaming time decreases by about 0.15 hours, on average, for users like those sampled. (correct answer)
  2. If a user gets 1 year older, that causes their weekly streaming time to drop by exactly 0.15 hours.
  3. At age 0, the model predicts 18 hours per week, so newborns would typically stream about 18 hours weekly.
  4. A 120-year-old is predicted to stream 18.00.15(120)18.0-0.15(120) hours per week, and this prediction is as reliable as those for ages 13–62.
  5. The intercept 18.0 means that 18% of users stream each week.

Explanation: This question tests interpretation of age-related regression models. The equation y^=18.00.15x\hat{y}=18.0-0.15x predicts weekly streaming hours from user age. The correct interpretation (A) states that for each additional year of age, predicted weekly streaming time decreases by about 0.15 hours on average. This properly acknowledges the associative and average nature of the relationship. Choice B incorrectly implies causation from aging itself. Choice C attempts to apply the model to newborns (age 0), far outside the observed range of 13-62 years. Choice D extrapolates to 120 years, well beyond the data range. Choice E completely misinterprets what the intercept represents. Linear models should only be used within the range of observed data - extrapolation to extreme ages produces unreliable and often nonsensical predictions.

Question 2

A botanist measured the amount of fertilizer applied to a plot (xx, in grams) and the plant height after 6 weeks (yy, in centimeters) for 9 plots, with fertilizer amounts ranging from 0 to 40 grams. A least-squares regression line is y^=12.4+0.48x\hat{y} = 12.4 + 0.48x. The purpose of this linear model is to summarize the linear association and predict typical plant height for fertilizer amounts within the observed range. Which interpretation of the model is correct?

  1. For each additional gram of fertilizer, the predicted plant height increases by about 0.48 cm, on average, for plots like those observed. (correct answer)
  2. Applying one more gram of fertilizer causes every plant to grow 0.48 cm taller than it otherwise would.
  3. If 0 grams of fertilizer are applied, the plant height will be exactly 12.4 cm.
  4. At 200 grams of fertilizer, the model can be used to predict plant height accurately because regression lines work for any xx.
  5. The intercept 12.4 means fertilizer explains 12.4% of plant height.

Explanation: The skill involves interpreting y^=12.4+0.48x\hat{y} = 12.4 + 0.48x for plant height from fertilizer (x from 0 to 40 grams). The slope shows 0.48 cm taller per gram on average in the range. Choice A correctly interprets without causation. Distractor B assumes causation for every plant. Choice C treats the intercept as exact for zero fertilizer. Limitation: no extrapolation beyond data, as choice D does to 200 grams. Slope doesn't equal explained variation.

Question 3

An environmental scientist measured water temperature (xx, in °C, from 6 to 24) and dissolved oxygen (yy, mg/L) at 12 sites in a river. The least-squares regression line predicting dissolved oxygen from temperature is y^=12.10.18x\hat{y}=12.1-0.18x. The purpose of this linear model is to describe the linear association and predict typical dissolved oxygen for temperatures in the observed range. Which interpretation of the model is correct?

  1. For each 1°C increase in water temperature, the predicted dissolved oxygen decreases by about 0.18 mg/L, on average, for sites like those measured. (correct answer)
  2. Raising the temperature by 1°C causes dissolved oxygen to decrease by exactly 0.18 mg/L at every site.
  3. At 0°C, dissolved oxygen will be exactly 12.1 mg/L, so the model is accurate for freezing conditions.
  4. Because the slope is negative, higher temperatures and dissolved oxygen are independent.
  5. The intercept 12.1 means 12.1% of the sites had 0 dissolved oxygen.

Explanation: This question assesses understanding of regression in environmental science. The equation y^=12.10.18x\hat{y}=12.1-0.18x predicts dissolved oxygen from water temperature. The correct answer (A) properly interprets the slope: for each 1°C increase in temperature, predicted dissolved oxygen decreases by about 0.18 mg/L on average. This uses appropriate language for observational data. Choice B incorrectly claims causation and exact effects at every site. Choice C attempts to extrapolate to 0°C, outside the observed range of 6-24°C. Choice D incorrectly claims independence when negative slope indicates negative association. Choice E completely misinterprets the intercept. While temperature likely does causally affect dissolved oxygen, the regression model itself only describes the observed association within the measured temperature range.

Question 4

A school counselor collected data from 12 students on the number of hours they studied for a final exam (xx, from 1 to 9 hours) and their exam score (yy, in points). A least-squares regression line was fit to predict score from study hours: y^=58.2+4.1x\hat{y}=58.2+4.1x. The purpose of this linear model is to summarize the linear association and predict typical exam score from study time within the observed range. Which interpretation of the model is correct?

  1. If a student studies 0 hours, the model guarantees the student will score exactly 58.2 points.
  2. For each additional hour studied, the predicted exam score increases by about 4.1 points, on average, for students similar to those in the data. (correct answer)
  3. Because the slope is positive, studying an extra hour causes every student's score to increase by 4.1 points.
  4. A student who studies 20 hours is predicted to score 58.2+4.1(20)58.2+4.1(20) points, so the model is accurate for any study time.
  5. About 58.2% of the variation in exam scores is explained by study hours.

Explanation: This question tests understanding of slope interpretation in a linear regression model. The regression equation y^=58.2+4.1x\hat{y}=58.2+4.1x models the relationship between study hours and exam scores. The correct interpretation (B) states that for each additional hour studied, the predicted exam score increases by about 4.1 points on average for students similar to those in the data. This properly acknowledges that the slope represents an average association, not a guarantee for individuals. Choice C incorrectly implies causation and exact outcomes for every student. Choice A misinterprets the intercept as a guarantee rather than a prediction. Choice D incorrectly extrapolates to 20 hours, which is far beyond the observed range of 1-9 hours. Choice E confuses the intercept with R-squared. Remember that regression models describe average relationships within the observed data range, not causal effects or guarantees for individuals.

Question 5

A real estate agent recorded the size of a house (xx, in hundreds of square feet) and its selling price (yy, in thousands of dollars) for 11 homes in a neighborhood. Sizes ranged from 12 to 28 (i.e., 1200 to 2800 sq ft). A least-squares regression line is y^=95+8.7x.\hat{y}=95+8.7x. The purpose of this linear model is to summarize the linear relationship and predict typical selling prices for houses within the observed size range. Which interpretation of the model is correct?

  1. A house that is 0 square feet would be predicted to sell for $95{,}000, so the model is unrealistic and cannot be used at all.
  2. For each additional 100 square feet of size, the predicted selling price increases by about 8,7008{,}700, on average, for houses like those observed. (correct answer)
  3. Increasing a home's size by 100 square feet causes the selling price to increase by 8,7008{,}700 for every home.
  4. A 3500-square-foot house (x = 35) is predicted to sell for 95+8.7(35)95+8.7(35) thousand dollars, so the model should be used for any home size.
  5. The intercept 95 means most houses in the neighborhood are about 95,00095{,}000.

Explanation: Interpreting the regression model \hat{y} = 95 + 8.7x for house prices based on size (x in hundreds of sq ft, from 12 to 28) is the key skill. The slope means each 100 sq ft increase is associated with about 8,7008,700 higher predicted price on average within observed sizes. Choice B correctly states this without causal language or extrapolation. Choice C is a distractor, wrongly implying causation from size to price. Choice A dismisses the model due to an unrealistic zero-size intercept, but intercepts can be useful even if extrapolated. Limitations: avoid using the model beyond data, as choice D does for 3500 sq ft. Regression captures linear trends but doesn't account for other variables affecting prices.

Question 6

A researcher recorded the distance from a city center (xx, in miles) and the monthly rent for a one-bedroom apartment (yy, in dollars) for 13 apartments, with distances ranging from 1 to 18 miles. A least-squares regression line is y^=185042x.\hat{y}=1850-42x. The purpose of this linear model is to summarize the linear association and predict typical rents for apartments within the observed distance range. Which interpretation of the model is correct?

  1. Each additional mile from the city center is associated with a decrease of about $42 in the predicted monthly rent, on average, for apartments like those observed. (correct answer)
  2. Moving an apartment 1 mile farther from the city center causes the rent to drop by $42.
  3. At 0 miles from the city center, every apartment will rent for exactly $1850.
  4. An apartment 40 miles away is predicted to rent for 185042(40)1850-42(40), so the model is valid far beyond the observed distances.
  5. The intercept 1850 means the average rent of all apartments in the city is $1850.

Explanation: The skill is interpreting \hat{y} = 1850 - 42x for rent versus distance (x from 1 to 18 miles). The slope shows each mile farther associates with $42 lower predicted rent on average in the range. Choice A is correct, avoiding causation and sticking to data. Choice B distracts by claiming direct causation from distance to rent drop. Choice C misinterprets the intercept as exact for zero miles. Limitation: no extrapolation, unlike choice D to 40 miles. Intercepts estimate averages but may not reflect reality outside data.

Question 7

A district analyzed 10 schools, recording average class size (xx, students per class, from 18 to 34) and average standardized test score (yy, points). A least-squares regression line was fit: y^=6103.5x\hat{y}=610-3.5x. The purpose of this linear model is to summarize the linear association and predict typical test score from class size within the observed range. Which interpretation of the model is correct?

  1. Increasing a school's average class size by 1 student will cause the school's average test score to drop by 3.5 points.
  2. For each additional student in average class size, the predicted average test score decreases by about 3.5 points, on average, for schools like those studied. (correct answer)
  3. Because the intercept is 610, a school with 0 students per class would score 610 points, and this is a reliable prediction.
  4. The equation shows that smaller classes are the only reason some schools have higher scores.
  5. Since the slope is negative, the correlation must be r=0r=0.

Explanation: This question examines proper interpretation of regression in educational policy context. The equation y^=6103.5x\hat{y}=610-3.5x predicts average test scores from average class size. The correct answer (B) properly interprets the slope: for each additional student in average class size, the predicted average test score decreases by about 3.5 points on average. This uses appropriate statistical language avoiding causal claims. Choice A incorrectly implies causation - while smaller classes might cause higher scores, the regression only shows association. Choice C misinterprets the intercept at 0 students per class as meaningful. Choice D wrongly claims class size is the only factor. Choice E incorrectly states that negative slope means zero correlation when it actually indicates negative correlation. Regression models describe associations in observational data but cannot prove causation without proper experimental design.

Question 8

An environmental scientist models ozone level (yy, in ppb) from traffic volume (xx, in thousands of cars per day) using data from days with traffic between 10 and 60 (thousand cars). The regression line is y^=18+1.1x\hat{y}=18+1.1x. The purpose of the linear model is to describe the association and predict typical ozone levels for traffic volumes in the observed range. Which interpretation of the model is correct?

  1. For each additional 1,000 cars per day (within 10–60 thousand), the predicted ozone level increases by about 1.1 ppb, on average. (correct answer)
  2. If traffic volume is 0, then the ozone level will be 18 ppb.
  3. An increase of 1.1 ppb in ozone causes traffic volume to increase by 1,000 cars per day.
  4. Because the slope is positive, increasing traffic causes ozone to increase by 1.1 ppb for every additional 1,000 cars.
  5. At 100 thousand cars per day, the model predicts 128 ppb, so it is appropriate to use the model at 100 thousand cars per day.

Explanation: This question tests understanding of slope interpretation in an environmental science context. The regression equation y^=18+1.1x\hat{y}=18+1.1x models predicted ozone levels from traffic volume (in thousands of cars), where the slope 1.1 represents the average change in predicted ozone per thousand cars. Choice A correctly states "for each additional 1,000 cars per day, the predicted ozone level increases by about 1.1 ppb, on average." Choice B incorrectly treats the intercept as an actual value rather than a prediction outside the data range. Choice C reverses causation, suggesting ozone causes traffic changes. Choice D claims direct causation from an observational study. Choice E extrapolates to 100 thousand cars, well beyond the observed range of 10-60 thousand. Regression models from observational data describe associations, not causal relationships, and should not be extrapolated beyond their data range.

Question 9

A researcher studied 13 cars and recorded vehicle weight (xx, in thousands of pounds, from 2.4 to 4.8) and highway fuel economy (yy, miles per gallon). The least-squares regression line predicting mpg from weight is y^=46.05.2x\hat{y}=46.0-5.2x. The purpose of this linear model is to describe the linear association and predict typical fuel economy for weights in the observed range. Which interpretation of the model is correct?

  1. For each additional 1,000 pounds of vehicle weight, the predicted highway fuel economy decreases by about 5.2 mpg, on average, for cars like those in the study. (correct answer)
  2. Reducing a car's weight by 1,000 pounds will cause its highway mpg to increase by exactly 5.2 for every car.
  3. A car that weighs 0 pounds would be predicted to get 46.0 mpg, and that prediction is meaningful because it comes from the model.
  4. Because the slope is negative, there is no relationship between weight and mpg.
  5. The intercept 46.0 means that 46% of cars get 0 mpg when weight is 0.

Explanation: This question tests understanding of regression interpretation in an automotive context. The equation y^=46.05.2x\hat{y}=46.0-5.2x predicts highway fuel economy from vehicle weight (in thousands of pounds). The correct interpretation (A) states that for each additional 1,000 pounds of weight, predicted highway mpg decreases by about 5.2 on average. This properly uses associative language and acknowledges the average nature of the relationship. Choice B incorrectly implies causation and exact effects. Choice C attempts to interpret the intercept at 0 weight, which is meaningless and far outside the observed range of 2.4-4.8 thousand pounds. Choice D incorrectly claims no relationship when negative slope indicates negative association. Choice E completely misinterprets the intercept. Remember that regression models describe patterns within realistic data ranges, not impossible scenarios like weightless cars.

Question 10

A manager tracked the number of customers served in an hour (xx) and the total tips earned that hour (yy, in dollars) for 18 hourly shifts, with xx ranging from 12 to 55 customers. A least-squares regression line is y^=8.5+0.62x\hat{y}=8.5+0.62x. The purpose of this linear model is to summarize the linear association and predict typical tips for shifts within the observed range. Which interpretation of the model is correct?

  1. Each additional customer served is associated with an increase of about $0.62 in the predicted total tips for that hour, on average, for shifts like those observed. (correct answer)
  2. Serving one more customer causes tips to increase by exactly $0.62 every hour.
  3. If 0 customers are served, the server will earn exactly $8.50 in tips.
  4. A shift with 120 customers can be predicted accurately using the model since the relationship is linear.
  5. The intercept 8.5 means most hours have about 8.5 customers.

Explanation: This question assesses interpreting \hat{y} = 8.5 + 0.62x for tips from customers served (x from 12 to 55). The slope indicates $0.62 more predicted tips per extra customer on average in the range. Choice A is right, limiting to association and data. Choice B wrongly claims causation for exact increases. Choice C misuses the intercept for zero customers. Key limitation: avoid extrapolation, unlike choice D to 120 customers. Intercepts may not be meaningful alone.

Question 11

A scientist collected data from 12 batteries on discharge time (xx, in hours, from 1.5 to 8.0) and operating temperature (yy, in °C, from 28 to 44). The least-squares regression line predicting temperature from discharge time is y^=26.5+2.1x.\hat{y}=26.5+2.1x. The purpose of this linear model is to summarize the linear association and predict typical operating temperature for discharge times within the observed range. Which interpretation of the model is correct?

  1. For each additional hour of discharge time, the predicted operating temperature increases by about 2.1°C, on average, for batteries with discharge times 1.5–8.0 hours. (correct answer)
  2. Increasing discharge time by 1 hour causes the operating temperature to rise by exactly 2.1°C for every battery.
  3. A battery with 0 hours of discharge time will have an operating temperature of exactly 26.5°C, so the model is accurate at x=0x=0.
  4. Because the intercept is 26.5, the temperature cannot go below 26.5°C.
  5. The model predicts the temperature at 20 hours discharge time with the same reliability as within 1.5–8.0 hours.

Explanation: This question evaluates understanding of slope interpretation in a technical context. The model y^=26.5+2.1x\hat{y}=26.5+2.1x has a slope of 2.1, meaning for each additional hour of discharge time, the predicted operating temperature increases by 2.1°C on average. Choice A correctly interprets this relationship and appropriately restricts it to the observed range of 1.5-8.0 hours. Choice B incorrectly implies causation and claims an exact temperature rise for every battery. Choice C inappropriately extrapolates to 0 hours discharge time and treats the prediction as accurate outside the data range. Choice D misinterprets the y-intercept as a physical constraint on temperature. Choice E incorrectly claims equal reliability for predictions at 20 hours as within the observed range - extrapolation far beyond observed data is much less reliable. Regression models are tools for understanding patterns within observed data, not for making predictions far outside that range.

Question 12

A student recorded the number of hours studied (xx) and the score on a quiz out of 100 (yy) for 12 classmates (hours ranged from 0.5 to 6). A least-squares regression line was fit to predict quiz score from hours studied: y^=52+6.5x\hat{y} = 52 + 6.5x. The purpose of this linear model is to summarize the linear association and predict typical quiz scores for study times within the observed range. Which interpretation of the model is correct?

  1. For each additional hour studied, the predicted quiz score increases by about 6.5 points, on average, for students with study times similar to those observed. (correct answer)
  2. If a student studies 0 hours, the model proves the student will score exactly 52 points on the quiz.
  3. Because the slope is positive, studying an extra hour causes quiz scores to increase by 6.5 points for any student.
  4. A student who studies 10 hours is predicted to score 52+6.5(10)=11752+6.5(10)=117 points, so the model is accurate for any number of hours.
  5. About 6.5% of quiz score is explained by hours studied because the slope is 6.5.

Explanation: This question assesses the skill of interpreting the slope and intercept in a linear regression model for predicting quiz scores from hours studied. The model is y^=52+6.5x\hat{y} = 52 + 6.5x, where the slope indicates that for each additional hour studied, the predicted quiz score increases by about 6.5 points on average within the observed range of 0.5 to 6 hours. Choice A correctly captures this associational interpretation without claiming causation or extrapolating beyond the data. A common distractor, like choice C, mistakenly infers causation from the positive slope, assuming that studying causes the score increase, which regression alone cannot prove. Another distractor, choice B, treats the intercept as a literal prediction for x=0, but intercepts often lack real-world meaning outside the data range. A mini-lesson on model limitations: linear regression describes associations but does not imply causation, and predictions should be restricted to the observed range of x to avoid unreliable extrapolations. Always contextualize interpretations to the sample studied, as results may not generalize.

Question 13

A teacher compared the number of pages a student read in a week (xx) with the student's score on a reading quiz (yy) for 16 students. Pages ranged from 10 to 80. The least-squares regression line is y^=58+0.35x\hat{y} = 58 + 0.35x. The purpose of this linear model is to describe the linear association and predict typical quiz scores for page counts within the observed range. Which interpretation of the model is correct?

  1. For each additional page read, the predicted quiz score increases by about 0.35 points, on average, for students with page counts similar to those observed. (correct answer)
  2. A student who reads 0 pages will score exactly 58 points on the quiz.
  3. Reading 10 more pages causes a student's quiz score to rise by 3.5 points, so increasing reading will always improve scores.
  4. If a student reads 200 pages, the model can be trusted to predict the quiz score because the relationship is linear.
  5. The slope 0.35 means 35% of the variation in quiz scores is explained by pages read.

Explanation: This question focuses on interpreting y^=58+0.35x\hat{y} = 58 + 0.35x, linking pages read (x from 10 to 80) to quiz scores. The slope indicates a 0.35-point increase per extra page on average within the range. Choice A is accurate, emphasizing association and data limits. Distractor C assumes causation, suggesting more reading always boosts scores, but that's not proven. Choice B treats the intercept as exact for zero pages, ignoring variability. Key limitation: no extrapolation, as choice D does to 200 pages. Models like this explain trends but not all variation, and slope isn't r-squared.

Question 14

An environmental club measured daily high temperature (xx, in °F) and the number of bottles of water sold at an outdoor booth (yy) for 15 days, with temperatures ranging from 60°F to 92°F. A least-squares regression line was found: y^=120+4.1x\hat{y} = -120 + 4.1x. The purpose of this linear model is to describe the relationship and predict typical sales for temperatures within the observed range. Which interpretation of the model is correct?

  1. When the temperature is 0°F, the model shows the booth will sell exactly 120 fewer bottles than usual.
  2. For each 1°F increase in temperature, the predicted number of bottles sold increases by about 4.1 bottles, on average, for days like those observed. (correct answer)
  3. Raising the temperature by 1°F causes bottle sales to increase by 4.1 bottles for any day.
  4. Because the intercept is negative, the linear model is invalid and cannot be used for prediction within 60°F to 92°F.
  5. At 100°F the booth will sell 120+4.1(100)=290-120+4.1(100)=290 bottles, so the model is reliable for all temperatures.

Explanation: This question tests the interpretation of a linear regression model relating temperature to water bottle sales, with the equation y^=120+4.1x\hat{y} = -120 + 4.1x. The slope means that for each 1°F increase, predicted sales rise by about 4.1 bottles on average for temperatures between 60°F and 92°F. Choice B is correct as it emphasizes association and limits predictions to observed conditions without causal claims. Choice C is a distractor that incorrectly assumes causation, stating that temperature changes directly cause sales increases, which correlation does not establish. Choice A misinterprets the negative intercept as an exact prediction for 0°F, ignoring that it's outside the data range and may not be meaningful. A key limitation of such models is extrapolation; for instance, predicting at 100°F as in choice E is unreliable because the relationship may not hold beyond observed data. Remember, regression models summarize observed patterns but require caution with intercepts that imply unrealistic scenarios.

Question 15

A counselor collected data on 10 students: number of absences in a semester (xx) and final course percentage (yy). Absences ranged from 0 to 12. A least-squares regression line to predict final percentage from absences is y^=932.4x.\hat{y}=93-2.4x. The purpose of this linear model is to summarize the linear association and predict typical final percentages for students with absence counts within the observed range. Which interpretation of the model is correct?

  1. Each additional absence is associated with a decrease of about 2.4 percentage points in the predicted final grade, on average, for students like those observed. (correct answer)
  2. A student with 0 absences will definitely earn exactly 93% in the course.
  3. Reducing absences by 1 causes the final grade to increase by 2.4 percentage points for every student.
  4. A student with 30 absences is predicted to earn 932.4(30)=21%93 - 2.4(30) = 21\%, so the model should be used for any absence count.
  5. Because the slope is negative, there is no relationship between absences and final grade.

Explanation: The skill here involves correctly interpreting the slope and intercept in a regression model predicting final grades from absences, given by y^=932.4x\hat{y} = 93 - 2.4x. The negative slope indicates that each additional absence is associated with a 2.4 percentage point decrease in predicted grade on average, within 0 to 12 absences. Choice A accurately reflects this without overstepping into causation or extrapolation. Choice C is a misleading distractor, claiming causation by suggesting reducing absences directly increases grades, but regression shows correlation, not cause. Choice B errs by treating the intercept as a guarantee for zero absences, whereas it's an estimate and actual scores vary. Limitations include avoiding predictions outside the data range, as choice D does by extrapolating to 30 absences, which could be inaccurate if the relationship isn't linear beyond observed values. Overall, these models are tools for description and prediction within limits, not for proving causal effects.

Question 16

A school nurse recorded the number of minutes a student spent on a treadmill test (xx) and the student's heart rate immediately afterward (yy, in beats per minute) for 12 students. Times ranged from 3 to 14 minutes. A least-squares regression line is y^=78+5.2x.\hat{y} = 78 + 5.2x. The purpose of this linear model is to describe the linear association and predict typical heart rates for treadmill times within the observed range. Which interpretation of the model is correct?

  1. For each additional minute on the treadmill, the predicted post-test heart rate increases by about 5.2 beats per minute, on average, for students like those observed. (correct answer)
  2. If a student runs 0 minutes, the model proves the heart rate will be exactly 78 bpm after the test.
  3. Running one extra minute causes a student's heart rate to increase by 5.2 bpm, so the relationship is causal.
  4. The intercept 78 means the average resting heart rate of all students is 78 bpm.
  5. A 30-minute treadmill time can be predicted accurately using the model because the regression line is linear.

Explanation: Interpreting y^=78+5.2x\hat{y} = 78 + 5.2x for heart rate after treadmill time (x from 3 to 14 minutes) tests this skill. The slope means each extra minute links to 5.2 bpm higher predicted rate on average in the range. Choice A properly frames it as association. Distractor C assumes causation from time to rate increase. Choice B sees the intercept as proof for zero minutes. Limitation: linearity doesn't guarantee extrapolation, as choice E suggests for 30 minutes. Models predict typical outcomes, not certainties.

Question 17

A manager tracked advertising spending (xx, in hundreds of dollars, from 2 to 25) and weekly sales (yy, in thousands of dollars) for 11 weeks. The least-squares regression line predicting sales from ad spending is y^=4.6+0.32x\hat{y}=4.6+0.32x. The purpose of this linear model is to describe the linear association and predict typical weekly sales for ad spending values in the observed range. Which interpretation of the model is correct?

  1. Spending an additional $100 on advertising is associated with an increase of about $0.32 thousand in predicted weekly sales, on average, for weeks like those observed. (correct answer)
  2. Because the slope is 0.32, increasing advertising by $100 causes weekly sales to rise by exactly $320 for every week.
  3. If advertising spending is $0, weekly sales will be exactly $4.6 thousand, and this is a reliable prediction.
  4. The model guarantees that sales cannot ever be below $4.6 thousand because the intercept is 4.6.
  5. Since the line fits the data, the model is accurate for predicting sales at x=100x=100 (i.e., $10,000) hundreds of dollars.

Explanation: This question tests understanding of slope interpretation in a business context. The regression equation y^=4.6+0.32x\hat{y}=4.6+0.32x predicts weekly sales (in thousands) from advertising spending (in hundreds of dollars). The correct interpretation (A) states that spending an additional $100 on advertising is associated with an increase of about 0.32thousand(0.32 thousand (320) in predicted weekly sales, on average. This properly acknowledges the associative nature and average relationship. Choice B incorrectly claims causation and exact effects. Choice C misinterprets the intercept at x=0x=0 as reliable when the data range is 2-25. Choice D misunderstands what the intercept represents. Choice E incorrectly extrapolates to x=100x=100 (i.e., $10,000), far beyond the observed range. Remember that regression models are descriptive tools for the observed data range, not prescriptive formulas guaranteeing specific outcomes.

Question 18

A fitness researcher recorded resting heart rate (yy, beats per minute) and weekly minutes of aerobic exercise (xx, from 0 to 240 minutes) for 14 adults. A scatterplot with the least-squares regression line is shown, with fitted equation y^=78.00.06x\hat{y}=78.0-0.06x. The purpose of the linear model is to describe the linear relationship and predict typical heart rate from exercise time within the observed range. Which interpretation of the model is correct?

  1. An additional minute of aerobic exercise per week is associated with about a 0.06 bpm decrease in the predicted resting heart rate, on average, for adults like those sampled. (correct answer)
  2. If someone exercises 0 minutes per week, their resting heart rate must be exactly 78 bpm.
  3. Exercising 1000 minutes per week would be predicted to produce a negative resting heart rate, so the model proves exercise is harmful.
  4. Because the slope is negative, increasing exercise causes resting heart rate to decrease by 0.06 bpm for everyone.
  5. The intercept 78.0 means 78% of the variation in resting heart rate is explained by exercise time.

Explanation: This question examines interpretation of a health-related regression model. The equation y^=78.00.06x\hat{y}=78.0-0.06x predicts resting heart rate from weekly exercise minutes. The correct answer (A) properly interprets the slope: an additional minute of exercise per week is associated with about a 0.06 bpm decrease in predicted resting heart rate, on average. This uses appropriate statistical language avoiding causal claims. Choice B misinterprets the intercept as an exact value for all non-exercisers. Choice C demonstrates the danger of extrapolation - 1000 minutes is far beyond the observed range of 0-240 minutes. Choice D incorrectly implies causation and exact effects for everyone. Choice E confuses the intercept with R-squared. Linear models describe average associations within the observed data range, not causal mechanisms or guarantees for individuals.

Question 19

A fitness coach records resting heart rate (yy, beats per minute) and weekly aerobic exercise time (xx, minutes) for 14 clients. A least-squares regression line is fit to predict resting heart rate from exercise time: y^=78.40.06x\hat{y}=78.4-0.06x. The purpose of this linear model is to describe the association and predict typical resting heart rate for clients with exercise times similar to those observed. Which interpretation of the model is correct?

  1. For each additional minute of weekly aerobic exercise, the predicted resting heart rate decreases by about 0.06 beats per minute, on average, for clients like those in the data. (correct answer)
  2. If a client does 0 minutes of exercise per week, the client will have a resting heart rate of exactly 78.4 bpm.
  3. Increasing weekly exercise by 10 minutes causes resting heart rate to drop by 0.6 bpm.
  4. Because the slope is negative, the relationship must be nonlinear.
  5. The model says 78.4% of resting heart rate is explained by exercise time.

Explanation: This question tests interpretation of a regression model with a negative slope in a health context. The equation y^=78.40.06x\hat{y}=78.4-0.06x indicates that for each additional minute of weekly exercise, the predicted resting heart rate decreases by 0.06 beats per minute on average. Choice A correctly interprets this relationship using appropriate statistical language and acknowledges the model applies to "clients like those in the data." Choice B incorrectly treats the y-intercept as an exact value, Choice C implies causation, Choice D incorrectly links negative slope to nonlinearity (linear models can have negative slopes), and Choice E confuses the y-intercept value with R-squared percentage. Remember that regression models describe average associations, not individual outcomes or causal effects, and their reliability depends on staying within the observed data range.

Question 20

A school counselor collects data from 12 students on weekly study time (xx, hours) and their quiz score (yy, points). A least-squares regression line is fit to predict quiz score from study time, with equation y^=58+3.2x\hat{y}=58+3.2x. The purpose of this linear model is to summarize the linear association and predict typical quiz scores for students with study times similar to those observed. Which interpretation of the model is correct?

  1. For each additional hour studied per week, the predicted quiz score increases by about 3.2 points, on average, for students with study times like those in the data. (correct answer)
  2. If a student studies 0 hours per week, the model proves the student will score exactly 58 points.
  3. Studying an extra hour per week causes quiz scores to increase by 3.2 points.
  4. A student who studies 20 hours per week will definitely score 58+3.2(20)=12258+3.2(20)=122 points because the line gives the exact score.
  5. About 3.2% of the variation in quiz scores is explained by study time.

Explanation: This question tests understanding of slope interpretation in linear regression models. The regression equation y^=58+3.2x\hat{y}=58+3.2x has a slope of 3.2, which represents the average change in predicted quiz score for each one-unit increase in study time. Choice A correctly interprets this as "for each additional hour studied per week, the predicted quiz score increases by about 3.2 points, on average." The key phrases "predicted," "on average," and "for students with study times like those in the data" properly acknowledge that this is a statistical model describing typical patterns, not exact values or causal relationships. Choices B and D incorrectly treat predictions as exact values, Choice C incorrectly implies causation, and Choice E confuses the slope with R-squared. Linear regression models describe associations and make predictions about typical values within the observed data range, not exact outcomes or causal effects.