Home

Tutoring

Subjects

Live Classes

Study Coach

Essay Review

On-Demand Courses

Colleges

Games


Sign up

Log in

Opening subject page...

Loading your content

Practice

  • All Subjects
  • Algebra Flashcards
  • SAT Math Practice Tests
  • Math Question of the Day
  • Live Classes
  • On-Demand Courses

Varsity Tutors

  • Find a Tutor
  • Test Prep
  • Online Classes
  • K-12 Learning
  • College Search
  • VarsityTutors.com

© 2026 Varsity Tutors. All rights reserved.

← Back to quizzes

Statistics Quiz

Statistics Quiz: Evaluate Model Fit With Residuals

Practice Evaluate Model Fit With Residuals in Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

Question 1 / 20

0 of 20 answered

A meteorologist fit the linear model y^=12+0.6x\hat{y}=12+0.6xy^​=12+0.6x to predict afternoon temperature (y, in °C) from morning temperature (x, in °C). The residuals show this pattern:

For small x values, residuals are mostly positive; for medium x values, residuals are near 0; for large x values, residuals are mostly negative.

What does the residual pattern suggest about the model choice?

Select an answer to continue

What this quiz covers

This quiz focuses on Evaluate Model Fit With Residuals, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A meteorologist fit the linear model y^=12+0.6x\hat{y}=12+0.6xy^​=12+0.6x to predict afternoon temperature (y, in °C) from morning temperature (x, in °C). The residuals show this pattern:

For small x values, residuals are mostly positive; for medium x values, residuals are near 0; for large x values, residuals are mostly negative.

What does the residual pattern suggest about the model choice?

  1. The residuals are randomly scattered around 0, suggesting the linear model is appropriate.
  2. Positive residuals mean the model overestimates afternoon temperature for small x values.
  3. The residuals show a systematic pattern, suggesting the linear model may not be appropriate and a different model form could fit better. (correct answer)
  4. Because some residuals are positive and some are negative, the model must be perfect overall.

Explanation: Using residuals to check model fit reveals if the linear form suits the data through error patterns. Residual is actual - predicted (y - ŷ), with positive signaling underestimation and negative overestimation. Random scatter means no predictable trends, just noise around zero. A pattern like positive for small x, zero in middle, negative for large x suggests systematic bias, possibly curvature. This described pattern indicates a poor fit for the temperature model, implying a nonlinear alternative. People often misinterpret sign changes as random when patterned, or think small residuals suffice despite trends. Always prioritize pattern detection over size in residual checks for robust model evaluation.

Question 2

A city planner fit the model y^=200+15x\hat{y}=200+15xy^​=200+15x to predict daily subway riders (y, in thousands) from the number of downtown events (x). The residuals for x = 0 through 6 events were:

x: 0, 1, 2, 3, 4, 5, 6 residual: 2, -1, 1, -2, 0, 2, -2

Which statement best describes how well the model fits the data?

  1. The model is a good fit because all residuals are positive, so the model is consistently close.
  2. The model is a poor fit because the residuals increase from negative to positive as x increases.
  3. The model is a good fit because the residuals are randomly scattered around 0 with no clear pattern. (correct answer)
  4. The model is a poor fit because negative residuals mean the model underestimates the number of riders.

Explanation: To evaluate model fit, residuals show if errors are random or patterned, guiding model choice. Residual is actual - predicted (y - ŷ); positive indicates underestimation, negative overestimation. Random scatter around 0 looks like irregular ups and downs with no trends. Patterns suggest missed elements, like curvature or heteroscedasticity. Here, residuals (2, -1, 1, -2, 0, 2, -2) fluctuate randomly without clear trends, supporting a good linear fit for subway riders. Misconceptions include thinking sign changes always mean patterns, but true randomness can have alternations; small sizes don't excuse patterns elsewhere. Focus on absence of patterns, not just residual magnitude, for confirming model fit.

Question 3

A student used the model y^=100−5x\hat{y}=100-5xy^​=100−5x to predict the number of pages (y) remaining in a book after x days. On day 6, the model predicted 70 pages remaining, but the actual number of pages remaining was 74.

For x = 6, what does the residual mean in context?

  1. The residual is -4, meaning the model predicted 4 fewer pages remaining than actually remained.
  2. The residual is 4, meaning the model predicted 4 fewer pages remaining than actually remained. (correct answer)
  3. The residual is 4, meaning the model predicted 4 more pages remaining than actually remained.
  4. The residual is -4, meaning the model predicted 4 more pages remaining than actually remained.

Explanation: Residuals evaluate model fit by quantifying and interpreting prediction inaccuracies. Defined as actual - predicted (y - ŷ), positive residuals mean the model underestimates, negative mean overestimates. Random scatter around 0 is unstructured variation without trends. Patterns indicate model flaws, such as unmodeled curvature. For x=6, residual is +4 (74 - 70), meaning the model predicted 4 fewer pages remaining than actual, showing underestimation. Misconceptions include sign reversal, confusing over- and underestimation interpretations. Check residuals for patterns, not just size, to ensure model appropriateness.

Question 4

A student modeled the relationship between the number of practice problems completed (x) and a quiz score (y) with the linear model y^=45+3x\hat{y}=45+3xy^​=45+3x. The residuals (defined as residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​) for several students are shown below.

x: 1, 2, 3, 4, 5, 6, 7 residual: -6, -4, -2, 0, 2, 4, 6

Which statement best describes how well the model fits the data?

  1. The model is a poor fit because positive residuals mean the model overestimates the quiz scores for larger x values.
  2. The model is a good fit because the residuals are randomly scattered around 0 with no clear trend.
  3. The model is a poor fit because the residuals show a systematic pattern (increasing from negative to positive), suggesting a different model form may be better. (correct answer)
  4. The model is a good fit because the residuals are small in magnitude, even though they show a clear pattern.

Explanation: Evaluating model fit with residuals involves examining the differences between actual and predicted values to assess if a linear model is appropriate. A residual is defined as actual value minus predicted value (y - ŷ), where a positive residual means the model underestimates the actual value and a negative residual means it overestimates. A good fit shows residuals randomly scattered around zero, meaning no discernible trend or pattern as x changes. In contrast, a clear pattern in residuals, such as a systematic increase or curve, implies the model misses some aspect of the data, like curvature or non-constant variability. In this case, the residuals steadily increase from -6 to +6 as x goes from 1 to 7, showing a clear linear trend rather than random scatter, indicating the linear model does not capture the relationship well. A common misconception is that small residuals alone indicate a good fit, but even small residuals with a pattern, like here, suggest a poor fit. Always check for patterns in residuals beyond just their size to decide if a different model might be better.

Question 5

A city planner modeled the relationship between distance from downtown xxx (miles) and average rent yyy (dollars) using y^=2200−120x\hat{y}=2200-120xy^​=2200−120x. The residual plot below shows residuals that increase in spread as xxx increases.

What does the residual pattern suggest about the model choice?

  1. The model is appropriate because residuals above 0 mean the points are below the regression line.
  2. The linear form is appropriate because the residuals are centered around 0 and show constant variability.
  3. The model is guaranteed to fit well because rent and distance are usually strongly correlated.
  4. The model is inappropriate because the residuals show a funnel shape, suggesting the variability changes with distance. (correct answer)

Explanation: Evaluating fit with residuals means plotting them to spot if the model captures the relationship properly. Residuals are defined as y - ŷ, where positive means actual > predicted (underestimation) and negative the reverse. Random scatter around 0 features even, patternless distribution with constant spread. A funnel pattern signals changing variability, suggesting the linear model is inappropriate and may need transformation. The plot shows increasing spread with x, indicating a poor choice for the rent model. A misconception is that residuals above 0 always mean points below the line or small residuals are good despite patterns. Remember to always scan for patterns like varying spread, beyond just residual magnitudes, in assessments.

Question 6

A nutritionist modeled the relationship between daily sugar intake xxx (grams) and an energy score yyy using y^=80−0.2x\hat{y}=80-0.2xy^​=80−0.2x. Residuals (y−y^y-\hat{y}y−y^​) for 8 people are shown.

Which statement best describes how well the model fits the data based on the residual pattern?

  1. The model is a good fit because most residuals are between -6 and 6, so the model must be accurate.
  2. The model is a poor fit because the residuals show a clear pattern: negative at low xxx, positive in the middle, then negative again. (correct answer)
  3. The model is a poor fit because positive residuals mean the model overestimates the energy score.
  4. The model is a good fit because the residuals are randomly scattered around 0 with no pattern.

Explanation: Residual plots assess model fit by revealing if the linear assumption holds through error patterns. A residual is actual minus predicted (y - ŷ), with positive indicating underestimation and negative indicating overestimation. Good fits show random scatter around 0, like points without trends or groupings. A pattern, such as negative-low, positive-middle, negative-high, implies missed curvature, making the model poor. The residuals here follow that wavy pattern, showing the energy score model fits poorly. People often think positive residuals mean overestimation universally or small residuals excuse patterns, but that's incorrect. Always prioritize checking for any patterns in residuals, not just their sizes, for accurate evaluations.

Question 7

A coach modeled the relationship between practice time xxx (hours) and free-throw percentage yyy using y^=60+2x\hat{y}=60+2xy^​=60+2x. Residuals were calculated as residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

Which statement best describes how well the model fits the data based on the residuals shown?

  1. The model is a good fit because the residuals show a clear increasing trend as xxx increases.
  2. The model is a poor fit because the residuals are randomly scattered around 0, so the model misses a pattern.
  3. The model is a good fit because the residuals are mixed above and below 0 with no clear pattern. (correct answer)
  4. The model is guaranteed to be a good fit because a linear model was used.

Explanation: Residual analysis for model fit checks if deviations from predictions are random, typically via a plot versus the predictor. Residuals are actual minus predicted (y - ŷ), where positive means the actual exceeds the prediction (underestimation) and negative means the opposite. Good fits exhibit random scatter around 0, with points mixed above and below without trends. Patterns indicate missed elements, like curvature or non-constant variance, calling for a better model. The residuals here are mixed above and below 0 with no clear pattern, supporting a good fit for the free-throw model. A misconception is that small residuals guarantee a good fit even with patterns, but patterns reveal underlying issues. Always examine residual patterns holistically, beyond just their magnitudes, for reliable assessments.

Question 8

An engineer modeled the relationship between machine age xxx (years) and yearly maintenance cost yyy (dollars) using y^=200+50x\hat{y}=200+50xy^​=200+50x. The residual plot below shows a clear curve.

Which statement best describes how well the model fits the data?

  1. The model is a poor fit because the residuals follow a curved pattern, suggesting the relationship may not be linear. (correct answer)
  2. The model is perfect because the residuals include values close to 0.
  3. The model is a good fit because the residuals are randomly scattered around 0 with no pattern.
  4. The model is a good fit because residuals above 0 mean the model overestimates, and that happens for some points.

Explanation: Assessing model fit with residuals involves looking for randomness in their distribution against the predictor. Defined as y - ŷ (actual minus predicted), positive residuals signal underestimation, and negative ones signal overestimation. Random scatter around 0 appears as unstructured points evenly around the zero line. A curved pattern suggests the model overlooks non-linearity, implying a poor fit and potential need for a curved model. The residual plot displays a clear curve, indicating the linear maintenance cost model fits poorly. Often, people confuse sign reversal or believe small residuals suffice despite patterns, but both can mislead. Focus on detecting patterns, not solely residual sizes, to ensure the model adequately represents the data.

Question 9

A biologist modeled the relationship between water temperature xxx (°C) and fish activity level yyy (arbitrary units) using y^=5+1.2x\hat{y}=5+1.2xy^​=5+1.2x. A residual plot is shown.

What does the residual pattern suggest about the model choice?

  1. The linear model is inappropriate because the residuals show a funnel shape, suggesting changing variability as xxx increases. (correct answer)
  2. The linear model must be inappropriate because some residuals are above 0, meaning the points are below the model.
  3. The linear model is perfect because the residuals alternate between positive and negative values.
  4. The linear model appears appropriate because the residuals are randomly scattered around 0 with roughly constant spread.

Explanation: To assess model fit, residuals are plotted against the predictor variable, helping identify if the linear model suits the data. A residual equals actual value minus predicted value (y - ŷ), with positive indicating underestimation and negative indicating overestimation. Random scatter around 0 looks like points haphazardly above and below the line, with even spread. A funnel-shaped pattern implies changing variability, meaning the linear model doesn't account for heteroscedasticity and may be inappropriate. The residual plot here shows a funnel shape, suggesting the fish activity model needs reevaluation. Commonly, people think positive residuals mean all points are below the line, but it actually means overestimation for those points. Prioritize checking for patterns like varying spread over residual size alone when evaluating fits.

Question 10

A business uses the model y^=8x+20\hat{y}=8x+20y^​=8x+20 to predict weekly sales (y, in hundreds of dollars) from advertising spending (x, in hundreds of dollars). Residuals are residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

Which statement best describes how well the model fits the data?

  1. The model is a good fit because positive residuals mean the model overestimates sales.
  2. The model is a good fit because the residuals are randomly scattered around 0 with no clear pattern. (correct answer)
  3. The model is a poor fit because the residuals show a clear curved pattern, suggesting a different model form may be better.
  4. The model is a poor fit because residuals close to 0 always indicate underestimation.

Explanation: Assessing model fit with residuals involves checking for patterns that would indicate model inadequacy. Residuals, calculated as actual minus predicted (y - ŷ), should ideally show random scatter around 0 if the linear model is appropriate. Random scatter means no systematic pattern emerges as you move across different x-values—the residuals appear unpredictable and evenly distributed above and below zero. When residuals display a clear pattern like a curve or changing spread, it signals the linear model is missing important structure in the data. In this sales prediction scenario, the residuals are randomly scattered around 0 with no discernible pattern, confirming the linear model adequately captures the relationship between advertising spending and sales. A common error is misinterpreting what positive residuals mean—they indicate underestimation, not overestimation. The key principle is that pattern detection, not residual size or sign alone, determines model appropriateness.

Question 11

A researcher uses y^=3x+5\hat{y}=3x+5y^​=3x+5 to predict the number of pages read (y) from the number of days in a reading program (x). Residuals are defined as residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

For x=10x=10x=10 days, the residual is +6+6+6. What does this residual mean in context?

  1. The model underestimated the pages read by 6 pages. (correct answer)
  2. The predicted number of pages read was 6 pages greater than the actual number because residuals are y^−y\hat{y}-yy^​−y.
  3. The participant read 6 fewer days than predicted by the model.
  4. The model overestimated the pages read by 6 pages.

Explanation: Interpreting individual residuals requires understanding both their calculation and contextual meaning. A residual is computed as actual minus predicted (residual = y - ŷ), measuring how far and in which direction the prediction missed. A positive residual of +6 means the actual value exceeded the predicted value by 6 units, indicating the model underestimated. For this reading study, with x = 10 days, the model predicted ŷ = 3(10) + 5 = 35 pages, but the participant actually read 41 pages (since 41 - 35 = +6). This means the model underestimated the participant's reading by 6 pages. A common error is confusing what the residual measures—it's about the response variable (pages read), not the predictor (days), and some incorrectly reverse the formula thinking residuals are ŷ - y. Remember the mnemonic: positive residual means the actual was positively surprising (higher than predicted), indicating underestimation.

Question 12

A mechanic models fuel efficiency (y, in mpg) from vehicle speed (x, in mph) using y^=0.2x+10\hat{y}=0.2x+10y^​=0.2x+10. Residuals are computed as residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

Which statement best describes how well the model fits the data?

  1. The model fits well because the residuals are randomly scattered around 0 with no clear pattern.
  2. The model is a good fit because the residuals are small in magnitude, even though they form a systematic pattern.
  3. The model is a poor fit because the residuals show a clear curved pattern, suggesting the relationship is not linear. (correct answer)
  4. The model is a poor fit because negative residuals mean the model underestimated mpg.

Explanation: When using residuals to evaluate model fit, the presence of patterns is more important than the magnitude of the residuals. A residual equals actual minus predicted (y - ŷ), with positive values indicating underestimation and negative values indicating overestimation. For a linear model to be appropriate, residuals should scatter randomly around 0 without forming any systematic pattern. A curved pattern in residuals—such as negative values at low x, positive in the middle, and negative again at high x—strongly suggests the true relationship is non-linear. In this fuel efficiency example, the residuals show a clear curved pattern, indicating that the relationship between speed and mpg is not adequately captured by a straight line, likely because fuel efficiency typically has a curved relationship with speed. A dangerous misconception is thinking small residuals guarantee a good fit—even tiny residuals forming a curve reveal model inadequacy. Always prioritize pattern detection over residual magnitude when assessing fit.

Question 13

A biologist models the mass of a plant (y, in grams) from the number of days since planting (x) using y^=2x+10\hat{y}=2x+10y^​=2x+10. Residuals are residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

Which statement best describes what the residual pattern suggests about the model choice?

  1. Because most residuals are within 1 gram, the model must be appropriate even if the spread changes with x.
  2. The residuals show increasing spread as x increases (a funnel shape), suggesting the linear model may not be equally accurate for all days. (correct answer)
  3. The residuals are random because some are positive and some are negative, so the model is appropriate.
  4. The residuals alternate signs, so the model must be perfect.

Explanation: Evaluating model fit requires examining not just whether residuals are scattered, but also whether their variability remains constant. A residual is the difference between actual and predicted values (y - ŷ), and ideally these should show random scatter with consistent spread across all x-values. When residuals display a funnel shape—starting with small spread and increasing as x increases—this violates the constant variance assumption of linear regression. This pattern suggests the model's accuracy changes with the predictor variable, being more precise for early days and less precise for later days. Such heteroscedasticity (changing variance) indicates the simple linear model may not be equally appropriate across the entire range of data. A common misconception is focusing only on residual magnitude or thinking that alternating signs guarantee a good fit, when the pattern of spread is equally important. The strategy is to check both for systematic patterns in the residuals' center and in their spread.

Question 14

A delivery company models delivery time (y, in minutes) from distance traveled (x, in miles) using y^=4x+15\hat{y}=4x+15y^​=4x+15. Residuals are computed as residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

Which statement best describes how well the model fits the data?

  1. The model is a poor fit because positive residuals mean the model overestimates delivery time.
  2. The model fits well because the residuals are all positive, meaning the model is consistently accurate.
  3. The model is a poor fit because the residuals show increasing spread as distance increases, so the model’s accuracy changes with x. (correct answer)
  4. The model fits well because the residuals are close to 0 at small distances, even if they spread out at larger distances.

Explanation: Evaluating model fit requires examining both the pattern and spread of residuals across the range of x-values. Residuals are calculated as actual minus predicted (y - ŷ), and should ideally maintain consistent variability regardless of x. When residuals show increasing spread as x increases—forming a funnel or megaphone shape—this indicates heteroscedasticity, violating a key assumption of linear regression. This pattern means the model's prediction accuracy depends on the distance: it may predict well for short deliveries but poorly for long ones, making the model unreliable for longer distances. Such changing variance suggests either a transformation is needed or a more complex model that accounts for this variability. A common misconception is thinking that positive residuals indicate good fit or that small residuals at low x-values validate the entire model. The critical insight is that constant variance across all x-values is essential for a properly fitting linear model.

Question 15

An engineer models the output of a machine (y, in units) from the number of hours it has been running since maintenance (x) using the linear model y^=100−2x\hat{y}=100-2xy^​=100−2x. Residuals are defined as residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

What does the residual pattern suggest about the model choice?

  1. The residuals are mostly positive, which means the model is overestimating output most of the time.
  2. Because the residuals are not all 0, the model cannot be used for prediction at all.
  3. The residuals show a curved pattern (positive, then near 0, then negative), suggesting a non-linear model may be more appropriate. (correct answer)
  4. The residuals are randomly scattered around 0, suggesting the linear model is appropriate.

Explanation: Using residuals to evaluate model choice involves identifying patterns that suggest alternative model forms. Residuals, computed as actual minus predicted (y - ŷ), should scatter randomly around 0 for an appropriate linear model. When residuals show a systematic curved pattern—such as starting positive, moving through zero, then becoming negative—this strongly indicates the underlying relationship is non-linear. This pattern suggests the true relationship has curvature that a straight line cannot capture, often occurring when there are diminishing returns or accelerating effects. In this machine output example, the curved residual pattern indicates that output doesn't decrease linearly with time since maintenance—perhaps degradation accelerates over time. A common misconception is thinking any non-zero residuals mean the model is useless; in reality, all models have residuals, but patterns in those residuals guide us toward better model choices. The strategy is recognizing that curved residual patterns specifically point toward polynomial or other non-linear models.

Question 16

A school models the number of absences (y) from the number of school days missed due to illness (x) using y^=1.1x+0.5\hat{y}=1.1x+0.5y^​=1.1x+0.5. Residuals are defined as residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

Which statement best describes how well the model fits the data?

  1. The model fits perfectly because the residuals are both positive and negative.
  2. The model is a poor fit because negative residuals mean the model underestimated absences.
  3. The model is a poor fit because the residuals show a clear curved pattern, suggesting the relationship may not be linear.
  4. The model fits well because the residuals are randomly scattered around 0 with no clear pattern. (correct answer)

Explanation: Assessing model fit with residuals focuses on detecting patterns that would indicate model inadequacy. A residual, calculated as actual minus predicted (y - ŷ), measures the prediction error for each data point. For a well-fitting linear model, these residuals should be randomly scattered around 0, showing no systematic pattern as x changes. Random scatter means the residuals appear unpredictable—sometimes positive, sometimes negative, with no discernible trend or curve. When residuals are randomly distributed, it confirms the linear model appropriately captures the relationship between variables. In this absence prediction model, the residuals show random scatter around 0 with no clear pattern, indicating the linear relationship adequately models how illness-related absences relate to total absences. A common error is thinking that having both positive and negative residuals guarantees perfect fit—what matters is the absence of systematic patterns, not just mixed signs.

Question 17

A city planner models daily water use (y, in thousands of gallons) from the day’s high temperature (x, in °F) with y^=1.5x−40\hat{y}=1.5x-40y^​=1.5x−40. Residuals are computed as residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

Which statement best describes how well the model fits the data?

  1. The model is a poor fit because residuals above 0 mean the actual water use is below the predicted use.
  2. The model fits perfectly because the residuals include both positive and negative values.
  3. The model is a poor fit because the residuals form a curved pattern, suggesting the relationship may not be linear.
  4. The model fits well because the residuals are randomly scattered around 0 with no clear pattern. (correct answer)

Explanation: To assess model fit using residuals, we examine whether they show random scatter or a systematic pattern. Residuals are calculated as actual minus predicted (y - ŷ), so positive residuals mean the model underestimated and negative residuals mean overestimation. A well-fitting linear model produces residuals that are randomly scattered around 0, showing no clear pattern as x increases. When residuals display a pattern like a curve or changing spread, it suggests the linear model is missing important features of the relationship. In this water usage example, the residuals are randomly scattered around 0 with no discernible pattern, indicating the linear model appropriately captures the relationship between temperature and water use. A common error is thinking that having both positive and negative residuals automatically means a perfect fit—what matters is whether they form a pattern. The key strategy is to look for systematic behavior in the residuals, not just their signs or magnitudes.

Question 18

A teacher uses the model y^=5x+50\hat{y}=5x+50y^​=5x+50 to predict a student’s exam score (y) from the number of hours studied (x). Residuals are defined as residual=y−y^\text{residual}=y-\hat{y}residual=y−y^​.

For a student who studied 6 hours, the residual is −4-4−4. What does this residual mean in context?

  1. The model overestimated the student’s score by 4 points. (correct answer)
  2. The student’s predicted score was 4 points higher than the actual score because residuals are computed as y^−y\hat{y}-yy^​−y.
  3. The student studied 4 fewer hours than the model predicted.
  4. The model underestimated the student’s score by 4 points.

Explanation: Understanding residuals requires careful attention to their definition and sign interpretation. A residual is calculated as actual minus predicted (residual = y - ŷ), which tells us how far off the prediction was and in which direction. A negative residual of -4 means the actual value was 4 units less than the predicted value, indicating the model overestimated. In this context, the model predicted the student would score ŷ = 5(6) + 50 = 80 points, but the actual score was only 76 points (since 76 - 80 = -4). This means the model overestimated the student's performance by 4 points. A common misconception is reversing the interpretation of signs or confusing the residual formula—some incorrectly think residuals are ŷ - y, which would flip all the signs. Remember: negative residual means overestimation, positive residual means underestimation, and this interpretation stays consistent regardless of context.

Question 19

A student fits the linear model y^=100−4x\hat{y}=100-4xy^​=100−4x to predict battery life (yyy, in hours) from screen brightness setting (xxx). The residuals (y−y^y-\hat{y}y−y^​) are shown. Which statement best describes how well the model fits the data?

  1. The model is a poor fit because the residuals are all between −1 and 1, so the model is too accurate to be real.
  2. The model is a poor fit because the residuals alternate signs, which always indicates curvature.
  3. The model fits perfectly because negative residuals mean the predictions are exact.
  4. The model fits reasonably well because the residuals are scattered around 0 with no clear pattern. (correct answer)

Explanation: Evaluating model fit with residuals requires distinguishing between random variation and systematic patterns. Residuals represent the difference between actual and predicted values (y - ŷ), with positive values indicating underestimation and negative indicating overestimation. A well-fitting model produces residuals that scatter randomly around 0, like leaves falling from a tree with no predictable pattern—some positive, some negative, with no systematic structure. In this battery life example, the residuals appear to be randomly distributed around 0 without showing curves, trends, or changing spread, suggesting the linear model adequately captures the relationship. A common misconception is that alternating signs in residuals always indicate a problem, but random alternation is actually expected in good fit—it's only problematic when the alternation follows a predictable pattern. The key strategy is to look for randomness in the residual plot: truly random scatter supports the model choice, while any discernible pattern suggests the need for model revision.

Question 20

A linear model is used to predict a runner’s heart rate (yyy, beats per minute) from running speed (xxx, mph): y^=60+8x\hat{y}=60+8xy^​=60+8x. Residuals (y−y^y-\hat{y}y−y^​) for several speeds are shown. Which statement best describes how well the model fits the data?

  1. The model fits well because the residuals get larger at higher speeds, which is expected for a good model.
  2. The model fits well because the residuals are all positive, so the model is consistent.
  3. The model must fit well because the residuals are centered above 0.
  4. The model is a poor fit because the residual spread increases as xxx increases (a funnel shape). (correct answer)

Explanation: This question examines a specific type of poor model fit revealed by residuals—changing variability or heteroscedasticity. Residuals equal actual minus predicted values (y - ŷ), and ideally should scatter randomly around 0 with consistent spread across all x-values. When residuals show a 'funnel shape'—small spread at low x-values that increases at higher x-values—it indicates the model's prediction accuracy deteriorates as x increases. In this heart rate example, the increasing residual spread suggests that heart rate becomes more variable at higher running speeds, which the linear model cannot capture. This violates the assumption of constant variance that underlies linear regression. A common misconception is thinking larger residuals at higher values are expected or acceptable, but good models should maintain consistent prediction accuracy across the entire range. The strategy for detecting this issue is to examine not just whether residuals center around 0, but whether their spread remains constant—widening spread signals the need for transformation or a different modeling approach.