College Statistics Quiz: Linearizing Transformations
20 questions · exam conditions
0:00
Linearizing TransformationsQuestion 1 of 20

A researcher fits the model y^=3.2+1.5x\widehat{\sqrt{y}} = 3.2 + 1.5x. What is the predicted value of yy when x=4x=4?

84.64
9.2
6.0
36.0
← Back to quizzes

College Statistics Quiz

College Statistics Quiz: Linearizing Transformations

Practice Linearizing Transformations in College Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Linearizing Transformations, giving you a quick way to practice the rules, question types, and explanations that matter most for College Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A researcher fits the model y^=3.2+1.5x\widehat{\sqrt{y}} = 3.2 + 1.5x. What is the predicted value of yy when x=4x=4?

  1. 84.64 (correct answer)
  2. 9.2
  3. 6.0
  4. 36.0
Explanation: When you encounter a regression model where the dependent variable has been transformed, you must work backwards through the transformation to find the actual predicted value. Here, the model predicts y^=3.2+1.5x\widehat{\sqrt{y}} = 3.2 + 1.5x, meaning the square root of y is being predicted, not y itself. When x=4x = 4, you first calculate: y^=3.2+1.5(4)=3.2+6.0=9.2\widehat{\sqrt{y}} = 3.2 + 1.5(4) = 3.2 + 6.0 = 9.2 This gives you the predicted value of y\sqrt{y}, which is 9.2. To find the predicted value of y, you must square both sides: y^=(9.2)2=84.64\widehat{y} = (9.2)^2 = 84.64 Looking at the wrong answers: Choice B (9.2) represents the predicted value of y\sqrt{y}, not y itself—this is the most common error students make with transformed variables. Choice C (6.0) is just the coefficient times x (1.5 × 4), ignoring the intercept entirely. Choice D (36.0) might result from incorrectly squaring just the coefficient term (6²) rather than the full predicted value. The correct answer is A (84.64) because you must complete the full transformation process. Study tip: With transformed dependent variables, always ask yourself "What is actually being predicted by the equation?" Then work backwards through the transformation. If the model predicts y\sqrt{y}, ln(y)\ln(y), or 1y\frac{1}{y}, you must apply the inverse transformation (squaring, exponentiating, or taking the reciprocal) to get your final answer for y.

Question 2

To model the relationship between the length of a fish (L) and its weight (W), a researcher fits a power model and obtains the regression equation log(W)^=2.1+3.05log(L)\widehat{\log(W)} = -2.1 + 3.05 \log(L). What is the correct interpretation of the coefficient 3.05?

  1. For each 1 cm increase in length, the fish's weight is predicted to increase by 3.05 grams.
  2. For each 1% increase in length, the fish's weight is predicted to increase by approximately 3.05%. (correct answer)
  3. For each 1 cm increase in length, the logarithm of the fish's weight is predicted to increase by 3.05.
  4. The model predicts that a fish of length 1 cm will have a weight of 3.05 grams.
Explanation: In a log-log regression model (power model), the slope represents the elasticity between the variables. A slope of 3.05 means that for a 1% increase in the explanatory variable (Length), the response variable (Weight) is predicted to increase by approximately 3.05%. This is a key interpretive difference between log-log models and other forms.

Question 3

A regression of the natural log of CEO salary on the natural log of company revenue yields the equation ln(salary)^=2.8+0.3ln(revenue)\widehat{\ln(\text{salary})} = 2.8 + 0.3 \ln(\text{revenue}). What is the predicted salary for a company with a revenue of e10e^{10} million dollars? Note that ln(e10)=10\ln(e^{10}) = 10.

  1. 5.85.8 million dollars
  2. e2.8+e3e^{2.8} + e^{3} million dollars
  3. e30.8e^{30.8} million dollars
  4. e5.8e^{5.8} million dollars (correct answer)
Explanation: First, find the value of the predictor variable, ln(revenue)\ln(\text{revenue}), which is given as 10. Substitute this into the regression equation: ln(salary)^=2.8+0.3(10)=2.8+3=5.8\widehat{\ln(\text{salary})} = 2.8 + 0.3(10) = 2.8 + 3 = 5.8. This is the predicted value on the log scale. To find the predicted salary, we must back-transform by taking the exponent: salary^=e5.8\widehat{\text{salary}} = e^{5.8}.

Question 4

An engineer models the cooling of a liquid over time. The relationship is expected to be exponential. After fitting a linear model to the transformed data, she obtains the equation ln(Temp)^=4.20.05Time\widehat{\ln(Temp)} = 4.2 - 0.05 \cdot \text{Time}. What is the predicted temperature at Time = 10 minutes?

  1. e3.740.4e^{3.7} \approx 40.4 degrees (correct answer)
  2. 4.20.05(10)=3.74.2 - 0.05(10) = 3.7 degrees
  3. e4.2e0.565.1e^{4.2} - e^{0.5} \approx 65.1 degrees
  4. ln(4.20.5)1.3\ln(4.2 - 0.5) \approx 1.3 degrees
Explanation: First, substitute Time = 10 into the regression equation to find the predicted value of the transformed variable: ln(Temp)^=4.20.05(10)=4.20.5=3.7\widehat{\ln(Temp)} = 4.2 - 0.05(10) = 4.2 - 0.5 = 3.7. This is the predicted natural log of the temperature. To find the predicted temperature, we must perform the inverse operation, which is exponentiation: Temp^=e3.740.4\widehat{Temp} = e^{3.7} \approx 40.4. Choice B is a common error where the student forgets to back-transform.

Question 5

An ecologist observes that the number of species in a particular ecosystem seems to grow rapidly at first and then level off as the area of the ecosystem increases. A scatterplot of the number of species (y) versus area (x) confirms this curvilinear pattern. Which of the following transformations is most likely to create a linear relationship?

  1. Plotting log(y)\log(y) against xx
  2. Plotting yy against log(x)\log(x) (correct answer)
  3. Plotting log(y)\log(y) against log(x)\log(x)
  4. Plotting yy against 1/x1/x
Explanation: The described pattern, rapid growth followed by a leveling off, is characteristic of a logarithmic relationship, modeled by an equation of the form y=a+blog(x)y = a + b \log(x). To linearize this relationship, one should plot the original response variable yy against the logarithm of the explanatory variable, log(x)\log(x).

Question 6

A physicist investigates the relationship between the pressure (P) and volume (V) of a gas at constant temperature. The data follows a curve consistent with Boyle's Law, which states that pressure and volume are inversely proportional. Which of the following scatterplots would most likely display a linear pattern?

  1. A plot of log(P)\log(P) versus VV
  2. A plot of PP versus log(V)\log(V)
  3. A plot of log(P)\log(P) versus log(V)\log(V)
  4. A plot of PP versus 1/V1/V (correct answer)
Explanation: Boyle's Law is P1/VP \propto 1/V, or P=k/VP = k/V for some constant k. This is a reciprocal relationship. To linearize this equation, we can plot PP against a new explanatory variable x=1/Vx = 1/V. The resulting relationship, P=kxP = kx, is linear with an intercept of 0. Therefore, a plot of PP versus 1/V1/V would be linear.

Question 7

A biologist models the growth of a bacterial colony using an exponential model, N=abtN = ab^t, where NN is the number of bacteria and tt is time in hours. After transforming the data, she fits a least-squares regression line and obtains the equation log10(N)^=2.3+0.15t\widehat{\log_{10}(N)} = 2.3 + 0.15t. Which of the following is the approximate exponential model for the bacterial growth?

  1. N^=2.3(1.41)t\hat{N} = 2.3(1.41)^t
  2. N^=199.5(1.41)t\hat{N} = 199.5(1.41)^t (correct answer)
  3. N^=2.3(0.15)t\hat{N} = 2.3(0.15)^t
  4. N^=199.5(0.15)t\hat{N} = 199.5(0.15)^t
Explanation: The transformation for an exponential model is log(N)=log(a)+tlog(b)\log(N) = \log(a) + t \log(b). Comparing this to the fitted equation log10(N)^=2.3+0.15t\widehat{\log_{10}(N)} = 2.3 + 0.15t, we have log10(a)=2.3\log_{10}(a) = 2.3 and log10(b)=0.15\log_{10}(b) = 0.15. To find aa and bb, we must back-transform: a=102.3199.5a = 10^{2.3} \approx 199.5 and b=100.151.41b = 10^{0.15} \approx 1.41. Thus, the model is N^=199.5(1.41)t\hat{N} = 199.5(1.41)^t.

Question 8

A dataset with a strong curved relationship has a correlation coefficient of r=0.78r = 0.78. An appropriate linearizing transformation is applied to the data, and a new linear model is fit. Which of the following would be the most plausible correlation coefficient for the regression on the transformed data?

  1. 0.65
  2. -0.78
  3. 0.97 (correct answer)
  4. 1.02
Explanation: A successful linearizing transformation strengthens the linear association between the variables. This means the absolute value of the correlation coefficient, r|r|, should increase and move closer to 1. Since the original correlation was positive (0.78), the new correlation should be a positive value closer to 1. 0.97 is the only plausible option. 0.65 would mean the linear relationship weakened. -0.78 would mean the direction of the relationship changed, which is unlikely for standard transformations. 1.02 is an impossible value for r.

Question 9

A regression of house price (in thousands of dollars) on square footage produces a residual plot that shows a distinct fan shape, with the spread of the residuals increasing as the fitted values increase. What is the primary issue this indicates and a common remedy?

  1. Non-linearity; add a quadratic term (square footage squared) to the model.
  2. Heteroscedasticity; perform a logarithmic transformation on the response variable (price). (correct answer)
  3. Influential outliers; identify and remove high-leverage points from the dataset.
  4. Collinearity; this cannot be determined without knowing the other predictor variables in the model.
Explanation: A fan-shaped residual plot indicates heteroscedasticity, which is the condition of non-constant variance in the errors. This violates a key assumption of linear regression. A common way to remedy this is to transform the response variable to stabilize the variance. The logarithmic transformation is often effective when the standard deviation of the residuals is proportional to the fitted values.

Question 10

A researcher is comparing two models to predict crop yield (Y) from the amount of fertilizer (F). Model A is a simple linear regression of Y on F, yielding R2=0.75R^2 = 0.75 and a curved residual plot. Model B is a regression of log(Y)\log(Y) on F, yielding R2=0.89R^2 = 0.89 and a residual plot with random scatter. Which of the following is the most valid conclusion?

  1. Model B is superior primarily because its R2R^2 value is significantly higher.
  2. The models cannot be compared because the response variables (Y and log(Y)\log(Y)) are on different scales.
  3. Model B is superior because its residual plot indicates that the conditions for linear regression are better satisfied. (correct answer)
  4. Model A is superior because its interpretation is more straightforward and it explains 75% of the variation in yield.
Explanation: While a higher R2R^2 is desirable, the primary diagnostic tool for model adequacy is the residual plot. The curved residual plot for Model A indicates that a linear model is not appropriate for the original data. The random scatter in the residuals for Model B indicates that the transformed model is a much better fit for the data structure and better satisfies the assumptions of linear regression. Therefore, the state of the residuals is the strongest reason to prefer Model B.

Question 11

A city's population growth is modeled over several decades. An analyst considers two models. Model A (Linear): Fits Population vs. Year. Result: Strong positive correlation, but the residual plot shows a clear curve, indicating the growth rate is increasing over time. Model B (Exponential): Fits log(Population) vs. Year. Result: The scatterplot of the transformed data is highly linear, and the residual plot shows no discernible pattern.

Based on the analyst's findings in the passage, which statement is the most statistically sound?

  1. Model B is preferable because the lack of a pattern in its residual plot suggests it better captures the underlying relationship. (correct answer)
  2. Model A is preferable because it directly predicts population without the need for back-transformation.
  3. Both models are equally valid, and the choice depends on whether the analyst prefers a simpler or more complex model.
  4. Neither model is useful until further analysis is done to check for outliers in the original dataset.
Explanation: When evaluating regression models, residual analysis is your primary tool for determining which model best captures the underlying relationship in your data. Residuals are the differences between observed and predicted values, and their patterns reveal whether your model assumptions are met. Model A shows a curved pattern in the residuals despite having a strong correlation. This curvature indicates systematic error—the linear model consistently under-predicts at certain points and over-predicts at others because population growth is actually accelerating over time, not increasing at a constant rate. A good model should have residuals that appear randomly scattered with no discernible pattern. Model B transforms the data by taking the logarithm of population, then fits a linear relationship. The resulting residual plot shows no pattern, indicating the exponential model (which becomes linear after log transformation) properly captures the underlying growth dynamics. Choice A is correct because the lack of pattern in Model B's residuals demonstrates it better represents the true relationship. Choice B is wrong because convenience of interpretation doesn't outweigh model validity—you can easily back-transform exponential predictions. Choice C is incorrect because the models aren't equally valid; residual analysis clearly shows one fits better than the other. Choice D is wrong because outlier analysis, while important, doesn't address the fundamental issue that Model A fails to capture the exponential nature of the growth. Remember: always examine residual plots when comparing models. Random scatter means good fit; patterns mean your model is missing something important about the relationship.

Question 12

After attempting to fit a linear model to data relating variables XX and YY, an analyst produces a residual plot. The plot shows that the residuals are all positive for small and large values of XX and all negative for intermediate values of XX. Which re-expression of the data is most likely to produce a more linear relationship?

  1. Fitting YY versus log(X)\log(X).
  2. Fitting log(Y)\log(Y) versus XX.
  3. Fitting YY versus X2X^2. (correct answer)
  4. Fitting YY versus 1/X1/X.
Explanation: The described residual pattern (positive-negative-positive) is a classic 'U' or 'parabolic' shape. This indicates that the original data has a parabolic relationship that is not captured by the linear model. To linearize such a relationship, one can fit the response variable YY against the square of the explanatory variable, X2X^2. This transforms the model from Y=b0+b1XY = b_0 + b_1 X to Y=b0+b1X2Y = b_0 + b_1 X^2, which is a linear model in the variable X2X^2.

Question 13

A regression model is fit to transformed data: ln(y)^=5.00.2x\widehat{\ln(y)} = 5.0 - 0.2x. Which of the following statements is a correct interpretation of the intercept?

  1. When x=0x=0, the predicted value of yy is 5.0.
  2. When ln(y)=0\ln(y)=0, the predicted value of xx is 25.
  3. When x=0x=0, the predicted value of yy is e5e^5. (correct answer)
  4. The model is invalid because the intercept cannot be interpreted in a transformed model.
Explanation: The intercept in a regression model is the predicted value of the response variable when the explanatory variable is zero. In this case, the response variable is ln(y)\ln(y). So, when x=0x=0, the predicted value of ln(y)\ln(y) is 5.0. To find the predicted value of yy, we must back-transform this value: y^=e5\hat{y} = e^5.

Question 14

Which of the following is a potential disadvantage of using transformations to linearize a relationship?

  1. The correlation coefficient of the transformed data will always be lower than the original data's correlation.
  2. Transformations make it impossible to use the model for prediction.
  3. The interpretation of the model's coefficients becomes less direct and must be considered in the context of the transformed scale. (correct answer)
  4. Applying a transformation always violates the assumption of normally distributed residuals.
Explanation: While transformations can greatly improve a model's fit and adherence to assumptions, they come at the cost of interpretability. A slope in a simple linear model represents the change in Y for a one-unit change in X. In a transformed model (e.g., log-log or semi-log), the slope represents a multiplicative or percentage change, which is less intuitive. Predictions also require a back-transformation step. The other choices are incorrect: transformations should increase correlation, are used for prediction, and can often help satisfy the normality assumption.

Question 15

An ecologist observes that the number of species in a particular ecosystem seems to grow rapidly at first and then level off as the area of the ecosystem increases. A scatterplot of the number of species (y) versus area (x) confirms this curvilinear pattern. Which of the following transformations is most likely to create a linear relationship?

  1. Plotting log(y)\log(y) against xx
  2. Plotting yy against log(x)\log(x) (correct answer)
  3. Plotting log(y)\log(y) against log(x)\log(x)
  4. Plotting yy against 1/x1/x
Explanation: The described pattern, rapid growth followed by a leveling off, is characteristic of a logarithmic relationship, modeled by an equation of the form y=a+blog(x)y = a + b \log(x). To linearize this relationship, one should plot the original response variable yy against the logarithm of the explanatory variable, log(x)\log(x).

Question 16

A physicist investigates the relationship between the pressure (P) and volume (V) of a gas at constant temperature. The data follows a curve consistent with Boyle's Law, which states that pressure and volume are inversely proportional. Which of the following scatterplots would most likely display a linear pattern?

  1. A plot of log(P)\log(P) versus VV
  2. A plot of PP versus log(V)\log(V)
  3. A plot of log(P)\log(P) versus log(V)\log(V)
  4. A plot of PP versus 1/V1/V (correct answer)
Explanation: Boyle's Law is P1/VP \propto 1/V, or P=k/VP = k/V for some constant k. This is a reciprocal relationship. To linearize this equation, we can plot PP against a new explanatory variable x=1/Vx = 1/V. The resulting relationship, P=kxP = kx, is linear with an intercept of 0. Therefore, a plot of PP versus 1/V1/V would be linear.

Question 17

The relationship between a star's luminosity (L) and its mass (M) can be described by a power model of the form L=aMpL = aM^p. An astronomer performs a linear regression on the logarithms of the data, resulting in the fitted line ln(L)^=0.12+3.5ln(M)\widehat{\ln(L)} = 0.12 + 3.5 \ln(M). Based on this model, what is the predicted luminosity of a star with a mass of M=2M=2 solar masses? Use ln(2)0.693\ln(2) \approx 0.693.

  1. 2.546
  2. 12.75 (correct answer)
  3. 7.24
  4. 11.02
Explanation: First, use the regression equation to predict ln(L)\ln(L). ln(L)^=0.12+3.5ln(2)0.12+3.5(0.693)=0.12+2.4255=2.5455\widehat{\ln(L)} = 0.12 + 3.5 \ln(2) \approx 0.12 + 3.5(0.693) = 0.12 + 2.4255 = 2.5455. This is the predicted natural log of the luminosity. To find the predicted luminosity L^\hat{L}, we must back-transform: L^=e2.545512.75\hat{L} = e^{2.5455} \approx 12.75. The distractor 2.546 is the result of forgetting to back-transform.

Question 18

A dataset with a strong curved relationship has a correlation coefficient of r=0.78r = 0.78. An appropriate linearizing transformation is applied to the data, and a new linear model is fit. Which of the following would be the most plausible correlation coefficient for the regression on the transformed data?

  1. 0.65
  2. -0.78
  3. 0.97 (correct answer)
  4. 1.02
Explanation: A successful linearizing transformation strengthens the linear association between the variables. This means the absolute value of the correlation coefficient, r|r|, should increase and move closer to 1. Since the original correlation was positive (0.78), the new correlation should be a positive value closer to 1. 0.97 is the only plausible option. 0.65 would mean the linear relationship weakened. -0.78 would mean the direction of the relationship changed, which is unlikely for standard transformations. 1.02 is an impossible value for r.

Question 19

To model the relationship between the length of a fish (L) and its weight (W), a researcher fits a power model and obtains the regression equation log(W)^=2.1+3.05log(L)\widehat{\log(W)} = -2.1 + 3.05 \log(L). What is the correct interpretation of the coefficient 3.05?

  1. For each 1 cm increase in length, the fish's weight is predicted to increase by 3.05 grams.
  2. For each 1% increase in length, the fish's weight is predicted to increase by approximately 3.05%. (correct answer)
  3. For each 1 cm increase in length, the logarithm of the fish's weight is predicted to increase by 3.05.
  4. The model predicts that a fish of length 1 cm will have a weight of 3.05 grams.
Explanation: In a log-log regression model (power model), the slope represents the elasticity between the variables. A slope of 3.05 means that for a 1% increase in the explanatory variable (Length), the response variable (Weight) is predicted to increase by approximately 3.05%. This is a key interpretive difference between log-log models and other forms.

Question 20

An engineer models the cooling of a liquid over time. The relationship is expected to be exponential. After fitting a linear model to the transformed data, she obtains the equation ln(Temp)^=4.20.05Time\widehat{\ln(Temp)} = 4.2 - 0.05 \cdot \text{Time}. What is the predicted temperature at Time = 10 minutes?

  1. e3.740.4e^{3.7} \approx 40.4 degrees (correct answer)
  2. 4.20.05(10)=3.74.2 - 0.05(10) = 3.7 degrees
  3. e4.2e0.565.1e^{4.2} - e^{0.5} \approx 65.1 degrees
  4. ln(4.20.5)1.3\ln(4.2 - 0.5) \approx 1.3 degrees
Explanation: First, substitute Time = 10 into the regression equation to find the predicted value of the transformed variable: ln(Temp)^=4.20.05(10)=4.20.5=3.7\widehat{\ln(Temp)} = 4.2 - 0.05(10) = 4.2 - 0.5 = 3.7. This is the predicted natural log of the temperature. To find the predicted temperature, we must perform the inverse operation, which is exponentiation: Temp^=e3.740.4\widehat{Temp} = e^{3.7} \approx 40.4. Choice B is a common error where the student forgets to back-transform.