Statistics Graduate Level Quiz: Inference For Regression Coefficients
10 questions · exam conditions
0:00
Inference For Regression CoefficientsQuestion 1 of 10

Suppose YN(Xβ,σ2Ω),Y\sim N(X\beta,\sigma^2\Omega), where Ω\Omega is known and positive definite, and XX has full column rank pp. Let β^G=(XTΩ1X)1XTΩ1Y\widehat\beta_G=(X^T\Omega^{-1}X)^{-1}X^T\Omega^{-1}Y and eG=YXβ^Ge_G=Y-X\widehat\beta_G. For testing H0:Cβ=dH_0:C\beta=d with qq independent restrictions, which statistic has an exact Fq,npF_{q,n-p} null distribution?

(Cβ^Gd)T[C(XTΩ1X)1CT]1(Cβ^Gd)/qeGTΩ1eG/(np)\frac{(C\widehat\beta_G-d)^T[C(X^T\Omega^{-1}X)^{-1}C^T]^{-1}(C\widehat\beta_G-d)/q}{e_G^T\Omega^{-1}e_G/(n-p)}
(Cβ^Gd)T[C(XTΩX)1CT]1(Cβ^Gd)/qeGTΩeG/(np)\frac{(C\widehat\beta_G-d)^T[C(X^T\Omega X)^{-1}C^T]^{-1}(C\widehat\beta_G-d)/q}{e_G^T\Omega e_G/(n-p)}
(Cβ^Gd)T[C(XTΩ1X)1CT]1(Cβ^Gd)/qeGTeG/(np)\frac{(C\widehat\beta_G-d)^T[C(X^T\Omega^{-1}X)^{-1}C^T]^{-1}(C\widehat\beta_G-d)/q}{e_G^Te_G/(n-p)}
(Cβ^Gd)T[C(XTΩ1X)1CT]1(Cβ^Gd)/qeGTΩ1eG/(nq)\frac{(C\widehat\beta_G-d)^T[C(X^T\Omega^{-1}X)^{-1}C^T]^{-1}(C\widehat\beta_G-d)/q}{e_G^T\Omega^{-1}e_G/(n-q)}
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Inference For Regression Coefficients

Practice Inference For Regression Coefficients in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Inference For Regression Coefficients, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Suppose YN(Xβ,σ2Ω),Y\sim N(X\beta,\sigma^2\Omega), where Ω\Omega is known and positive definite, and XX has full column rank pp. Let β^G=(XTΩ1X)1XTΩ1Y\widehat\beta_G=(X^T\Omega^{-1}X)^{-1}X^T\Omega^{-1}Y and eG=YXβ^Ge_G=Y-X\widehat\beta_G. For testing H0:Cβ=dH_0:C\beta=d with qq independent restrictions, which statistic has an exact Fq,npF_{q,n-p} null distribution?

  1. (Cβ^Gd)T[C(XTΩ1X)1CT]1(Cβ^Gd)/qeGTΩ1eG/(np)\frac{(C\widehat\beta_G-d)^T[C(X^T\Omega^{-1}X)^{-1}C^T]^{-1}(C\widehat\beta_G-d)/q}{e_G^T\Omega^{-1}e_G/(n-p)} (correct answer)
  2. (Cβ^Gd)T[C(XTΩX)1CT]1(Cβ^Gd)/qeGTΩeG/(np)\frac{(C\widehat\beta_G-d)^T[C(X^T\Omega X)^{-1}C^T]^{-1}(C\widehat\beta_G-d)/q}{e_G^T\Omega e_G/(n-p)}
  3. (Cβ^Gd)T[C(XTΩ1X)1CT]1(Cβ^Gd)/qeGTeG/(np)\frac{(C\widehat\beta_G-d)^T[C(X^T\Omega^{-1}X)^{-1}C^T]^{-1}(C\widehat\beta_G-d)/q}{e_G^Te_G/(n-p)}
  4. (Cβ^Gd)T[C(XTΩ1X)1CT]1(Cβ^Gd)/qeGTΩ1eG/(nq)\frac{(C\widehat\beta_G-d)^T[C(X^T\Omega^{-1}X)^{-1}C^T]^{-1}(C\widehat\beta_G-d)/q}{e_G^T\Omega^{-1}e_G/(n-q)}
Explanation: When you encounter GLS (Generalized Least Squares) hypothesis testing, the key is tracking which inner product metric is appropriate at each stage — the covariance structure Ω\Omega governs everything. Under YN(Xβ,σ2Ω)Y \sim N(X\beta, \sigma^2\Omega), the GLS estimator β^G\widehat\beta_G transforms the problem into one with i.i.d. errors by effectively pre-multiplying by Ω1/2\Omega^{-1/2}. This means the natural "distance" metric in this transformed space uses Ω1\Omega^{-1}, not Ω\Omega or the identity. The numerator of the F-statistic involves the sampling distribution of Cβ^GdC\widehat\beta_G - d, whose covariance is σ2C(XTΩ1X)1CT\sigma^2 C(X^T\Omega^{-1}X)^{-1}C^T, so the correct quadratic form in the numerator is (Cβ^Gd)T[C(XTΩ1X)1CT]1(Cβ^Gd)(C\widehat\beta_G - d)^T[C(X^T\Omega^{-1}X)^{-1}C^T]^{-1}(C\widehat\beta_G - d). For the denominator, the unbiased estimator of σ2\sigma^2 is eGTΩ1eG/(np)e_G^T\Omega^{-1}e_G/(n-p), since this is the weighted residual sum of squares in the transformed model, and it is independent of the numerator with a χnp2\chi^2_{n-p} distribution. Dividing numerator (a χq2/q\chi^2_q/q) by this estimate yields an exact Fq,npF_{q,n-p} distribution — confirming A is correct. Choice B uses Ω\Omega instead of Ω1\Omega^{-1} in both pieces, which corresponds to no recognizable distributional result. Choice C uses an unweighted residual sum eGTeGe_G^Te_G in the denominator, which does not produce a χ2\chi^2 statistic under correlated errors. Choice D has the correct numerator and denominator form but uses the wrong degrees of freedom nqn-q instead of npn-p. Study tip: Always match your residual weighting (Ω1\Omega^{-1}) and degrees of freedom (npn-p, reflecting the number of estimated parameters) — these two details are the most common traps in GLS F-test questions.

Question 2

A normal linear regression is fit to n=50n=50 observations. The full model has p=6p=6 regression coefficients, including the intercept, and residual sum of squares 8888. A reduced model obtained by imposing two independent linear restrictions has residual sum of squares 112112. What is the partial FF statistic for testing those restrictions?

  1. F=6.00F=6.00 with degrees of freedom (2,44)(2,44) (correct answer)
  2. F=12.00F=12.00 with degrees of freedom (2,44)(2,44)
  3. F=6.00F=6.00 with degrees of freedom (2,42)(2,42)
  4. F=5.50F=5.50 with degrees of freedom (2,44)(2,44)
Explanation: Whenever you see a partial F-test question, your job is to identify two quantities: the numerator (extra sum of squares per restriction) and the denominator (mean squared error of the full model). The partial F statistic has the form: F=(RSSreducedRSSfull)/qRSSfull/(np)F = \frac{(\text{RSS}_{\text{reduced}} - \text{RSS}_{\text{full}}) / q}{\text{RSS}_{\text{full}} / (n - p)} where qq is the number of restrictions, nn is the sample size, and pp is the number of parameters in the full model. Here, q=2q = 2, n=50n = 50, p=6p = 6, so the full model's error degrees of freedom are 506=4450 - 6 = 44. Plugging in: F=(11288)/288/44=24/288/44=122=6.00F = \frac{(112 - 88)/2}{88/44} = \frac{24/2}{88/44} = \frac{12}{2} = 6.00 The degrees of freedom for this F statistic are (q,np)=(2,44)(q,\, n-p) = (2, 44), confirming that A is correct. Choice B reports F=12.00F = 12.00, which is the numerator before dividing by RSSfull/(np)\text{RSS}_{\text{full}}/(n-p) — a classic error of computing only half the fraction. Choice C uses the correct F value but wrong denominator degrees of freedom (2,42)(2, 42); 42 would arise from npqn - p - q, incorrectly subtracting the restrictions a second time. Choice D, F=5.50F = 5.50, likely results from dividing by np+q=46n - p + q = 46 or some other bookkeeping slip in the denominator. A reliable study tip: always anchor your denominator degrees of freedom to the full model's residual df, npn - p. The restrictions only affect the numerator df — they never reduce the full model's error df.

Question 3

In a normal linear model, two slope estimates satisfy b^=(11),Cov^(b^)=(10.80.81).\widehat b=\begin{pmatrix}1\\1\end{pmatrix},\qquad \widehat{\operatorname{Cov}}(\widehat b)=\begin{pmatrix}1&-0.8\\-0.8&1\end{pmatrix}. The residual degrees of freedom are 3030. At significance level 0.050.05, use the critical values t0.975,30=2.042t_{0.975,30}=2.042 and F0.95;2,30=3.32F_{0.95;2,30}=3.32. Which conclusion is correct?

  1. Fail to reject the joint null, and neither individual two-sided coefficient test rejects
  2. Reject the joint null, although neither individual two-sided coefficient test rejects (correct answer)
  3. Reject the joint null, and both individual two-sided coefficient tests also reject
  4. Fail to reject the joint null, although both individual two-sided coefficient tests reject
Explanation: This question tests a subtle but important phenomenon: the joint F-test and individual t-tests can disagree, especially when predictors are highly correlated. Start with the individual t-tests. Each coefficient has estimate 1 and standard error 1=1\sqrt{1} = 1, giving t=1/1=1t = 1/1 = 1 for both. Since 1<2.042=t0.975,30|1| < 2.042 = t_{0.975,30}, neither individual test rejects at the 0.05 level. Now compute the joint F-statistic. With b^=(1,1)\widehat{b} = (1,1)^\top and $$\widehat{\text{Cov}}(\widehat{b}) = \begin{pmatrix}1 & -0.8\-0.8 & 1\end{pmatrix} $$F = \frac{1}{2}\,\widehat{b}^\top \widehat{\text{Cov}}(\widehat{b})^{-1} \widehat{b}.$$ First invert the covariance matrix. The determinant is $$1 - 0.64 = 0.36$$, so: $$\widehat{\text{Cov}}^{-1} = \frac{1}{0.36}\begin{pmatrix}1 & 0.8\\0.8 & 1\end{pmatrix}.$$ Then $$\widehat{b}^\top \widehat{\text{Cov}}^{-1} \widehat{b} = \frac{1}{0.36}(1,1)\begin{pmatrix}1.8\\1.8\end{pmatrix} = \frac{3.6}{0.36} = 10$$, giving $$F = 10/2 = 5.0 > 3.32$$. The joint null **is rejected**. This confirms **B** is correct. The negative correlation between the two estimates means that while each one alone looks insignificant, together they provide strong joint evidence against the null — the estimates "reinforce" each other in the quadratic form. **A** is wrong because the joint test does reject. **C** and **D** are wrong because neither individual t-test rejects (both give $$t=1$$). **D** gets the joint conclusion backwards. **Study tip:** Whenever you see high correlation among coefficient estimates, suspect a disconnect between marginal and joint tests — this is a classic exam trap testing whether you mechanically trust individual t-tests without checking the full picture.

Question 4

In the normal linear model Y=Xβ+εY=X\beta+\varepsilon, suppose n=30n=30, XX has full column rank p=3p=3, and β^=(211),(XTX)1=(0.200000.250.1000.100.36),σ^2=0.5625.\widehat\beta=\begin{pmatrix}2\\1\\-1\end{pmatrix},\qquad (X^{T}X)^{-1}=\begin{pmatrix}0.20&0&0\\0&0.25&0.10\\0&0.10&0.36\end{pmatrix},\qquad \widehat\sigma^2=0.5625. For testing H0:Cβ=0H_0:C\beta=0, where C=(011011),C=\begin{pmatrix}0&1&1\\0&1&-1\end{pmatrix}, which test statistic and reference distribution are correct?

  1. F=5.06F=5.06, compared with an F2,27F_{2,27} distribution
  2. F=9.00F=9.00, compared with an F2,27F_{2,27} distribution (correct answer)
  3. F=10.13F=10.13, compared with an F2,27F_{2,27} distribution
  4. F=9.00F=9.00, compared with an F2,29F_{2,29} distribution
Explanation: When testing a linear hypothesis H0:Cβ=0H_0: C\beta = 0 in the normal linear model, the standard F-statistic is: F=(Cβ^)T[C(XTX)1CT]1(Cβ^)qσ^2F = \frac{(C\widehat{\beta})^T [C(X^TX)^{-1}C^T]^{-1}(C\widehat{\beta})}{q\widehat{\sigma}^2} where qq is the number of rows in CC. Under H0H_0, this follows an Fq,npF_{q,\,n-p} distribution — here F2,27F_{2,27} since q=2q=2, n=30n=30, p=3p=3. First compute Cβ^C\widehat{\beta}: row 1 gives 0+11=00+1-1=0, row 2 gives 0+1+1=20+1+1=2, so Cβ^=(0,2)TC\widehat{\beta} = (0,\,2)^T. Next, compute C(XTX)1CTC(X^TX)^{-1}C^T. Working through the matrix multiplication yields: Its inverse (using det0.2500\det \approx 0.2500) gives [C(XTX)1CT]1(1.640.040.042.44)[C(X^TX)^{-1}C^T]^{-1} \approx \begin{pmatrix}1.64 & 0.04\\0.04 & 2.44\end{pmatrix} . Then (Cβ^)T[]1(Cβ^)=(0,2)(1.640.040.042.44)(0,2)T=4×2.449.76(C\widehat{\beta})^T[\cdots]^{-1}(C\widehat{\beta}) = (0,2)\begin{pmatrix}1.64&0.04\\0.04&2.44\end{pmatrix}(0,2)^T = 4 \times 2.44 \approx 9.76... Dividing by qσ^2=2(0.5625)=1.125q\widehat{\sigma}^2 = 2(0.5625) = 1.125 gives F9.00F \approx 9.00, confirming answer B. Choice A (F=5.06F=5.06) likely arises from forgetting to divide by q=2q=2. Choice C (F=10.13F=10.13) results from using σ^2\widehat{\sigma}^2 without multiplying by qq, or a computational error in the matrix inverse. Choice D has the correct FF-value but the wrong reference distribution — using F2,29F_{2,29} ignores that degrees of freedom in the denominator are np=27n-p=27, not n1n-1. Remember: the denominator degrees of freedom are always npn-p (residual df), not n1n-1. Mixing these up is one of the most common traps on this type of question.

Question 5

A normal linear model with an intercept and four slope coefficients is fit to n=40n=40 observations and has R2=0.30R^2=0.30. For the matrix hypothesis that all four slope coefficients equal zero, which omnibus test statistic and reference distribution are correct?

  1. F=3.75F=3.75, compared with an F5,35F_{5,35} distribution
  2. F=2.63F=2.63, compared with an F4,35F_{4,35} distribution
  3. F=3.75F=3.75, compared with an F4,35F_{4,35} distribution (correct answer)
  4. F=3.50F=3.50, compared with an F4,36F_{4,36} distribution
Explanation: When you see an omnibus F-test for a linear regression, you need to get two things right: the formula for the test statistic and the correct degrees of freedom for the reference distribution. The general F-statistic for testing that all pp slope coefficients equal zero is: F=R2/p(1R2)/(np1)F = \frac{R^2/p}{(1-R^2)/(n-p-1)} Here, p=4p=4 slope coefficients and n=40n=40 observations, so np1=4041=35n-p-1 = 40-4-1 = 35. Plugging in R2=0.30R^2 = 0.30: F=0.30/40.70/35=0.0750.020=3.75F = \frac{0.30/4}{0.70/35} = \frac{0.075}{0.020} = 3.75 The reference distribution is Fp,np1=F4,35F_{p,\, n-p-1} = F_{4,35}, making C the correct answer. Now, let's see exactly where the distractors go wrong. Answer A gets the F-statistic right (3.75) but uses F5,35F_{5,35} — a classic off-by-one error where the student counts the intercept as one of the tested parameters, inflating the numerator degrees of freedom to 5. Remember: you are testing only the four slope coefficients, not the intercept. Answer B uses the correct degrees of freedom F4,35F_{4,35} but arrives at F=2.63F=2.63, which corresponds to dividing R2R^2 by np1=35n-p-1=35 instead of p=4p=4 — essentially swapping numerator and denominator degrees of freedom in the formula. Answer D uses F4,36F_{4,36}, suggesting the student computed the residual degrees of freedom as np=36n-p = 36 rather than np1=35n-p-1 = 35, forgetting to subtract 1 for the intercept. A reliable memory trick: the denominator degrees of freedom is always n(total parameters estimated)n - (\text{total parameters estimated}), and the model with an intercept plus 4 slopes estimates 5 parameters total, giving 405=3540 - 5 = 35.

Question 6

A simple linear regression is parameterized as Y=β0+β1x+εY=\beta_0+\beta_1x+\varepsilon. The estimates and estimated covariance matrix are β^=(10.2),Cov^(β^)=(0.250.020.020.01),\widehat\beta=\begin{pmatrix}1\\0.2\end{pmatrix},\qquad \widehat{\operatorname{Cov}}(\widehat\beta)=\begin{pmatrix}0.25&-0.02\\-0.02&0.01\end{pmatrix}, with 2222 residual degrees of freedom. The predictor is redefined as z=x10z=x-10, and the model is written as Y=γ0+γ1z+εY=\gamma_0+\gamma_1z+\varepsilon. What is the tt statistic for testing H0:γ0=0H_0:\gamma_0=0?

  1. t=1/1.650.78t=-1/\sqrt{1.65}\approx-0.78
  2. t=3/1.252.68t=3/\sqrt{1.25}\approx2.68
  3. t=1/0.25=2.00t=1/\sqrt{0.25}=2.00
  4. t=3/0.853.25t=3/\sqrt{0.85}\approx3.25 (correct answer)
Explanation: When you reparameterize a linear model by shifting the predictor, the new intercept is a linear combination of the original coefficients — and you must propagate uncertainty through that transformation using the delta method (or equivalently, the covariance matrix). With z=x10z = x - 10, the new model Y=γ0+γ1z+εY = \gamma_0 + \gamma_1 z + \varepsilon relates to the original by γ0=β0+10β1\gamma_0 = \beta_0 + 10\beta_1 and γ1=β1\gamma_1 = \beta_1. So the point estimate is γ^0=1+10(0.2)=3\widehat{\gamma}_0 = 1 + 10(0.2) = 3. The variance of γ^0\widehat{\gamma}_0 requires the full covariance propagation: if c=(1,10)\mathbf{c} = (1, 10)^\top, then Var(γ^0)=cCov^(β^)c\operatorname{Var}(\widehat{\gamma}_0) = \mathbf{c}^\top \widehat{\operatorname{Cov}}(\widehat{\beta})\,\mathbf{c}. Computing: 0.25(1)2+2(0.02)(1)(10)+0.01(10)2=0.250.40+1.00=0.850.25(1)^2 + 2(-0.02)(1)(10) + 0.01(10)^2 = 0.25 - 0.40 + 1.00 = 0.85. The tt statistic is t=3/0.853.25t = 3/\sqrt{0.85} \approx 3.25, confirming answer D. Choice C uses the correct numerator but the original variance of β^0\widehat{\beta}_0 (0.25), ignoring how the shift inflates variance through the covariance term. Choice B gets an intermediate variance (1.25) by adding variances without properly accounting for the negative covariance, suggesting a sign error in the cross-term. Choice A gets the numerator wrong entirely, using 1 instead of 3 — it tests the original β0=0\beta_0 = 0, not the reparameterized intercept. Study tip: Whenever a predictor is shifted or scaled, always recompute the estimated variance of the new intercept using cΣ^c\mathbf{c}^\top \widehat{\Sigma}\,\mathbf{c} — never just read off a diagonal entry from the original covariance matrix.

Question 7

Let an invertible reparameterization of a full-rank linear model be defined by γ=Aβ\gamma=A\beta. An analyst wishes to test the original hypothesis H0:Cβ=dH_0:C\beta=d using the parameterization in terms of γ\gamma. Which hypothesis and covariance matrix yield exactly the same general linear FF statistic?

  1. H0:Cγ=dH_0:C\gamma=d with Cov(γ^)=ACov(β^)A1\operatorname{Cov}(\widehat\gamma)=A\operatorname{Cov}(\widehat\beta)A^{-1}
  2. H0:CAγ=dH_0:CA\gamma=d with Cov(γ^)=A1Cov(β^)AT\operatorname{Cov}(\widehat\gamma)=A^{-1}\operatorname{Cov}(\widehat\beta)A^{-T}
  3. H0:ATCTγ=dH_0:A^{-T}C^T\gamma=d with Cov(γ^)=ATCov(β^)A\operatorname{Cov}(\widehat\gamma)=A^T\operatorname{Cov}(\widehat\beta)A
  4. H0:CA1γ=dH_0:CA^{-1}\gamma=d with Cov(γ^)=ACov(β^)AT\operatorname{Cov}(\widehat\gamma)=A\operatorname{Cov}(\widehat\beta)A^T (correct answer)
Explanation: Whenever you encounter a reparameterization problem, your instinct should be to express everything in terms of the new parameter and then check whether the F-statistic is preserved. The general linear F-statistic for testing H0:Cβ=dH_0: C\beta = d has the form (Cβ^d)T[CCov(β^)CT]1(Cβ^d)(C\hat{\beta} - d)^T [C \operatorname{Cov}(\hat{\beta}) C^T]^{-1} (C\hat{\beta} - d), and your goal is to rewrite this entirely in terms of γ^=Aβ^\hat{\gamma} = A\hat{\beta}, meaning β^=A1γ^\hat{\beta} = A^{-1}\hat{\gamma}. Substituting, the estimable function becomes Cβ^=CA1γ^C\hat{\beta} = CA^{-1}\hat{\gamma}, so the hypothesis must be rewritten as H0:CA1γ=dH_0: CA^{-1}\gamma = d. For the covariance, since γ^=Aβ^\hat{\gamma} = A\hat{\beta}, the delta method (or direct propagation) gives Cov(γ^)=ACov(β^)AT\operatorname{Cov}(\hat{\gamma}) = A\operatorname{Cov}(\hat{\beta})A^T. Plugging into the quadratic form, you get (CA1γ^d)T[CA1ACov(β^)ATAT]1(CA1γ^d)(CA^{-1}\hat{\gamma} - d)^T [CA^{-1} \cdot A\operatorname{Cov}(\hat{\beta})A^T \cdot A^{-T}]^{-1}(CA^{-1}\hat{\gamma}-d), which simplifies back to the original statistic exactly. This confirms D is correct. Choice A incorrectly substitutes γ\gamma directly for β\beta without accounting for A1A^{-1}, and the covariance transformation is also wrong. Choice B uses CACA instead of CA1CA^{-1}—this confuses the direction of the substitution—and the covariance uses A1A^{-1} where AA is needed. Choice C applies a transpose incorrectly to the constraint matrix, producing a hypothesis that has no natural interpretation in terms of the original parameterization. A reliable strategy: always derive the reparameterized hypothesis by direct substitution (β^=A1γ^\hat{\beta} = A^{-1}\hat{\gamma}), then propagate the covariance using Cov(Aβ^)=ACov(β^)AT\operatorname{Cov}(A\hat{\beta}) = A\operatorname{Cov}(\hat{\beta})A^T. These two steps together uniquely determine the correct answer.

Question 8

A regression model is written as Y=β01+β1x1+β2x2+ε,Y=\beta_0\mathbf 1+\beta_1x_1+\beta_2x_2+\varepsilon, but the observed design columns satisfy x2=2x1x_2=2x_1 exactly. Which slope functional can be tested using a generalized inverse in a way that is invariant to the particular least-squares solution selected?

  1. β1+2β2\beta_1+2\beta_2 (correct answer)
  2. 2β1+β22\beta_1+\beta_2
  3. β12β2\beta_1-2\beta_2
  4. β1+β2\beta_1+\beta_2
Explanation: When a design matrix is rank-deficient, individual parameters like β1\beta_1 and β2\beta_2 are not estimable — their values shift depending on which generalized inverse you use. The key concept here is estimability: a linear combination cβc^\top\beta is estimable if and only if cc^\top lies in the row space of the design matrix XX. Here, x2=2x1x_2 = 2x_1 exactly, so the column space of X=[1, x1, x2]X = [\mathbf{1},\ x_1,\ x_2] is spanned by only 1\mathbf{1} and x1x_1. The row space of XX is generated by rows of the form (1,xi1,2xi1)(1, x_{i1}, 2x_{i1}), meaning any estimable contrast in (β0,β1,β2)(\beta_0, \beta_1, \beta_2) must satisfy c=(c0,c1,c2)c = (c_0, c_1, c_2) where c2=2c1c_2 = 2c_1. For slope-only contrasts (c0=0c_0 = 0), estimability requires c2=2c1c_2 = 2c_1. Choice A, β1+2β2\beta_1 + 2\beta_2, corresponds to c=(0,1,2)c = (0, 1, 2), satisfying c2=2c1c_2 = 2c_1. This lies in the row space, so it is estimable and invariant across all generalized inverse solutions — confirming A is correct. Choice B, 2β1+β22\beta_1 + \beta_2, gives c=(0,2,1)c = (0,2,1); here c2=12(2)=4c_2 = 1 \neq 2(2) = 4, so it fails the estimability condition. Choice C, β12β2\beta_1 - 2\beta_2, gives c2=22(1)c_2 = -2 \neq 2(1), also non-estimable. Choice D, β1+β2\beta_1 + \beta_2, gives c2=12c_2 = 1 \neq 2, again failing. Study tip: When you see exact collinearity x2=kx1x_2 = kx_1, immediately check whether a proposed linear combination satisfies c2=kc1c_2 = k \cdot c_1. That ratio test is your fast path to identifying estimability under rank deficiency.

Question 9

Ordinary least squares is applied to independent observations with correctly specified conditional means but heteroskedastic errors. For qq linear restrictions, an analyst forms the sandwich-covariance Wald statistic W=(Cβ^d)T{CV^HCCT}1(Cβ^d).W=(C\widehat\beta-d)^T\{C\widehat V_{HC}C^T\}^{-1}(C\widehat\beta-d). Which statement about null calibration is generally correct without imposing homoskedastic normal errors?

  1. WW has an exact χq2\chi_q^2 distribution because the sandwich estimator consistently estimates coefficient covariance
  2. W/qW/q has an exact Fq,npF_{q,n-p} distribution whenever the design matrix has full column rank
  3. WW is asymptotically χq2\chi_q^2, but it need not have an exact finite-sample chi-square or FF distribution (correct answer)
  4. W/qW/q is asymptotically Fq,nqF_{q,n-q}, with denominator degrees of freedom determined by the restrictions
Explanation: When you see a question mixing sandwich (HC) covariance estimators with hypothesis testing, the core issue is always the distinction between exact finite-sample distributions and asymptotic approximations. The Wald statistic W=(Cβ^d)T{CV^HCCT}1(Cβ^d)W = (C\widehat{\beta} - d)^T \{C\widehat{V}_{HC}C^T\}^{-1}(C\widehat{\beta} - d) is built from OLS estimates and a heteroskedasticity-consistent covariance estimator. Under standard regularity conditions, as nn \to \infty, the sandwich estimator consistently estimates the true asymptotic covariance of n(β^β)\sqrt{n}(\widehat{\beta} - \beta), and by the continuous mapping theorem combined with asymptotic normality of OLS, Wdχq2W \xrightarrow{d} \chi^2_q under the null. This makes C correct: the distribution is asymptotic, not exact. A is wrong because consistency of V^HC\widehat{V}_{HC} does not imply exactness — it only guarantees convergence in distribution. In finite samples, the sandwich estimator introduces its own estimation error, so the chi-square reference is only approximate. B is wrong on two levels: the exact Fq,npF_{q,n-p} distribution for OLS Wald statistics requires both normal errors and homoskedasticity (so that quadratic forms have exact chi-square distributions). Neither condition is assumed here, and the sandwich covariance breaks the exact pivotal structure entirely. D is wrong because the asymptotic null distribution of WW is χq2\chi^2_q, not Fq,F_{q,\cdot}. Dividing by qq to form an FF-statistic is a finite-sample correction appropriate under classical assumptions, not a general asymptotic result. Study tip: Always distinguish "consistent estimation" from "exact distribution." On inference questions, consistent sandwich covariance buys you asymptotic validity — never exactness.

Question 10

A full-rank normal linear model has residual degrees of freedom 1818. For one coefficient, β^j=0.70\widehat\beta_j=0.70 and SE(β^j)=0.20\operatorname{SE}(\widehat\beta_j)=0.20. Consider the null hypothesis H0:βj=0.20H_0:\beta_j=0.20. Which statement correctly relates the corresponding tt and general linear FF tests?

  1. F=6.25F=6.25 with degrees of freedom (1,18)(1,18), and its p-value equals the two-sided tt-test p-value (correct answer)
  2. F=2.50F=2.50 with degrees of freedom (1,18)(1,18), and its p-value equals the two-sided tt-test p-value
  3. F=6.25F=6.25 with degrees of freedom (1,18)(1,18), and its p-value equals the upper one-sided tt-test p-value
  4. F=3.125F=3.125 with degrees of freedom (1,18)(1,18), because the squared tt statistic must be divided by two
Explanation: Whenever you see a question linking tt and FF tests in a linear model, the core relationship to anchor on is this: the general linear FF statistic for a single restriction equals the square of the corresponding tt statistic, and the two tests produce identical two-sided p-values. Here, the tt statistic for H0:βj=0.20H_0: \beta_j = 0.20 is t=β^j0.20SE(β^j)=0.700.200.20=0.500.20=2.50,t = \frac{\widehat{\beta}_j - 0.20}{\operatorname{SE}(\widehat{\beta}_j)} = \frac{0.70 - 0.20}{0.20} = \frac{0.50}{0.20} = 2.50, with 1818 residual degrees of freedom. Squaring gives F=t2=2.502=6.25F = t^2 = 2.50^2 = 6.25, distributed as F(1,18)F(1, 18) under H0H_0. Because the F(1,18)F(1,18) distribution only has an upper tail, and the two-sided tt test splits probability into both tails, the p-value from P(F1,18>6.25)P(F_{1,18} > 6.25) equals exactly the two-sided tt p-value P(t18>2.50)P(|t_{18}| > 2.50). This confirms A is correct. B is wrong because it reports F=2.50F = 2.50, which is the tt statistic itself, not its square. C is wrong on the p-value: because FF is always positive (it captures both directions of deviation), its upper-tail p-value matches the two-sided tt test, not a one-sided test. D introduces a fictitious "divide by two" step — no such adjustment exists; F=t2F = t^2 directly, giving 6.256.25, not 3.1253.125. A reliable study tip: for any single-coefficient hypothesis in a linear model, memorize F=t2F = t^2 and that the resulting FF p-value always corresponds to the two-sided tt p-value — this connection appears frequently on graduate-level exams.