Statistics Graduate Level Quiz: Sandwich Variance And Robust Se
10 questions · exam conditions
0:00
Sandwich Variance And Robust SeQuestion 1 of 10

An estimator θ^\widehat{\theta} solves an estimating equation based on independent observations. At the true parameter, the scalar sensitivity and score variance are A=2A=2 and B=9B=9, respectively, where n(θ^θ0)\sqrt{n}(\widehat{\theta}-\theta_0) has sandwich variance A1BA1A^{-1}BA^{-1}. Which expression is the asymptotic variance of θ^\widehat{\theta} itself?

94n\frac{9}{4n}, because the sandwich variance for the scaled estimator must be divided by nn.
92n\frac{9}{2n}, because the score variance is premultiplied by only one inverse sensitivity.
49n\frac{4}{9n}, because the sensitivity and score variance exchange roles when scaling by nn.
94n\frac{9}{4\sqrt{n}}, because the variance of the scaled estimator is divided by n\sqrt{n}.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Sandwich Variance And Robust Se

Practice Sandwich Variance And Robust Se in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Sandwich Variance And Robust Se, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

An estimator θ^\widehat{\theta} solves an estimating equation based on independent observations. At the true parameter, the scalar sensitivity and score variance are A=2A=2 and B=9B=9, respectively, where n(θ^θ0)\sqrt{n}(\widehat{\theta}-\theta_0) has sandwich variance A1BA1A^{-1}BA^{-1}. Which expression is the asymptotic variance of θ^\widehat{\theta} itself?

  1. 94n\frac{9}{4n}, because the sandwich variance for the scaled estimator must be divided by nn. (correct answer)
  2. 92n\frac{9}{2n}, because the score variance is premultiplied by only one inverse sensitivity.
  3. 49n\frac{4}{9n}, because the sensitivity and score variance exchange roles when scaling by nn.
  4. 94n\frac{9}{4\sqrt{n}}, because the variance of the scaled estimator is divided by n\sqrt{n}.
Explanation: Whenever you encounter sandwich variance problems, your first instinct should be to carefully track which quantity has the stated asymptotic distribution, then scale appropriately to find the variance of the estimator itself. Here, you're told that n(θ^θ0)\sqrt{n}(\widehat{\theta} - \theta_0) has asymptotic variance A1BA1A^{-1}BA^{-1}. Plugging in A=2A = 2 and B=9B = 9, this sandwich variance equals 12912=94\frac{1}{2} \cdot 9 \cdot \frac{1}{2} = \frac{9}{4}. But this is the variance of the scaled quantity n(θ^θ0)\sqrt{n}(\widehat{\theta} - \theta_0), not of θ^\widehat{\theta} itself. Since Var(nX)=nVar(X)\text{Var}(\sqrt{n}\cdot X) = n \cdot \text{Var}(X), you divide by nn to recover the asymptotic variance of θ^\widehat{\theta}: 94n\frac{9}{4n}. That confirms A is correct. B is wrong because it applies only one inverse sensitivity — computing A1B=92A^{-1}B = \frac{9}{2} — which mistakes the sandwich formula for a one-sided expression. The correct sandwich always applies A1A^{-1} on both sides. C inverts the roles of AA and BB, computing B1AB11n=49nB^{-1}AB^{-1} \cdot \frac{1}{n} = \frac{4}{9n}. This reflects a conceptual swap of sensitivity and score variance, which has no theoretical justification. D correctly computes the sandwich value 94\frac{9}{4} but then divides by n\sqrt{n} instead of nn, confusing the standard deviation scaling of n\sqrt{n} with the variance scaling of nn. Study tip: Always identify what the asymptotic distribution describes — the scaled or unscaled estimator — before computing any variance. The n\sqrt{n} versus nn distinction is a classic trap on graduate-level exams.

Question 2

A researcher estimates an ordinary least squares regression of an outcome on a treatment indicator and observed covariates. The conditional error variance depends on the covariates, and treatment remains correlated with an omitted determinant of the outcome after conditioning on the included covariates. The researcher reports heteroskedasticity-robust standard errors.

Which conclusion is most appropriate under this data-generating process?

  1. The robust standard errors consistently estimate uncertainty around the causal coefficient because they accommodate both heteroskedasticity and omitted-variable bias.
  2. The robust standard errors may estimate sampling variability around a pseudo-true projection coefficient, but they do not make that coefficient causally consistent. (correct answer)
  3. The robust standard errors are invalid solely because the conditional error variance is nonconstant, although the coefficient remains causally consistent.
  4. The robust standard errors remove the coefficient bias asymptotically, but conventional homoskedastic standard errors would leave the bias unchanged.
Explanation: When a question combines robust standard errors with omitted-variable bias (OVB), you need to keep two separate problems firmly distinguished: what the coefficient estimates and how precisely it is estimated. Heteroskedasticity-robust standard errors address only the second problem. When the error variance is nonconstant, conventional OLS standard errors are inconsistent estimators of the true sampling variability. Robust standard errors correct this by using a sandwich estimator that remains valid asymptotically regardless of the error variance structure. Crucially, however, robust standard errors operate on whatever coefficient OLS produces — they do not touch the coefficient itself. When treatment is correlated with an omitted determinant of the outcome even after conditioning on included covariates, OLS converges to a pseudo-true (or "projection") coefficient rather than the causal parameter. More precisely, the probability limit of the OLS estimator is β^pβ+bias,\hat{\beta} \xrightarrow{p} \beta + \text{bias}, where the bias term reflects the omitted variable's influence. Robust standard errors then consistently estimate the sampling variability around this biased limit — they do their job correctly, but that job is not causal identification. This is exactly what B states. A is wrong because robust standard errors cannot "accommodate" OVB — they have no mechanism to remove coefficient bias. C reverses the logic entirely: the coefficient is not causally consistent (OVB remains), and robust standard errors are not invalid due to heteroskedasticity — that is precisely what they are designed for. D is wrong because neither robust nor conventional standard errors affect the coefficient's probability limit; standard error choice never removes bias. A reliable study habit: always ask two separate questions — what does the estimator converge to? and are the standard errors valid around that limit? They are independent concerns, and exams frequently conflate them as a trap.

Question 3

A panel contains GG independent firms, each observed for TT periods. Regression errors may have arbitrary heteroskedasticity and serial correlation within a firm, but errors are independent across firms. The coefficient estimator is computed using all GTGT observations.

Which asymptotic framework most directly justifies firm-clustered sandwich standard errors?

  1. TT\to\infty with fixed GG, because a large number of observations within each firm makes the empirical cluster covariance consistent.
  2. GG\to\infty with bounded or suitably controlled TT, because firms rather than firm-period observations provide the independent score contributions. (correct answer)
  3. GTGT\to\infty under any allocation between GG and TT, because only the total number of regression observations determines sandwich consistency.
  4. Fixed GG and fixed TT with normally distributed errors, because normality makes the cluster-level covariance estimator exactly unbiased.
Explanation: When thinking about clustered standard errors, the key question is: what is the independent unit of variation? The sandwich estimator's consistency relies on a law of large numbers over independent "chunks" — and those chunks must grow in number for the asymptotics to work. In a clustered panel, firms are independent of each other, but observations within a firm are not. This means each firm contributes one independent score vector (the sum of its TT period-level scores). The sandwich estimator essentially averages GG of these independent contributions. For the outer product of scores to converge to the true variance, you need the number of independent units — firms — to grow. This is exactly why B is correct: as GG \to \infty, the law of large numbers applies across firms, making the cluster-robust variance estimator consistent, regardless of whether TT is fixed or grows slowly. A is wrong because letting TT \to \infty with fixed GG gives you more data within each cluster, but you still only have GG independent contributions to the sandwich. You can never average away estimation error with a fixed number of independent units — you'd need a separate theory (like time-series HAC) entirely. C is wrong because total observations GTGT \to \infty is insufficient on its own. If GG stays fixed while TT grows, the independence structure is violated for the cluster sandwich; the estimator may not be consistent. D is wrong because normality with fixed GG and TT doesn't rescue the sandwich — you'd have too few clusters to average over, and the covariance estimator remains noisy regardless of distributional assumptions. Study tip: Always ask "how many independent units are there?" — that number, not total observations, drives sandwich consistency in clustered settings.

Question 4

A parametric likelihood model is fitted by maximum likelihood, but the assumed conditional density is misspecified. Let θ\theta^* maximize the expected log likelihood under the true distribution. At θ\theta^*, define HH as the negative expected Hessian of the log likelihood and JJ as the variance of the score.

Under standard misspecified maximum-likelihood regularity conditions, which covariance statement is correct?

  1. The estimator targets the true structural parameter and has covariance H1/nH^{-1}/n because the information identity continues to hold under misspecification.
  2. The estimator targets θ\theta^* and has covariance H1JH1/nH^{-1}JH^{-1}/n, which reduces to H1/nH^{-1}/n if the information identity holds. (correct answer)
  3. The estimator targets θ\theta^* and has covariance J1HJ1/nJ^{-1}HJ^{-1}/n, because the score variance supplies the outer bread matrices.
  4. The estimator targets the true structural parameter and has covariance J/nJ/n because only the empirical variability of the score remains relevant.
Explanation: When a likelihood model is misspecified, the key insight is that the MLE converges not to the true structural parameter, but to the pseudo-true parameter θ\theta^*, defined as the minimizer of the Kullback-Leibler divergence between the true distribution and the assumed model. Once you accept this, the asymptotic covariance follows from a careful application of the delta method to the score equations. Under correct specification, the information identity H=JH = J holds, collapsing the sandwich to H1/nH^{-1}/n. Under misspecification, this equality breaks down. The score is still asymptotically normal, but its variance JJ no longer matches the curvature matrix HH. Expanding the MLE around θ\theta^* via a Taylor argument yields the sandwich (Huber-White) covariance: Var(θ^)=H1JH1/n\text{Var}(\hat{\theta}) = H^{-1} J H^{-1} / n Here H1H^{-1} is the "bread" (appearing on both sides) and JJ is the "meat." This confirms B is correct — and B also correctly notes that when H=JH = J, the expression reduces to the familiar H1/nH^{-1}/n. A is wrong on two counts: the estimator does not target the true structural parameter under misspecification, and the information identity does not hold, so H1/nH^{-1}/n is the wrong variance. C inverts the roles of HH and JJ, placing J1J^{-1} as bread and HH as meat — a direct reversal of the correct sandwich formula. D incorrectly claims the target is the true parameter and that only J/nJ/n matters, ignoring the curvature correction entirely. Study tip: Memorize the sandwich as "bread-meat-bread" = H1JH1/nH^{-1}JH^{-1}/n. Any answer that flips or drops a component is a trap.

Question 5

In a completely randomized experiment with a fixed treatment fraction, the difference in sample means is estimated by OLS using only an intercept and a treatment indicator. Potential outcomes are treated as fixed, and individual treatment effects may vary. The researcher reports an unpooled heteroskedasticity-robust standard error.

How does the large-sample robust variance estimator generally relate to the finite-population randomization variance?

  1. It is inconsistent unless outcome variances are identical across treatment arms because robust regression requires homoskedastic potential outcomes.
  2. It is generally anti-conservative because the randomization variance adds an unidentifiable treatment-effect heterogeneity term omitted by the robust estimator.
  3. It equals the randomization variance exactly even with heterogeneous treatment effects because random assignment implies a zero covariance between potential outcomes.
  4. It estimates the identifiable within-arm variance terms and is generally conservative because the randomization variance subtracts an unidentifiable treatment-effect heterogeneity term. (correct answer)
Explanation: When you see a question linking OLS robust standard errors to randomization-based inference, your job is to track what each variance formula actually estimates and whether anything is systematically left out or added. Under the Neyman randomization framework with fixed potential outcomes, the exact finite-population variance of the difference-in-means estimator is: Vrand=S12n1+S02n0Sτ2nV_{\text{rand}} = \frac{S_1^2}{n_1} + \frac{S_0^2}{n_0} - \frac{S_{\tau}^2}{n} where S12,S02S_1^2, S_0^2 are finite-population variances of treated and control potential outcomes, and Sτ2S_{\tau}^2 is the variance of individual treatment effects τi=Yi(1)Yi(0)\tau_i = Y_i(1) - Y_i(0). That last term is unidentifiable — you never observe both potential outcomes for the same unit. The heteroskedasticity-robust (Eicker-Huber-White) variance estimator consistently estimates only the first two terms: S12n1+S02n0\frac{S_1^2}{n_1} + \frac{S_0^2}{n_0}. Since Sτ20S_{\tau}^2 \geq 0, the robust estimator omits a non-negative term that the true variance subtracts, making it an overestimate — i.e., conservative. This confirms D is correct. A is wrong because the robust estimator is perfectly consistent for the within-arm variance components; homoskedasticity is not required. B gets the direction exactly backwards — the robust estimator is conservative (over-estimates), not anti-conservative, because it fails to subtract Sτ2/nS_{\tau}^2/n, not because it adds something spurious. C is wrong because zero covariance between potential outcomes does not collapse the variance formula; the treatment-effect heterogeneity term survives under all assignment mechanisms. Your study tip: always ask whether a variance formula includes or excludes Sτ2S_{\tau}^2. Its sign in the randomization variance (subtracted) is the key to remembering that robust SEs are conservative, not liberal.

Question 6

A generalized estimating equation is used for longitudinal binary outcomes. The marginal mean model is correctly specified, but the chosen working correlation matrix is incorrect. Subjects are independent, the number of subjects tends to infinity, and each subject contributes a bounded number of observations.

Which statement most accurately describes inference for the regression coefficient?

  1. The coefficient estimator is generally inconsistent because a correct working correlation is necessary for unbiased generalized estimating equations.
  2. Both covariance estimators are automatically valid because bounded cluster size makes misspecification of within-subject correlation asymptotically irrelevant.
  3. The coefficient estimator remains consistent, but only the model-based covariance is valid because the sandwich requires the working correlation to be correct.
  4. The coefficient estimator can remain consistent, and the subject-clustered sandwich covariance is valid, although the working-correlation model-based covariance may fail. (correct answer)
Explanation: Whenever you see a GEE question involving misspecified working correlation, your anchor concept should be the separation of concerns in GEE: the estimating equations' unbiasedness depends only on the mean model, not the correlation structure. GEE solves i=1nDiTVi1(Yiμi)=0\sum_{i=1}^n D_i^T V_i^{-1}(Y_i - \mu_i) = 0, where ViV_i is constructed from the working correlation. As long as E[Yiμi]=0E[Y_i - \mu_i] = 0 (mean model correctly specified), the estimating equations remain unbiased regardless of whether ViV_i matches the true correlation. With independent subjects and nn \to \infty, a standard law-of-large-numbers argument ensures consistency of β^\hat{\beta}. The sandwich (robust) covariance V^sandwich=A^1B^A^1\hat{V}_{sandwich} = \hat{A}^{-1}\hat{B}\hat{A}^{-1} is specifically designed to remain valid under working-correlation misspecification — it uses the empirical "meat" B^\hat{B} to correct for the mismatch. The model-based covariance, however, assumes ViV_i is exactly correct; when it's wrong, this estimator is inconsistent for the true variance of β^\hat{\beta}. This makes D the correct answer. A is wrong because consistency of β^\hat{\beta} doesn't require a correct working correlation — only a correctly specified mean model. B is wrong in both claims: bounded cluster size does not make correlation misspecification irrelevant, and the model-based covariance remains invalid under misspecification. C has the covariance story exactly backwards — it's the sandwich, not the model-based estimator, that survives working-correlation misspecification. A useful rule of thumb: in GEE, the sandwich covariance is your insurance policy against a wrong working correlation, while the model-based covariance is only trustworthy when the working correlation is correct.

Question 7

A time-series regression has conditionally mean-zero errors, but the regression score exhibits weak serial dependence. The coefficient estimator is consistent. A researcher compares the usual heteroskedasticity-robust covariance estimator with a heteroskedasticity-and-autocorrelation-consistent covariance estimator.

Which statement gives the strongest justification for using the heteroskedasticity-and-autocorrelation-consistent estimator?

  1. A fixed bandwidth suffices for consistency because weak dependence implies autocovariances beyond any fixed lag are exactly zero in large samples.
  2. It corrects endogeneity in the regression coefficients by reweighting observations according to estimated serial correlations in the residuals.
  3. It estimates the long-run score covariance by incorporating weighted lagged score covariances, with the bandwidth growing slowly relative to the sample size. (correct answer)
  4. It treats the entire time series as one independent cluster, so the standard many-cluster sandwich consistency argument applies directly.
Explanation: Whenever you see a question about variance estimation in time-series settings, anchor yourself to one core idea: the object you need to estimate is the long-run variance of the score, not just its contemporaneous variance. In OLS with serially correlated scores, the asymptotic covariance of β^\hat{\beta} depends on S=j=ΓjS = \sum_{j=-\infty}^{\infty} \Gamma_j, where Γj=E[ststj]\Gamma_j = E[s_t s_{t-j}'] is the lag-jj score autocovariance. A heteroskedasticity-and-autocorrelation-consistent (HAC) estimator — such as Newey-West — approximates this sum by computing weighted sample autocovariances up to some truncation bandwidth MM. Crucially, for consistency, MM \to \infty but M/n0M/n \to 0, so the bandwidth grows slowly relative to the sample size. This is exactly what option C describes, making it the strongest, technically correct justification. Option A is wrong because weak dependence does not imply autocovariances are exactly zero beyond any fixed lag. It only implies they decay to zero asymptotically; a fixed bandwidth would omit non-negligible contributions, causing inconsistency. Option B confuses two separate problems. HAC estimation addresses variance estimation for inference, not endogeneity in the coefficient estimates themselves. Serial correlation in residuals does not bias β^\hat{\beta} when errors are conditionally mean-zero — the passage already tells you the estimator is consistent. Option D mischaracterizes the cluster-robust sandwich estimator. Treating the whole time series as one cluster actually requires strong within-cluster homogeneity assumptions and does not recover the long-run variance the way a proper HAC kernel does. Your study tip: always distinguish coefficient consistency (a mean condition) from inference validity (a variance condition). HAC fixes the latter, not the former.

Question 8

In a sequence of linear regressions, one observation retains substantial leverage as the sample size grows, so the maximum diagonal element of the hat matrix does not approach zero. A researcher considers HC0 and HC3 heteroskedasticity-robust covariance estimators.

Which statement best describes the role of HC3 in this setting?

  1. HC3 inflates residual contributions from high-leverage observations and may reduce HC0's downward bias, but it does not universally restore consistency under persistent leverage. (correct answer)
  2. HC3 removes high-leverage observations from the estimating equation and therefore guarantees consistency of both coefficients and standard errors.
  3. HC3 is asymptotically identical to HC0 even with persistent leverage because all leverage corrections vanish whenever the sample size increases.
  4. HC3 corrects only serial correlation, so it provides no adjustment to residual contributions associated with high-leverage observations.
Explanation: When a question involves heteroskedasticity-robust covariance estimators and leverage, your anchor concept should be how hat matrix diagonal elements hiih_{ii} appear in each estimator's formula and what happens when those elements stay large as nn \to \infty. HC0 scales squared residuals directly: e^i2\hat{e}_i^2. HC3 instead uses e^i2/(1hii)2\hat{e}_i^2 / (1 - h_{ii})^2, a jackknife-motivated correction that upweights residuals from high-leverage observations. In normal asymptotics, hii0h_{ii} \to 0 for all observations, so this correction vanishes and HC0 and HC3 converge. The catch here is that one observation retains high leverage — maxihii↛0\max_i h_{ii} \not\to 0 — which violates the standard regularity condition. HC3's inflation of that observation's residual contribution does partially counteract HC0's downward bias in finite samples, but because the leverage persists indefinitely, HC3 cannot guarantee consistency in this non-standard setting either. That logic confirms A as correct. B is wrong because HC3 never removes observations — it reweights residual contributions. Claiming it "guarantees consistency" misrepresents both its mechanism and its theoretical limits. C is wrong in the opposite direction: it assumes hii0h_{ii} \to 0 universally, which is precisely the condition the passage tells you fails. HC0 and HC3 do not become asymptotically equivalent under persistent leverage. D confuses HC3 with HAC (Newey-West style) estimators designed for serial correlation; HC3 makes no serial-correlation adjustment and is explicitly about heteroskedasticity with leverage weighting. A useful rule of thumb: whenever you see "persistent leverage" or maxhii↛0\max h_{ii} \not\to 0, treat standard asymptotic guarantees as suspect for all sandwich estimators, not just HC0.

Question 9

For a two-parameter estimator, the estimated sensitivity and score covariance are A^=(2001)\widehat{A}=\begin{pmatrix}2&0\\0&1\end{pmatrix} and B^=(4119)\widehat{B}=\begin{pmatrix}4&1\\1&9\end{pmatrix} . The covariance estimator is A^1B^A^T/n\widehat{A}^{-1}\widehat{B}\widehat{A}^{-T}/n.

What is the sandwich standard error for the estimated contrast θ^1θ^2\widehat{\theta}_1-\widehat{\theta}_2?

  1. 3n\frac{3}{\sqrt{n}}, after including the covariance contribution between the two parameter estimates. (correct answer)
  2. 10n\frac{\sqrt{10}}{\sqrt{n}}, after adding the two marginal sandwich variances and ignoring covariance.
  3. 11n\frac{\sqrt{11}}{\sqrt{n}}, after applying the contrast directly to the score covariance matrix.
  4. 3/2n\frac{\sqrt{3/2}}{\sqrt{n}}, after using the inverse sensitivity as the covariance matrix.
Explanation: When working with the sandwich estimator, your goal is to compute the full covariance matrix of your parameter estimates, then apply the linear contrast to extract the relevant variance. The sandwich covariance matrix is V^=A^1B^A^T/n\widehat{V} = \widehat{A}^{-1}\widehat{B}\widehat{A}^{-T}/n. Since A^\widehat{A} is diagonal, A^1=diag(1/2,1)\widehat{A}^{-1} = \text{diag}(1/2, 1), and computing A^1B^A^T\widehat{A}^{-1}\widehat{B}\widehat{A}^{-T} gives (11/219)\begin{pmatrix}1&1/2\\1&9\end{pmatrix}... let's be precise: the (i,j)(i,j) entry scales row ii by 1/Aii1/A_{ii} and column jj by 1/Ajj1/A_{jj}, yielding (4/41/21/29)=(11/21/29)\begin{pmatrix}4/4 & 1/2 \\ 1/2 & 9\end{pmatrix} = \begin{pmatrix}1 & 1/2 \\ 1/2 & 9\end{pmatrix}. For the contrast c=(1,1)Tc = (1, -1)^T, the variance is cTV^c/1=(1,1)(11/21/29)(11)/nc^T \widehat{V} c/1 = (1,-1)\begin{pmatrix}1&1/2\\1/2&9\end{pmatrix}\begin{pmatrix}1\\-1\end{pmatrix}/n. This equals (11/21/2+9)/n=9/n(1 - 1/2 - 1/2 + 9)/n = 9/n, giving standard error 3/n3/\sqrt{n}. Answer A is correct. Answer B is wrong because it sums only the diagonal entries of V^\widehat{V} (i.e., 1+9=101 + 9 = 10), ignoring the off-diagonal covariance term 2(1/2)-2(1/2), which crucially reduces the variance. Answer C incorrectly applies the contrast directly to B^\widehat{B} without the sensitivity scaling, skipping the A^1\widehat{A}^{-1} transformation entirely. Answer D uses only A^1A^T\widehat{A}^{-1}\widehat{A}^{-T}, treating the sensitivity inverse itself as the covariance and discarding B^\widehat{B} altogether. The key study tip: whenever you see a sandwich estimator with a linear contrast, always compute the full covariance matrix first, then apply cTV^cc^T \widehat{V} c. Off-diagonal terms can substantially change the answer, and forgetting them is the most common trap on these questions.

Question 10

An estimator satisfies Avar(θ^)=Σ/n\operatorname{Avar}(\widehat{\theta})=\Sigma/n, where θ^=(θ^1,θ^2)T\widehat{\theta}=(\widehat{\theta}_1,\widehat{\theta}_2)^T, θ=(2,4)T\theta=(2,4)^T, and Σ=(4229)\Sigma=\begin{pmatrix}4&2\\2&9\end{pmatrix} . The parameter of interest is the ratio g(θ)=θ1/θ2g(\theta)=\theta_1/\theta_2.

Using the delta method with the sandwich covariance, what is the asymptotic standard error of g(θ^)g(\widehat{\theta})?

  1. 338n\frac{\sqrt{33}}{8\sqrt{n}}, treating the covariance contribution as positive rather than negative.
  2. 58n\frac{5}{8\sqrt{n}}, using both marginal variances but omitting their covariance.
  3. 178n\frac{\sqrt{17}}{8\sqrt{n}}, using both marginal variances and their positive covariance. (correct answer)
  4. 12n\frac{1}{2\sqrt{n}}, treating the denominator parameter as fixed rather than estimated.
Explanation: Whenever you encounter a question involving a function of a vector estimator, the delta method is your go-to tool. The key formula is: if Avar(θ^)=Σ/n\operatorname{Avar}(\widehat{\theta}) = \Sigma/n, then Avar(g(θ^))=g(θ)TΣg(θ)/n\operatorname{Avar}(g(\widehat{\theta})) = \nabla g(\theta)^T \Sigma \nabla g(\theta) / n, where g\nabla g is the gradient evaluated at the true parameter. For g(θ)=θ1/θ2g(\theta) = \theta_1/\theta_2, the gradient is g=(1θ2, θ1θ22)T\nabla g = \left(\frac{1}{\theta_2},\ -\frac{\theta_1}{\theta_2^2}\right)^T. Plugging in θ=(2,4)T\theta = (2,4)^T gives g=(14, 18)T\nabla g = \left(\frac{1}{4},\ -\frac{1}{8}\right)^T. Now compute the sandwich product with $$\Sigma = \begin{pmatrix}4&2\2&9\end{pmatrix} $$\Sigma \nabla g = \begin{pmatrix}4(1/4)+2(-1/8)\\2(1/4)+9(-1/8)\end{pmatrix} = \begin{pmatrix}3/4\\-1/8\end{pmatrix} gTΣg=(1/4)(3/4)+(1/8)(1/8)=3/16+1/64=17/64\nabla g^T \Sigma \nabla g = (1/4)(3/4) + (-1/8)(-1/8) = 3/16 + 1/64 = 17/64 So the asymptotic variance is 17/(64n)17/(64n), giving an asymptotic standard error of 17/(8n)\sqrt{17}/(8\sqrt{n}), confirming C. Choice A arises from using a positive sign on the cross-term (1/8)(3/4)(-1/8)(3/4) instead of correctly tracking the negative gradient component, yielding 33/6433/64 instead of 17/6417/64. Choice B ignores the covariance term entirely, summing only the two diagonal contributions (1/4)24+(1/8)29=25/64(1/4)^2 \cdot 4 + (1/8)^2 \cdot 9 = 25/64. Choice D treats θ2\theta_2 as fixed, so only the variance of θ^1\widehat{\theta}_1 enters, giving variance 4/(16n)=1/(4n)4/(16n) = 1/(4n). Your study tip: always write out the full gradient vector explicitly before computing the sandwich form — sign errors on gradient components are the most common source of mistakes in delta method problems.