Statistics Graduate Level Quiz: M Estimators
10 questions · exam conditions
0:00
M EstimatorsQuestion 1 of 10

Suppose X1,,XnX_1,\ldots,X_n are independent observations from a normal distribution with mean θ0\theta_0 and variance 11. An M-estimator is the consistent root near θ0\theta_0 of i=1n(Xiθ)3=0\sum_{i=1}^n(X_i-\theta)^3=0. What is the asymptotic variance of n(θ^nθ0)\sqrt n(\widehat\theta_n-\theta_0)?

13\tfrac{1}{3}
35\tfrac{3}{5}
53\tfrac{5}{3}
1515
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: M Estimators

Practice M Estimators in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on M Estimators, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Suppose X1,,XnX_1,\ldots,X_n are independent observations from a normal distribution with mean θ0\theta_0 and variance 11. An M-estimator is the consistent root near θ0\theta_0 of i=1n(Xiθ)3=0\sum_{i=1}^n(X_i-\theta)^3=0. What is the asymptotic variance of n(θ^nθ0)\sqrt n(\widehat\theta_n-\theta_0)?

  1. 13\tfrac{1}{3}
  2. 35\tfrac{3}{5}
  3. 53\tfrac{5}{3} (correct answer)
  4. 1515
Explanation: When you see an M-estimator question, your default tool is the sandwich variance formula: for the estimating equation ψ(Xi,θ)=0\sum \psi(X_i, \theta) = 0, the asymptotic variance of n(θ^nθ0)\sqrt{n}(\hat{\theta}_n - \theta_0) is E[ψ2](E[ψ])2\frac{E[\psi^2]}{(E[\psi'])^2}, where ψ=ψ/θ\psi' = \partial\psi/\partial\theta. Here, ψ(X,θ)=(Xθ)3\psi(X,\theta) = (X-\theta)^3, so ψ(X,θ)=3(Xθ)2\psi'(X,\theta) = -3(X-\theta)^2. Evaluated at θ0\theta_0, with Z=Xθ0N(0,1)Z = X - \theta_0 \sim N(0,1): Denominator: E[ψ]=3E[Z2]=3(1)=3E[\psi'] = -3E[Z^2] = -3(1) = -3, so (E[ψ])2=9(E[\psi'])^2 = 9. Numerator: E[ψ2]=E[Z6]E[\psi^2] = E[Z^6]. For a standard normal, E[Z2k]=(2k1)!!E[Z^{2k}] = (2k-1)!!, so E[Z6]=5!!=15E[Z^6] = 5!! = 15. The asymptotic variance is therefore 159=53\frac{15}{9} = \frac{5}{3}, confirming answer C. Now for the distractors. A (13\tfrac{1}{3}) likely comes from inverting only the denominator term E[ψ]=3E[\psi'] = -3 without squaring it and ignoring the numerator entirely — a misremembering of the formula. B (35\tfrac{3}{5}) is the reciprocal of the correct answer; students who flip numerator and denominator in the sandwich formula land here. D (1515) forgets to divide by (E[ψ])2=9(E[\psi'])^2 = 9 at all — it's just the raw sixth moment. Study tip: Always memorize the standard normal moments — E[Z2]=1E[Z^2]=1, E[Z4]=3E[Z^4]=3, E[Z6]=15E[Z^6]=15 — and write the sandwich formula clearly before plugging in. Swapping numerator and denominator is the most common computational error on M-estimator problems.

Question 2

For observations 4,0,1,9-4,0,1,9, a location M-estimator minimizes i=14ρ(Xiθ)\sum_{i=1}^4 \rho(X_i-\theta), where ρ\rho is the Huber loss with threshold 22. Thus its score is ψ(u)=max{2,min(u,2)}\psi(u)=\max\{-2,\min(u,2)\}. Which value solves the corresponding estimating equation?

  1. θ^=0\widehat\theta=0
  2. θ^=12\widehat\theta=\tfrac{1}{2} (correct answer)
  3. θ^=1\widehat\theta=1
  4. θ^=32\widehat\theta=\tfrac{3}{2}
Explanation: When working with M-estimators, the key insight is that you don't minimize the loss directly — you solve the estimating equation i=14ψ(Xiθ)=0\sum_{i=1}^4 \psi(X_i - \theta) = 0, where ψ\psi is the derivative of ρ\rho. For Huber loss with threshold c=2c=2, this means ψ(u)=max{2,min(u,2)}\psi(u) = \max\{-2, \min(u, 2)\}: residuals within [2,2][-2, 2] contribute their actual value, while residuals beyond that are winsorized to ±2\pm 2. To find the solution, evaluate ψ(Xiθ)\sum \psi(X_i - \theta) for each candidate. Try θ=12\theta = \tfrac{1}{2} (answer B): the residuals are 40.5=4.5-4 - 0.5 = -4.5, 00.5=0.50 - 0.5 = -0.5, 10.5=0.51 - 0.5 = 0.5, 90.5=8.59 - 0.5 = 8.5. Applying ψ\psi: the extreme residuals 4.5-4.5 and 8.58.5 get clipped to 2-2 and 22, while the middle residuals pass through. The sum is (2)+(0.5)+(0.5)+(2)=0(-2) + (-0.5) + (0.5) + (2) = 0. ✓ This confirms B is correct. For A (θ=0\theta=0): residuals are 4,0,1,9-4, 0, 1, 9, giving ψ\psi-values 2,0,1,2-2, 0, 1, 2, which sum to 101 \neq 0. For C (θ=1\theta=1): residuals 5,1,0,8-5, -1, 0, 8 yield 2,1,0,2-2, -1, 0, 2, summing to 10-1 \neq 0. For D (θ=32\theta=\tfrac{3}{2}): residuals 5.5,1.5,0.5,7.5-5.5, -1.5, -0.5, 7.5 yield 2,1.5,0.5,2-2, -1.5, -0.5, 2, summing to 20-2 \neq 0. Your strategy: always plug candidates into the estimating equation rather than trying to solve algebraically. Huber M-estimators won't have closed-form solutions in general, so checking the score equation directly is the reliable approach.

Question 3

Let P(X=0)=0.3P(X=0)=0.3, P(X=2)=0.5P(X=2)=0.5, and P(X=5)=0.2P(X=5)=0.2. An estimator minimizes the empirical quantile loss iρ0.6(Xiθ)\sum_i\rho_{0.6}(X_i-\theta), where ρτ(u)=u{τI(u<0)}\rho_\tau(u)=u\{\tau-I(u<0)\}. Which statement correctly describes its population target?

  1. The target is 00 because it is the largest value strictly below cumulative probability 0.60.6.
  2. The target is 22 even though the ordinary derivative-based score need not equal zero there. (correct answer)
  3. The target is 55 because cumulative probability first exceeds 0.60.6 after the point 22.
  4. No population minimizer exists because no value makes E{0.6I(X<θ)}=0E\{0.6-I(X<\theta)\}=0.
Explanation: Whenever you see a question involving quantile loss minimization, your first instinct should be to identify the population quantile being targeted — specifically, the τ\tau-th quantile, which minimizes E[ρτ(Xθ)]E[\rho_\tau(X - \theta)]. For τ=0.6\tau = 0.6, the population minimizer is the 0.6-quantile, defined as Q(0.6)=inf{θ:F(θ)0.6}Q(0.6) = \inf\{\theta : F(\theta) \geq 0.6\}. Computing the CDF: F(0)=0.3F(0) = 0.3, F(2)=0.8F(2) = 0.8, F(5)=1.0F(5) = 1.0. Since F(2)=0.80.6F(2) = 0.8 \geq 0.6 and F(0)=0.3<0.6F(0) = 0.3 < 0.6, the infimum of values where F0.6F \geq 0.6 is exactly θ=2\theta = 2. So B is correct — the population target is 22. The subtle but important point B makes is that the "score equation" E[0.6I(X<θ)]=0E[0.6 - I(X < \theta)] = 0 need not hold at a discrete mass point. At θ=2\theta = 2: E[0.6I(X<2)]=0.6P(X<2)=0.60.3=0.30E[0.6 - I(X < 2)] = 0.6 - P(X < 2) = 0.6 - 0.3 = 0.3 \neq 0. The subgradient at θ=2\theta = 2 contains zero, which is the correct optimality condition for non-smooth losses — not the ordinary derivative. A is wrong because θ=0\theta = 0 is the largest value below cumulative probability 0.6, but that makes it the 0.3-quantile, not the 0.6-quantile. C is wrong in its reasoning: the quantile is the point where FF first reaches 0.6, which is 22, not 55. D is wrong because the absence of an exact zero subgradient doesn't mean no minimizer exists — subgradient optimality is the right criterion here. Your key takeaway: for discrete distributions, the quantile minimizes the check loss via a subgradient condition, not a classical score equation. Always use inf{F(θ)τ}\inf\{F(\theta) \geq \tau\}, not a derivative argument.

Question 4

In the scalar regression model Y=Xβ0+εY=X\beta_0+\varepsilon, consider the M-estimating equation i=1nXiψ(YiXiβ)=0\sum_{i=1}^n X_i\psi(Y_i-X_i\beta)=0 for a nonlinear score ψ\psi. Which condition most directly guarantees that the population estimating equation is unbiased at β0\beta_0 for arbitrary distributions of XX?

  1. E{ψ(ε)}=0E\{\psi(\varepsilon)\}=0, without any restriction on how ε\varepsilon depends on XX.
  2. E(X)=0E(X)=0, without any restriction on the conditional distribution of ε\varepsilon.
  3. E{ψ(ε)X}=0E\{\psi(\varepsilon)\mid X\}=0 almost surely, together with finite required moments. (correct answer)
  4. E(εX)=0E(\varepsilon\mid X)=0 almost surely, regardless of the particular nonlinear function ψ\psi.
Explanation: When analyzing M-estimating equations, your goal is to ensure the population version of the equation equals zero at the true parameter β0\beta_0. The population estimating equation is E{Xψ(YXβ0)}=E{Xψ(ε)}=0E\{X\psi(Y - X\beta_0)\} = E\{X\psi(\varepsilon)\} = 0, and you need this to hold for arbitrary distributions of XX. The cleanest way to verify this uses the tower property (iterated expectations): E{Xψ(ε)}=E{E(Xψ(ε)X)}=E{XE(ψ(ε)X)}E\{X\psi(\varepsilon)\} = E\{E(X\psi(\varepsilon)\mid X)\} = E\{X \cdot E(\psi(\varepsilon)\mid X)\}. If E{ψ(ε)X}=0E\{\psi(\varepsilon)\mid X\} = 0 almost surely, then each inner expectation is zero regardless of the distribution of XX, making the full expectation zero. This is exactly option C, and it works universally because the condition eliminates the inner term before XX's distribution even matters. Option A gives you only the marginal condition E{ψ(ε)}=0E\{\psi(\varepsilon)\} = 0, which is insufficient when ε\varepsilon and XX are dependent. You can construct cases where E{ψ(ε)}=0E\{\psi(\varepsilon)\} = 0 but E{Xψ(ε)}0E\{X\psi(\varepsilon)\} \neq 0 due to covariance between XX and ψ(ε)\psi(\varepsilon). Option B requiring E(X)=0E(X) = 0 is a condition on XX alone and doesn't control E{Xψ(ε)}E\{X\psi(\varepsilon)\} — a zero-mean XX can still covary with ψ(ε)\psi(\varepsilon). Option D is the classical unbiasedness condition for the linear score (OLS), where ψ(ε)=ε\psi(\varepsilon) = \varepsilon. For a nonlinear ψ\psi, knowing E(εX)=0E(\varepsilon \mid X) = 0 tells you nothing about E{ψ(ε)X}E\{\psi(\varepsilon) \mid X\} in general. Study tip: Whenever a score function is nonlinear, immediately shift your thinking from marginal to conditional moment conditions — the conditional version is almost always the operative requirement, and the tower property is your primary tool for verifying it.

Question 5

An estimator solves i=1nψ(Xi,θ)=0\sum_{i=1}^n\psi(X_i,\theta)=0. A researcher replaces this equation by i=1nc(θ)ψ(Xi,θ)=0\sum_{i=1}^n c(\theta)\psi(X_i,\theta)=0, where c(θ)c(\theta) is continuously differentiable and nonzero near the population root θ0\theta_0. What is the effect of this replacement?

  1. It preserves the local roots and also preserves the correctly computed sandwich asymptotic variance. (correct answer)
  2. It preserves the local roots but multiplies the sandwich asymptotic variance by c(θ0)2c(\theta_0)^2.
  3. It generally changes the local roots because the factor depends on the unknown parameter.
  4. It preserves the asymptotic variance but reverses the roots whenever c(θ0)<0c(\theta_0)<0.
Explanation: When analyzing M-estimators, the key question is: what determines the local roots and the asymptotic variance? Both depend on the estimating equation only through its zero-crossings and its local behavior near θ0\theta_0. Because c(θ)c(\theta) is nonzero near θ0\theta_0, the equation c(θ)ψ(Xi,θ)=0\sum c(\theta)\psi(X_i,\theta)=0 has exactly the same local roots as ψ(Xi,θ)=0\sum\psi(X_i,\theta)=0. Multiplying by a nonzero scalar function cannot create or destroy zeros. Now consider the sandwich variance. For an M-estimator the asymptotic variance is V=A1BATV = A^{-1}BA^{-T}, where A=E[θψ]A = E[\partial_\theta \psi] and B=E[ψψT]B = E[\psi\psi^T]. When you replace ψ\psi by c(θ)ψc(\theta)\psi, both AA and BB are scaled: the new "bread" becomes c(θ0)Ac(\theta_0)A and the new "meat" becomes c(θ0)2Bc(\theta_0)^2 B. The sandwich then gives (c(θ0)A)1(c(θ0)2B)(c(θ0)A)T=c(θ0)1c(θ0)2c(θ0)1A1BAT=V(c(\theta_0)A)^{-1}(c(\theta_0)^2 B)(c(\theta_0)A)^{-T} = c(\theta_0)^{-1}c(\theta_0)^2 c(\theta_0)^{-1}A^{-1}BA^{-T} = V. The c(θ0)c(\theta_0) factors cancel exactly, so the correctly computed sandwich variance is unchanged. This confirms A is correct. Choice B is the classic trap: it forgets that c(θ0)c(\theta_0) appears in both the bread and the meat, so the cancellation is exact. Choice C is wrong because multiplying by a parameter-dependent but nonzero function never changes the roots — the zeros of a product are only where factors are zero. Choice D conflates sign-changes with variance changes; even if c(θ0)<0c(\theta_0)<0, the roots and variance are preserved (the sign of the estimating equation is irrelevant to its zeros). Study tip: Whenever a question asks about rescaling an estimating equation, immediately check whether the scaling factor cancels in both numerator and denominator of the sandwich — it almost always does if the factor is evaluated at the same θ0\theta_0.

Question 6

For the estimating equation U(θ)=i=1n(Xiθ)3=0U(\theta)=\sum_{i=1}^n(X_i-\theta)^3=0, suppose that at a starting value θ(0)\theta^{(0)} the residuals satisfy i(Xiθ(0))3=24\sum_i(X_i-\theta^{(0)})^3=24 and i(Xiθ(0))2=12\sum_i(X_i-\theta^{(0)})^2=12. What is the first Newton–Raphson iterate?

  1. θ(1)=θ(0)2\theta^{(1)}=\theta^{(0)}-2
  2. θ(1)=θ(0)23\theta^{(1)}=\theta^{(0)}-\tfrac{2}{3}
  3. θ(1)=θ(0)+2\theta^{(1)}=\theta^{(0)}+2
  4. θ(1)=θ(0)+23\theta^{(1)}=\theta^{(0)}+\tfrac{2}{3} (correct answer)
Explanation: When you encounter a question about Newton–Raphson iteration, your first instinct should be to write down the update formula precisely: θ(1)=θ(0)U(θ(0))U(θ(0))\theta^{(1)} = \theta^{(0)} - \frac{U(\theta^{(0)})}{U'(\theta^{(0)})}, where U(θ)U'(\theta) is the derivative of the estimating equation with respect to θ\theta. Here, U(θ)=i(Xiθ)3U(\theta) = \sum_i (X_i - \theta)^3. Differentiating with respect to θ\theta gives U(θ)=3i(Xiθ)2U'(\theta) = -3\sum_i (X_i - \theta)^2. At θ(0)\theta^{(0)}, you're told U(θ(0))=24U(\theta^{(0)}) = 24 and i(Xiθ(0))2=12\sum_i (X_i - \theta^{(0)})^2 = 12, so U(θ(0))=3(12)=36U'(\theta^{(0)}) = -3(12) = -36. Plugging in: θ(1)=θ(0)2436=θ(0)+23\theta^{(1)} = \theta^{(0)} - \frac{24}{-36} = \theta^{(0)} + \frac{2}{3}, confirming answer D. Each wrong answer reflects a specific error. A results from forgetting the factor of 3-3 in the derivative entirely and mishandling the sign, producing 2-2. B drops the negative sign on UU' but keeps the magnitude, yielding 23-\tfrac{2}{3} — a sign error that's easy to make if you forget the chain rule contributes a 1-1 when differentiating (Xiθ)(X_i - \theta). C uses +2+2, which could come from incorrectly computing the derivative as 3(12)=36-3(12) = -36 but then dividing 2424 by 1212 alone (ignoring the 3), giving the wrong magnitude with the right sign direction. As a study tip, always differentiate U(θ)U(\theta) carefully before substituting — the negative from the chain rule on (Xiθ)(X_i - \theta) is the most common source of sign errors in Newton–Raphson problems on graduate exams.

Question 7

Consider the criterion Qn(θ)=θ4/4θ2/2θXnQ_n(\theta)=\theta^4/4-\theta^2/2-\theta\overline X_n, where E(X)=0E(X)=0 and Xn0\overline X_n\to0 in probability. Its estimating equation is θ3θXn=0\theta^3-\theta-\overline X_n=0. Which conclusion about consistency is most accurate?

  1. A global minimizer need not converge to one unique value because the population criterion has equal minima at 1-1 and 11. (correct answer)
  2. The root near 00 is the unique consistent minimizer because it solves the limiting estimating equation.
  3. Every root is consistent for 00 because the random term Xn\overline X_n converges to zero.
  4. The estimator must diverge because the limiting estimating equation has more than one finite root.
Explanation: When analyzing M-estimators or extremum estimators, you need to distinguish between the sample criterion and its population (limiting) counterpart. Consistency arguments depend critically on the structure of the limiting criterion, not just the sample version. Here, as Xn0\overline{X}_n \to 0, the sample criterion Qn(θ)=θ4/4θ2/2θXnQ_n(\theta) = \theta^4/4 - \theta^2/2 - \theta\overline{X}_n converges to the population criterion Q(θ)=θ4/4θ2/2Q(\theta) = \theta^4/4 - \theta^2/2. Taking the derivative of this limit gives θ3θ=θ(θ1)(θ+1)=0\theta^3 - \theta = \theta(\theta-1)(\theta+1) = 0, which has roots at θ{1,0,1}\theta \in \{-1, 0, 1\}. Checking second derivatives, Q(θ)=3θ21Q''(\theta) = 3\theta^2 - 1, reveals that θ=±1\theta = \pm 1 are local (and in fact global) minima with equal criterion values Q(±1)=1/4Q(\pm 1) = -1/4, while θ=0\theta = 0 is a local maximum. Because the limiting criterion achieves its minimum at two distinct points, a global minimizer of QnQ_n can converge to either 1-1 or 11 depending on the realized path of Xn\overline{X}_n, making the global minimizer inconsistent for any single value. Answer A correctly captures this non-uniqueness. Answer B is wrong because θ=0\theta = 0 solves the limiting estimating equation but is a maximizer, not a minimizer — a critical error in reasoning. Answer C is wrong because Xn0\overline{X}_n \to 0 does not collapse the two minima into one; the symmetry-breaking is insufficient when both minima are exactly equal in the limit. Answer D is wrong because multiple roots do not imply divergence; they imply non-uniqueness of the limit point. As a study tip: always verify whether candidate roots of the limiting estimating equation are minima or maxima, and check whether the minimum is uniquely attained — both conditions are necessary for a consistent global minimizer.

Question 8

Suppose an M-estimator satisfies n(θ^nθ0)N(0,V)\sqrt n(\widehat\theta_n-\theta_0)\mathrel{\Rightarrow}N(0,V), where θ0>0\theta_0>0. The parameter is reexpressed as η=θ2\eta=\theta^2, and the estimator is η^n=θ^n2\widehat\eta_n=\widehat\theta_n^2. What is the limiting distribution of n(η^nη0)\sqrt n(\widehat\eta_n-\eta_0)?

  1. N(0,V)N(0,V), because reparameterization does not change the root of an estimating equation.
  2. N(0,2θ0V)N(0,2\theta_0V), because the transformed variance is multiplied by the first derivative.
  3. N(0,θ04V)N(0,\theta_0^4V), because the original variance is multiplied by the squared parameter.
  4. N(0,4θ02V)N(0,4\theta_0^2V), because the transformed variance uses the squared derivative. (correct answer)
Explanation: When you see a question involving a transformation of an asymptotically normal estimator, your first instinct should be the Delta Method: if n(θ^nθ0)N(0,V)\sqrt{n}(\hat{\theta}_n - \theta_0) \Rightarrow N(0, V) and gg is differentiable at θ0\theta_0, then n(g(θ^n)g(θ0))N(0,[g(θ0)]2V)\sqrt{n}(g(\hat{\theta}_n) - g(\theta_0)) \Rightarrow N(0, [g'(\theta_0)]^2 V). Here, g(θ)=θ2g(\theta) = \theta^2, so g(θ)=2θg'(\theta) = 2\theta, giving g(θ0)=2θ0g'(\theta_0) = 2\theta_0. The asymptotic variance of the transformed estimator is (2θ0)2V=4θ02V(2\theta_0)^2 \cdot V = 4\theta_0^2 V, confirming that D is correct: n(η^nη0)N(0,4θ02V)\sqrt{n}(\hat{\eta}_n - \eta_0) \Rightarrow N(0, 4\theta_0^2 V). A is wrong because reparameterization absolutely changes the variance — the estimating-equation root is invariant (that's the invariance property of M-estimators), but the asymptotic distribution of the transformed estimator still picks up the Jacobian factor from the Delta Method. B confuses squaring the derivative with just multiplying by it. The variance transformation uses [g(θ0)]2[g'(\theta_0)]^2, not g(θ0)g'(\theta_0) alone. Multiplying by 2θ02\theta_0 rather than 4θ024\theta_0^2 is a common algebraic slip. C arrives at θ04V\theta_0^4 V, which would correspond to g(θ0)=θ02g'(\theta_0) = \theta_0^2 — the derivative of θ3/3\theta^3/3, not θ2\theta^2. This reflects a failure to actually differentiate gg and instead substituting the function value itself. Strategy tip: Whenever a transformation is involved, immediately write down g(θ)g(\theta), compute g(θ0)g'(\theta_0), and apply [g(θ0)]2V[g'(\theta_0)]^2 V. The squaring of the derivative is the step most students miss under pressure.

Question 9

Let the distribution of XX be continuous, with finite mean μ\mu and unique median mm, where μm\mu\ne m. For a fixed constant λ>0\lambda>0, define θ^n\widehat\theta_n to minimize n1i=1n{(Xiθ)2+λXiθ}n^{-1}\sum_{i=1}^n\{(X_i-\theta)^2+\lambda|X_i-\theta|\}. Under standard uniform convergence and uniqueness conditions, to what population value does θ^n\widehat\theta_n converge?

  1. It converges to μ\mu, the solution of 2(θμ)=02(\theta-\mu)=0, because the quadratic term dominates and the absolute-deviation term contributes nothing to the first-order condition.
  2. It converges to mm, the solution of E{sign(θX)}=0E\{\operatorname{sign}(\theta-X)\}=0, because the absolute-deviation term dominates and the quadratic term contributes nothing to the first-order condition.
  3. It converges to the unique solution of 2(θμ)+λE{sign(θX)}=02(\theta-\mu)+\lambda E\{\operatorname{sign}(\theta-X)\}=0, which reflects contributions from both the quadratic and absolute-deviation components. (correct answer)
  4. It converges to the unique solution of 2(θμ)λE{sign(θX)}=02(\theta-\mu)-\lambda E\{\operatorname{sign}(\theta-X)\}=0, because differentiating Xθ|X-\theta| with respect to θ\theta produces a negative sign.
Explanation: When you encounter a loss function that combines multiple terms, your first instinct should be to differentiate the expected loss with respect to θ\theta and set it to zero — each term contributes its own first-order condition, and neither can simply be ignored. Here, the population objective is Q(θ)=E{(Xθ)2+λXθ}Q(\theta) = E\{(X-\theta)^2 + \lambda|X-\theta|\}. Differentiating with respect to θ\theta: the quadratic term gives ddθE(Xθ)2=2E(Xθ)=2(θμ)\frac{d}{d\theta}E(X-\theta)^2 = -2E(X-\theta) = 2(\theta - \mu), and the absolute-value term gives ddθE(λXθ)=λE{sign(θX)}\frac{d}{d\theta}E(\lambda|X-\theta|) = \lambda E\{\text{sign}(\theta - X)\}, since ddθXθ=sign(θX)\frac{d}{d\theta}|X - \theta| = \text{sign}(\theta - X). Setting the sum to zero yields 2(θμ)+λE{sign(θX)}=02(\theta - \mu) + \lambda E\{\text{sign}(\theta - X)\} = 0, which is exactly what C states. By the law of large numbers and standard M-estimation theory, θ^n\widehat\theta_n converges to the unique root of this equation. A is wrong because it claims the quadratic term dominates and discards the absolute-deviation contribution entirely — both terms remain present in the first-order condition for any fixed λ>0\lambda > 0. B makes the symmetric error: it discards the quadratic term, which would only be appropriate if the loss were pure LAD (least absolute deviations), not a combined criterion. D has the sign wrong. Differentiating Xθ|X - \theta| with respect to θ\theta yields +sign(θX)+\text{sign}(\theta - X), not sign(θX)-\text{sign}(\theta - X). The chain rule introduces a factor of 1-1 from (Xθ)θ\frac{\partial(X-\theta)}{\partial\theta}, but that negative is already absorbed into sign(θX)\text{sign}(\theta - X) rather than sign(Xθ)\text{sign}(X - \theta). Study tip: For any combined loss, always differentiate term by term and keep every contribution — "dominance" arguments only apply in limiting regimes (e.g., λ0\lambda \to 0 or λ\lambda \to \infty), not for fixed λ\lambda.

Question 10

For a location M-estimator defined by iψ(Xiθ)=0\sum_i\psi(X_i-\theta)=0, suppose A=E{ψ(Xθ0)}A=E\{\psi'(X-\theta_0)\} exists and is finite and nonzero. Which statement about its influence function is correct?

  1. The influence function is bounded whenever the loss ρ\rho is convex and symmetric, regardless of whether the score ψ\psi itself is bounded.
  2. The influence function is bounded whenever ψ\psi is bounded; the nonzero sensitivity constant AA only rescales it without affecting boundedness. (correct answer)
  3. The influence function is bounded whenever the estimator is Fisher consistent at the assumed model distribution, because Fisher consistency implies a zero score expectation.
  4. The influence function is bounded only when ψ\psi is strictly increasing and differentiable everywhere on the real line, since these properties ensure a well-defined sensitivity.
Explanation: When analyzing influence functions for M-estimators, your central focus should be the explicit formula connecting the influence function to the defining equation. For a location M-estimator solving iψ(Xiθ)=0\sum_i \psi(X_i - \theta) = 0, the influence function at a point xx is derived via the implicit function theorem and equals: IF(x)=ψ(xθ0)A,A=E{ψ(Xθ0)}\text{IF}(x) = \frac{\psi(x - \theta_0)}{A}, \quad A = E\{\psi'(X - \theta_0)\} This formula makes the key insight immediate: the influence function is simply ψ\psi rescaled by the constant AA. Since AA is assumed finite and nonzero, it acts purely as a scaling factor. Therefore, the influence function is bounded if and only if ψ\psi itself is bounded — confirming that B is correct. The constant AA rescales the range of ψ\psi but cannot introduce or remove boundedness. Choice A is wrong because convexity and symmetry of the loss ρ\rho do not guarantee that ψ=ρ\psi = \rho' is bounded. For example, the squared loss is convex and symmetric, yet its score ψ(x)=x\psi(x) = x is unbounded. Choice C confuses two distinct properties. Fisher consistency means Eθ0[ψ(Xθ0)]=0E_{\theta_0}[\psi(X - \theta_0)] = 0, which is a centering condition ensuring the estimating equation is unbiased. It says nothing about whether ψ\psi takes large values at outlying points — boundedness of ψ\psi is a separate requirement. Choice D is incorrect because strict monotonicity and global differentiability of ψ\psi are sufficient for uniqueness and regularity of the estimator, not for boundedness of the influence function. A strictly increasing ψ(x)=x\psi(x) = x yields an unbounded influence function. A useful study tip: whenever you see an influence function question for M-estimators, write down IF(x)=ψ(xθ0)/A\text{IF}(x) = \psi(x-\theta_0)/A immediately — then every robustness property of the influence function reduces directly to a property of ψ\psi.