Statistics Graduate Level Quiz: Influence Functions
10 questions · exam conditions
0:00
Influence FunctionsQuestion 1 of 10

Let T(F)T(F) be the median of a distribution FF having a unique median mm and a density continuous and positive at mm. Suppose f(m)=0.20f(m)=0.20. A small contamination mass is placed at a point z>mz>m.

For Fε=(1ε)F+εΔzF_{\varepsilon}=(1-\varepsilon)F+\varepsilon\Delta_z, what is the influence function IF(z;T,F)\operatorname{IF}(z;T,F)?

2.5-2.5, because contamination above the median shifts the distribution function downward at the original median.
2.52.5, because the median must move upward to restore cumulative probability one-half.
5.05.0, because the contamination effect is scaled by the reciprocal of the density at the median.
0.100.10, because half of the contamination mass contributes to the first-order median shift.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Influence Functions

Practice Influence Functions in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Influence Functions, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Let T(F)T(F) be the median of a distribution FF having a unique median mm and a density continuous and positive at mm. Suppose f(m)=0.20f(m)=0.20. A small contamination mass is placed at a point z>mz>m.

For Fε=(1ε)F+εΔzF_{\varepsilon}=(1-\varepsilon)F+\varepsilon\Delta_z, what is the influence function IF(z;T,F)\operatorname{IF}(z;T,F)?

  1. 2.5-2.5, because contamination above the median shifts the distribution function downward at the original median.
  2. 2.52.5, because the median must move upward to restore cumulative probability one-half. (correct answer)
  3. 5.05.0, because the contamination effect is scaled by the reciprocal of the density at the median.
  4. 0.100.10, because half of the contamination mass contributes to the first-order median shift.
Explanation: Whenever you encounter influence function questions, your anchor should be the formal definition: IF(z;T,F)=limε0T(Fε)T(F)ε\operatorname{IF}(z; T, F) = \lim_{\varepsilon \to 0} \frac{T(F_\varepsilon) - T(F)}{\varepsilon}, which measures the first-order sensitivity of a functional to infinitesimal contamination at point zz. For the sample median, the derivation proceeds by solving for where the contaminated distribution Fε=(1ε)F+εΔzF_\varepsilon = (1-\varepsilon)F + \varepsilon\Delta_z achieves cumulative probability one-half. Setting Fε(mε)=1/2F_\varepsilon(m_\varepsilon) = 1/2 and differentiating implicitly with respect to ε\varepsilon at ε=0\varepsilon = 0 yields the closed-form result: IF(z;T,F)=sgn(zm)2f(m)\operatorname{IF}(z; T, F) = \frac{\operatorname{sgn}(z - m)}{2f(m)} Since z>mz > m, sgn(zm)=+1\operatorname{sgn}(z - m) = +1, so IF(z;T,F)=12×0.20=10.40=2.5\operatorname{IF}(z; T, F) = \frac{1}{2 \times 0.20} = \frac{1}{0.40} = 2.5. The median must shift upward to compensate for the extra mass above it — confirming B is correct. Choice A gets the sign wrong. Contamination above the median pushes the median upward, not downward; the distribution function at the original median increases, forcing the quantile to move right, not left. Choice C arrives at 5.05.0 by using 1/f(m)1/f(m) alone, forgetting the factor of 22 in the denominator — a classic algebra slip. Choice D produces 0.100.10 by multiplying ε/2\varepsilon/2 thinking proportionally, which confuses a finite-sample heuristic with the correct asymptotic derivative formula. Your study tip: always memorize the median's influence function as sgn(zm)2f(m)\frac{\operatorname{sgn}(z-m)}{2f(m)}. The factor of 2 reflects the two-sided nature of the median condition, and f(m)f(m) in the denominator means the influence is larger when the density is flat — intuition worth retaining.

Question 2

Define T(F)=EF(X2)/EF(X)T(F)=E_F(X^2)/E_F(X), where EF(X)=2E_F(X)=2 and EF(X2)=5E_F(X^2)=5. Consider contamination at z=4z=4.

Using the first-order contamination derivative, what is IF(4;T,F)\operatorname{IF}(4;T,F)?

  1. 1.51.5, because only the contamination effect on the numerator remains after centering.
  2. 8.08.0, because the contamination contributes its squared value relative to the first moment.
  3. 5.55.5, because the numerator derivative is divided by the original first moment.
  4. 3.03.0, because both the numerator and denominator change under contamination. (correct answer)
Explanation: When you see a functional of the form T(F)=EF(X2)/EF(X)T(F) = E_F(X^2)/E_F(X), recognize it as a ratio functional. The influence function requires applying the quotient rule to the contaminated functional T(Fϵ)=EFϵ(X2)/EFϵ(X)T(F_\epsilon) = E_{F_\epsilon}(X^2)/E_{F_\epsilon}(X), where Fϵ=(1ϵ)F+ϵδzF_\epsilon = (1-\epsilon)F + \epsilon\delta_z. Under contamination by a point mass at zz, the numerator becomes (1ϵ)EF(X2)+ϵz2(1-\epsilon)E_F(X^2) + \epsilon z^2 and the denominator becomes (1ϵ)EF(X)+ϵz(1-\epsilon)E_F(X) + \epsilon z. Differentiating T(Fϵ)T(F_\epsilon) with respect to ϵ\epsilon at ϵ=0\epsilon=0 via the quotient rule gives: IF(z;T,F)=(z2EF(X2))EF(X)EF(X2)(zEF(X))[EF(X)]2\operatorname{IF}(z;T,F) = \frac{(z^2 - E_F(X^2)) \cdot E_F(X) - E_F(X^2)\cdot(z - E_F(X))}{[E_F(X)]^2} Plugging in z=4z=4, EF(X)=2E_F(X)=2, EF(X2)=5E_F(X^2)=5: =(165)25(42)4=22104=124=3.0= \frac{(16-5)\cdot 2 - 5\cdot(4-2)}{4} = \frac{22 - 10}{4} = \frac{12}{4} = 3.0 This confirms D is correct. Choice A incorrectly ignores the denominator's contribution entirely. Choice B simply reports z2/EF(X)=8z^2/E_F(X) = 8, confusing the raw contaminated value with the influence function derivative. Choice C computes only (z2EF(X2))/EF(X)=11/2=5.5(z^2 - E_F(X^2))/E_F(X) = 11/2 = 5.5, applying only the numerator's centered derivative without accounting for the denominator's shift. As a study rule: whenever T(F)T(F) is a ratio, always apply the full quotient rule for the IF — forgetting the denominator's derivative is the single most common error on these problems.

Question 3

A location functional T(F)T(F) is defined by EF{ψ(XT(F))}=0E_F\{\psi(X-T(F))\}=0, where ψ(u)=max(1.5,min(u,1.5))\psi(u)=\max(-1.5,\min(u,1.5)). At the model distribution, T(F)=0T(F)=0, XX is standard normal, and PF(X1.5)=0.8664P_F(|X|\le 1.5)=0.8664.

Approximately what is the influence function at a contamination point z=4z=4?

  1. 0.5780.578, obtained by dividing the expected derivative by the clipped contamination score.
  2. 1.5001.500, obtained by using the clipped score without accounting for sensitivity.
  3. 1.7311.731, obtained by scaling the clipped score by the inverse expected derivative. (correct answer)
  4. 4.6174.617, obtained by scaling the unclipped residual by the inverse expected derivative.
Explanation: When you encounter influence function (IF) questions, the key formula to internalize is: IF(z;T,F)=ψ(zT(F))EF[ψ(XT(F))]\text{IF}(z; T, F) = \frac{\psi(z - T(F))}{E_F[\psi'(X - T(F))]} This comes from differentiating the estimating equation EF[ψ(XT(F))]=0E_F[\psi(X - T(F))] = 0 with respect to contamination at point zz. The numerator captures how the score function evaluates the outlier; the denominator measures the estimator's average sensitivity. At T(F)=0T(F) = 0 and z=4z = 4, the numerator is ψ(4)=min(4,1.5)=1.5\psi(4) = \min(4, 1.5) = 1.5 — the score is clipped. For the denominator, since ψ\psi is the Huber-type clip function, ψ(u)=1\psi'(u) = 1 when u1.5|u| \le 1.5 and 00 otherwise. So EF[ψ(X)]=PF(X1.5)=0.8664E_F[\psi'(X)] = P_F(|X| \le 1.5) = 0.8664. The influence function is therefore: IF(4)=1.50.86641.731\text{IF}(4) = \frac{1.5}{0.8664} \approx 1.731 This confirms C is correct. A is wrong because it inverts the formula — dividing the expected derivative by the score, rather than the score by the expected derivative. B simply reads off the clipped score (1.5) and ignores the denominator entirely, failing to account for how steeply the estimator responds to perturbations. D uses the raw, unclipped value z=4z = 4 in the numerator, as if ψ\psi were the identity function — a fundamental misreading of what the redescending/bounded ψ\psi does to outliers. Study tip: For M-estimator influence functions, always write down the estimating equation first, then differentiate. The denominator E[ψ]E[\psi'] is never 1 unless ψ\psi is the identity — don't skip it.

Question 4

Two location estimators are under consideration. The first has a bounded influence function at the model distribution, whereas the second has an unbounded influence function. No additional information about their finite-sample behavior is available.

Which conclusion is justified solely from this information?

  1. The first estimator has a positive finite-sample breakdown point, whereas the second estimator has breakdown point zero.
  2. The first estimator has bounded first-order sensitivity to infinitesimal point-mass contamination at the model. (correct answer)
  3. The first estimator has smaller asymptotic variance than the second estimator under every distribution in the model.
  4. The first estimator remains uniformly accurate under contamination by any fixed positive fraction of arbitrary observations.
Explanation: When studying robustness theory, the key is distinguishing precisely what each concept measures. The influence function (IF) quantifies the asymptotic effect of an infinitesimal point-mass contamination on an estimator's value. Formally, if TT is a functional and FF is the model distribution, the IF at point xx is: IF(x;T,F)=limϵ0T((1ϵ)F+ϵδx)T(F)ϵ\text{IF}(x; T, F) = \lim_{\epsilon \to 0} \frac{T((1-\epsilon)F + \epsilon\delta_x) - T(F)}{\epsilon} A bounded influence function means this derivative is uniformly bounded over all contamination points xx, which is precisely a statement about first-order sensitivity to infinitesimal contamination. This is exactly what answer B states — it's a direct, definitional consequence of having a bounded IF, making B correct. Answer A is tempting but unjustified. Bounded IF and positive breakdown point are related robustness ideas, but neither implies the other. An estimator can have a bounded IF yet still have breakdown point zero (e.g., certain M-estimators with bounded scores but non-robust scale choices). The IF is a local, infinitesimal concept; breakdown point is a global, finite-sample concept. You cannot infer one from the other without additional information. Answer C is false — boundedness of the IF says nothing about asymptotic variance. In fact, high-breakdown estimators often trade efficiency for robustness, meaning the first estimator could have larger variance. Answer D overreaches significantly. A bounded IF only guarantees infinitesimal-contamination stability; it provides no uniform guarantee over a fixed positive contamination fraction ϵ>0\epsilon > 0. Study tip: Always match robustness concepts to their precise scope — IF is infinitesimal and asymptotic, breakdown point is finite and global. Exam distractors frequently blur this boundary.

Question 5

Let T(F)={EF(X)}2T(F)=\{E_F(X)\}^2, and suppose the true distribution has EF(X)=0E_F(X)=0 and 0<VarF(X)=σ2<0<\operatorname{Var}_F(X)=\sigma^2<\infty. The plug-in estimator is T(Fn)=Xˉn2T(F_n)=\bar X_n^2.

Which statement correctly describes the influence function and the first nondegenerate asymptotic behavior of this estimator?

  1. The influence function is zero, and nXˉn2n\bar X_n^2 converges in distribution to σ2χ12\sigma^2\chi_1^2. (correct answer)
  2. The influence function is 2X2X, and nXˉn2\sqrt n\,\bar X_n^2 converges to a centered normal distribution.
  3. The influence function is X2σ2X^2-\sigma^2, and nXˉn2\sqrt n\,\bar X_n^2 has asymptotic variance 2σ42\sigma^4.
  4. The influence function is zero, and nXˉn2n\bar X_n^2 converges in probability to the constant σ2\sigma^2.
Explanation: When a functional T(F)T(F) is smooth, its plug-in estimator inherits a n\sqrt{n}-rate via the delta method, driven by a nonzero influence function. But when the influence function vanishes, the estimator is second-order, and you need to look one level deeper. The influence function of T(F)={EF(X)}2T(F) = \{E_F(X)\}^2 is the Gateaux derivative: IF(x;T,F)=ddϵϵ=0T((1ϵ)F+ϵδx)=2EF(X)(xEF(X)).\text{IF}(x; T, F) = \frac{d}{d\epsilon}\Big|_{\epsilon=0} T((1-\epsilon)F + \epsilon\delta_x) = 2E_F(X)\cdot(x - E_F(X)). Since EF(X)=0E_F(X) = 0 by assumption, the influence function equals zero. This means the standard n\sqrt{n} central limit theorem gives a degenerate (zero) limit for nXˉn2\sqrt{n}\,\bar{X}_n^2, so you must rescale by nn instead. By the CLT, nXˉndN(0,σ2)\sqrt{n}\,\bar{X}_n \xrightarrow{d} N(0, \sigma^2), and therefore nXˉn2=(nXˉn)2dσ2χ12,n\bar{X}_n^2 = \left(\sqrt{n}\,\bar{X}_n\right)^2 \xrightarrow{d} \sigma^2 \chi_1^2, confirming answer A. Answer B is wrong on both counts: the influence function is not 2X2X (that would require EF(X)0E_F(X)\neq 0), and nXˉn20\sqrt{n}\,\bar{X}_n^2 \to 0 in probability, not a normal. Answer C confuses T(F)T(F) with the variance functional VarF(X)=E[X2](E[X])2\text{Var}_F(X) = E[X^2] - (E[X])^2; the influence function X2σ2X^2 - \sigma^2 belongs to a different functional entirely. Answer D gets the rescaling right but claims convergence in probability to a constant — that would require a degenerate limit, but σ2χ12\sigma^2\chi_1^2 is a genuine random variable, not a constant. Study tip: Whenever EF(X)=0E_F(X) = 0 makes the first-order influence function vanish, immediately switch to a second-order analysis — rescale by nn, apply the continuous mapping theorem to the squared CLT limit, and expect a chi-squared (not normal) distribution.

Question 6

Let XX be Bernoulli with success probability p=0.25p=0.25. The parameter of interest is the log-odds functional θ(F)=log{p/(1p)}\theta(F)=\log\{p/(1-p)\}, where p=EF(X)p=E_F(X).

What is the influence function for θ\theta at the contamination point z=1z=1?

  1. 0.750.75, because this is the influence of a success on the probability functional.
  2. 3.003.00, because the probability influence is divided only by the success probability.
  3. 4.004.00, because the probability influence is scaled by the derivative of the log-odds. (correct answer)
  4. 5.335.33, because the reciprocal Bernoulli variance is the influence of one success.
Explanation: Whenever you encounter influence function questions, your instinct should be to apply the chain rule through the statistical functional. The influence function of a transformed functional θ(F)=g(μ(F))\theta(F) = g(\mu(F)) is: IF(z;θ,F)=g(μ)IF(z;μ,F)\text{IF}(z; \theta, F) = g'(\mu) \cdot \text{IF}(z; \mu, F) Here, μ(F)=p=EF[X]\mu(F) = p = E_F[X] and g(p)=log{p/(1p)}g(p) = \log\{p/(1-p)\}. The influence function for the mean functional at point zz is simply IF(z;μ,F)=zp\text{IF}(z; \mu, F) = z - p. At z=1z = 1 and p=0.25p = 0.25, this equals 10.25=0.751 - 0.25 = 0.75. The derivative of the log-odds is g(p)=1/[p(1p)]g'(p) = 1/[p(1-p)]. At p=0.25p = 0.25: g(0.25)=1/(0.25×0.75)=1/0.18755.33g'(0.25) = 1/(0.25 \times 0.75) = 1/0.1875 \approx 5.33. Multiplying gives 5.33×0.75=4.005.33 \times 0.75 = 4.00, confirming C is correct. Choice A stops too early — 0.750.75 is the influence function for the mean functional, not the log-odds. It ignores the chain rule transformation entirely. Choice B computes 0.75/0.25=3.000.75 / 0.25 = 3.00, dividing only by pp rather than by the full Bernoulli variance p(1p)p(1-p), which misapplies the derivative of the log-odds. Choice D reports 5.335.33, which is g(p)g'(p) alone — the derivative of the log-odds evaluated at pp — but forgets to multiply by the mean's influence function zp=0.75z - p = 0.75. As a study tip: always decompose functional influence functions using the chain rule — identify the inner functional's IF first, then scale by the outer function's derivative. This two-step process prevents every trap present in this question.

Question 7

A regular estimator TnT_n has the asymptotic linear representation TnT(F)=n1i=1nIF(Xi;T,F)+op(n1/2)T_n-T(F)=n^{-1}\sum_{i=1}^n \operatorname{IF}(X_i;T,F)+o_p(n^{-1/2}). At the distribution of interest, IF(X;T,F)=X22\operatorname{IF}(X;T,F)=X^2-2, EF(X2)=2E_F(X^2)=2, and EF(X4)=10E_F(X^4)=10.

What is the asymptotic variance of n{TnT(F)}\sqrt{n}\{T_n-T(F)\}?

  1. 22, because the influence function is centered by the second moment.
  2. 1010, because the fourth moment determines the limiting variance directly.
  3. 88, because the cross term contributes twice the squared second moment.
  4. 66, because the limiting variance is the second moment of the influence function. (correct answer)
Explanation: Whenever you encounter an asymptotic linear representation, your first instinct should be: the asymptotic variance of n(TnT(F))\sqrt{n}(T_n - T(F)) equals the variance of the influence function. By the CLT applied to the i.i.d. sum, n(TnT(F))dN(0,VarF(IF(X;T,F)))\sqrt{n}(T_n - T(F)) \xrightarrow{d} N(0, \text{Var}_F(\operatorname{IF}(X;T,F))). So the task reduces to computing VarF(X22)\text{Var}_F(X^2 - 2). Since EF(X2)=2E_F(X^2) = 2, the influence function IF(X;T,F)=X22\operatorname{IF}(X;T,F) = X^2 - 2 is already mean-zero (a requirement for valid influence functions). Its variance is: VarF(X22)=EF[(X22)2]=EF[X44X2+4]=EF(X4)4EF(X2)+4=108+4=6.\text{Var}_F(X^2 - 2) = E_F[(X^2-2)^2] = E_F[X^4 - 4X^2 + 4] = E_F(X^4) - 4E_F(X^2) + 4 = 10 - 8 + 4 = 6. This confirms D is correct: the asymptotic variance is 66, the second moment of the influence function (i.e., its variance, since it's centered). A is wrong because it claims the variance equals 2=EF(X2)2 = E_F(X^2), confusing the mean of X2X^2 with the variance of the influence function — these are entirely different quantities. B incorrectly equates the asymptotic variance with EF(X4)=10E_F(X^4) = 10, ignoring that you must subtract the squared mean of X2X^2 (and cross terms) when computing the variance. C arrives at 88 by subtracting only 4EF(X2)=84E_F(X^2) = 8 from 1010, forgetting to add back the constant term +4+4 in the expansion. Your study tip: always expand EF[(IF)2]E_F[(\operatorname{IF})^2] fully using Var(Y)=E[Y2](E[Y])2\text{Var}(Y) = E[Y^2] - (E[Y])^2, and verify the influence function is mean-zero before proceeding.

Question 8

For an empirical distribution FnF_n, an estimator is well approximated by a smooth functional T(Fn)T(F_n). One observation xix_i is replaced by a new value zz, producing the empirical distribution FnrepF_n^{\mathrm{rep}}.

Which expression gives the appropriate first-order approximation to T(Fnrep)T(Fn)T(F_n^{\mathrm{rep}})-T(F_n)?

  1. n1{IF(z;T,Fn)IF(xi;T,Fn)}n^{-1}\{\operatorname{IF}(z;T,F_n)-\operatorname{IF}(x_i;T,F_n)\}, because replacement both adds and removes mass. (correct answer)
  2. n1IF(z;T,Fn)n^{-1}\operatorname{IF}(z;T,F_n), because the replacement adds contamination at the new observation.
  3. (n1)1IF(xi;T,Fn)(n-1)^{-1}\operatorname{IF}(x_i;T,F_n), because the deleted observation determines the entire first-order change.
  4. n1{IF(z;T,Fn)+IF(xi;T,Fn)}n^{-1}\{\operatorname{IF}(z;T,F_n)+\operatorname{IF}(x_i;T,F_n)\}, because both observations contribute perturbations of equal sign.
Explanation: When working with influence functions and empirical distributions, the key is tracking exactly how the distribution changes under each operation — adding mass, removing mass, or both simultaneously. When you replace observation xix_i with a new value zz, you are doing two things at once: removing a point mass of size 1/n1/n at xix_i and adding a point mass of size 1/n1/n at zz. The influence function IF(x;T,F)\operatorname{IF}(x; T, F) captures the first-order effect of adding an infinitesimal point mass at xx. By linearity of the first-order expansion, removing mass at xix_i contributes n1IF(xi;T,Fn)-n^{-1}\operatorname{IF}(x_i; T, F_n) and adding mass at zz contributes +n1IF(z;T,Fn)+n^{-1}\operatorname{IF}(z; T, F_n), giving the net approximation n1{IF(z;T,Fn)IF(xi;T,Fn)}n^{-1}\{\operatorname{IF}(z;T,F_n)-\operatorname{IF}(x_i;T,F_n)\}. This confirms A is correct. Choice B is wrong because it only accounts for the addition of zz and completely ignores that xix_i is simultaneously removed — a half-treatment of the perturbation. Choice C is wrong on two counts: it uses (n1)1(n-1)^{-1} instead of n1n^{-1}, and it ignores the contribution from zz entirely, focusing only on the deleted point. Choice D is a sign error — it adds both influence function values instead of subtracting, which would correspond to adding mass at both locations simultaneously rather than a replacement. A reliable memory anchor: replacement = addition minus deletion. Whenever a question involves swapping one observation for another, write out the two separate perturbations explicitly and track their signs before combining — this prevents both the sign error in D and the incomplete accounting in B and C.

Question 9

In a random-design linear regression model, the ordinary least-squares functional satisfies E{X(YXTβ)}=0E\{X(Y-X^{\mathsf T}\beta)\}=0, and M=E(XXT)M=E(XX^{\mathsf T}) is nonsingular. Consider contamination at a point (x,y)(x,y) whose residual r=yxTβr=y-x^{\mathsf T}\beta remains fixed and nonzero while the norm of xx increases.

What does the influence-function calculation imply about ordinary least squares in this sequence of contamination points?

  1. Its influence approaches zero because the fixed residual becomes negligible relative to the increasing covariate norm.
  2. Its influence remains bounded because only the residual, rather than the covariate value, enters the estimating equation.
  3. Its influence is generally unbounded because M1xrM^{-1}xr grows with leverage when the residual is fixed. (correct answer)
  4. Its influence is undefined because influence functions cannot be formed for random-design regression functionals.
Explanation: When analyzing robustness of an estimator, your first instinct should be to examine its influence function — the functional derivative measuring how sensitive the estimator is to infinitesimal contamination at a single point. For OLS in the random-design setting, the regression functional β(F)=M1E(XY)\beta(F) = M^{-1}E(XY) yields an influence function proportional to M1xrM^{-1}x r, where r=yxTβr = y - x^\mathsf{T}\beta is the residual at the contamination point (x,y)(x, y). This is exactly why C is correct. When rr stays fixed but x\|x\| \to \infty, the term M1xrM^{-1}xr grows without bound — its norm scales with x\|x\|. This is precisely the leverage effect: a point far from the center of the covariate space exerts unbounded influence on the fitted coefficients, even with a modest residual. OLS has no mechanism to down-weight high-leverage points, so its influence function is unbounded in the covariate direction. A reverses the logic. The fixed residual doesn't "cancel out" the growing covariate norm — both enter multiplicatively in M1xrM^{-1}xr, and the growing norm dominates. B is factually wrong: the estimating equation E{X(YXTβ)}=0E\{X(Y - X^\mathsf{T}\beta)\} = 0 explicitly involves XX, so the covariate absolutely enters the influence calculation. D is false — influence functions are perfectly well-defined for smooth statistical functionals, including random-design regression, as long as the functional is Gâteaux-differentiable at the model distribution. As a study strategy: whenever a question pairs "fixed residual" with "growing covariate," immediately think leverage unboundedness. This is a classic argument for why robust alternatives like M-estimators with bounded ψ\psi-functions are preferred over OLS when high-leverage contamination is possible.

Question 10

Consider the variance functional V(F)=(xμF)2dF(x)V(F)=\int (x-\mu_F)^2\,dF(x). At the distribution of interest, μF=1\mu_F=1 and V(F)=4V(F)=4.

What is the influence function of VV at a contamination point z=4z=4?

  1. 33, obtained from the contamination point's deviation from the population mean.
  2. 55, obtained by centering the squared deviation by the population variance. (correct answer)
  3. 99, obtained from the squared deviation of the contamination point from the mean.
  4. 5-5, obtained by subtracting the squared deviation from the population variance.
Explanation: When you encounter influence function questions, your goal is to quantify how a point mass contamination at zz shifts a functional — essentially, the derivative of the functional along the direction of that contamination. For the variance functional V(F)=(xμF)2dF(x)V(F) = \int (x - \mu_F)^2 \, dF(x), the influence function is derived by considering V(Fϵ)=V((1ϵ)F+ϵδz)V(F_\epsilon) = V((1-\epsilon)F + \epsilon\delta_z) and differentiating with respect to ϵ\epsilon at ϵ=0\epsilon = 0. Working through this carefully (accounting for the fact that μF\mu_F also shifts with contamination), the influence function evaluates to (zμF)2V(F)(z - \mu_F)^2 - V(F). Plugging in z=4z = 4, μF=1\mu_F = 1, and V(F)=4V(F) = 4: (41)24=94=5(4-1)^2 - 4 = 9 - 4 = 5. This is choice B — and the description "centering the squared deviation by the population variance" captures precisely this structure. Choice A gives 33, which is just zμF=41z - \mu_F = 4 - 1, the raw deviation — not squared, not centered. This ignores almost all of the influence function's structure. Choice C gives 9=(zμF)29 = (z - \mu_F)^2, which is the squared deviation alone; this student forgot to subtract V(F)V(F), missing the centering step entirely. Choice D gives 5-5, which reverses the subtraction — subtracting the squared deviation from the variance rather than the variance from the squared deviation. A useful memory device: the influence function of variance looks like a "centered" squared deviation — (zμ)2σ2(z - \mu)^2 - \sigma^2 — analogous to how variance itself centers squared deviations around their mean. Always remember to subtract V(F)V(F).