Statistics Graduate Level Quiz: Common Pitfalls
10 questions · exam conditions
0:00
Common PitfallsQuestion 1 of 10

For 0<θ<10<\theta<1, define fθ(x)=θ11{0<x<θ}f_\theta(x)=\theta^{-1}\mathbf{1}\{0<x<\theta\} on (0,1)(0,1), and define f0(x)=0f_0(x)=0. For every fixed x>0x>0, fθ(x)f0(x)f_\theta(x)\to f_0(x) as θ0\theta\downarrow0.

Which statement correctly explains why one cannot conclude that 01fθ(x)dx01f0(x)dx\int_0^1 f_\theta(x)\,dx\to\int_0^1 f_0(x)\,dx?

Fatou's lemma is inapplicable because the functions are nonnegative but fail to converge almost everywhere.
The monotone convergence theorem applies, but it yields convergence to one rather than convergence to zero.
The bounded convergence theorem applies because each individual function is bounded on (0,1)(0,1).
Pointwise convergence is insufficient because no common integrable dominating function controls the moving spike.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Common Pitfalls

Practice Common Pitfalls in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Common Pitfalls, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

For 0<θ<10<\theta<1, define fθ(x)=θ11{0<x<θ}f_\theta(x)=\theta^{-1}\mathbf{1}\{0<x<\theta\} on (0,1)(0,1), and define f0(x)=0f_0(x)=0. For every fixed x>0x>0, fθ(x)f0(x)f_\theta(x)\to f_0(x) as θ0\theta\downarrow0.

Which statement correctly explains why one cannot conclude that 01fθ(x)dx01f0(x)dx\int_0^1 f_\theta(x)\,dx\to\int_0^1 f_0(x)\,dx?

  1. Fatou's lemma is inapplicable because the functions are nonnegative but fail to converge almost everywhere.
  2. The monotone convergence theorem applies, but it yields convergence to one rather than convergence to zero.
  3. The bounded convergence theorem applies because each individual function is bounded on (0,1)(0,1).
  4. Pointwise convergence is insufficient because no common integrable dominating function controls the moving spike. (correct answer)
Explanation: Whenever you encounter a question about exchanging limits and integrals, your first instinct should be to ask: which convergence theorem applies, and are its hypotheses satisfied? The classic theorems — Dominated Convergence, Monotone Convergence, Bounded Convergence — each require conditions beyond mere pointwise convergence. Here, each fθf_\theta is a "moving spike": a rectangle of height θ1\theta^{-1} and width θ\theta, so 01fθ(x)dx=1\int_0^1 f_\theta(x)\,dx = 1 for every θ>0\theta > 0. Yet pointwise, fθ(x)0f_\theta(x) \to 0 for every fixed x>0x > 0, since eventually x>θx > \theta. So fθ10=f0\int f_\theta \to 1 \neq 0 = \int f_0. The culprit is the lack of a dominating function. Any integrable gg that dominates this family would need g(x)θ1g(x) \geq \theta^{-1} for x(0,θ)x \in (0,\theta) for all small θ\theta, which forces gg to be non-integrable near zero. Without such a dominator, the Dominated Convergence Theorem fails, and you cannot pass the limit inside the integral. This is precisely what D captures. A is wrong because the functions do converge pointwise almost everywhere (everywhere on (0,1)(0,1) in fact) — the issue is not a failure of a.e. convergence. B is wrong because the Monotone Convergence Theorem requires a monotone increasing sequence of functions, and fθf_\theta is not monotone as θ0\theta \downarrow 0. C is wrong because the Bounded Convergence Theorem requires a uniform bound across all functions; here the spike height θ1\theta^{-1} \to \infty, so no such uniform bound exists. As a study rule: moving-spike examples are the canonical illustration that pointwise convergence alone never justifies swapping limits and integrals — always verify a dominating function exists.

Question 2

Let Sn=k=1nξkS_n=\sum_{k=1}^n \xi_k, where the ξk\xi_k are independent and satisfy P(ξk=1)=P(ξk=1)=1/2P(\xi_k=1)=P(\xi_k=-1)=1/2. Let τ=inf{n0:Sn=1}\tau=\inf\{n\geq0:S_n=1\}. The symmetric random walk is recurrent, so P(τ<)=1P(\tau<\infty)=1 and Sτ=1S_\tau=1 almost surely.

Why does applying optional stopping to conclude E(Sτ)=E(S0)=0E(S_\tau)=E(S_0)=0 produce an incorrect result?

  1. The process is not a martingale because its increments are discrete rather than continuously distributed.
  2. The stopping time is not measurable with respect to the natural filtration at any finite time.
  3. Almost-sure finiteness is insufficient; E(τ)=E(\tau)=\infty and the stopped martingale lacks the needed uniform integrability. (correct answer)
  4. Optional stopping applies only when the terminal value is random rather than equal to a fixed constant.
Explanation: Whenever you encounter a question about the Optional Stopping Theorem (OST), your first instinct should be to check the conditions under which it applies — not just whether a martingale exists. The OST guarantees E(Sτ)=E(S0)E(S_\tau) = E(S_0) only when additional regularity conditions hold beyond P(τ<)=1P(\tau < \infty) = 1. Here, SnS_n is indeed a martingale (mean-zero, symmetric increments), and τ\tau is almost surely finite. But the classical OST requires something stronger: either E(τ)<E(\tau) < \infty, or the stopped process {Snτ}\{S_{n \wedge \tau}\} is uniformly integrable. For the simple symmetric random walk hitting level 1, it is a classical result that E(τ)=E(\tau) = \infty. Without finite expected stopping time, the martingale can "drift" over an arbitrarily long horizon before stopping, and the stopped process fails to be uniformly integrable. Consequently, you cannot exchange the limit and expectation, and E(Sτ)=10=E(S0)E(S_\tau) = 1 \neq 0 = E(S_0). This makes C correct. A is wrong because the OST applies perfectly well to discrete distributions — the symmetric random walk is a textbook martingale regardless of its discrete increments. B is wrong because τ\tau is a standard stopping time and is measurable with respect to the natural filtration; the issue is integrability, not measurability. D is wrong and essentially fabricated — OST places no restriction on whether the terminal value is fixed or random. Your study tip: whenever OST is invoked, immediately ask three questions — Is it a martingale? Is E(τ)<E(\tau) < \infty? Is the stopped process uniformly integrable? Failing any one condition invalidates the conclusion.

Question 3

For independent and identically distributed observations (Yi,Xi)(Y_i,X_i), let β\beta^* minimize E[(YXβ)2]E[(Y-X^\top\beta)^2], and define u=YXβu=Y-X^\top\beta^*. Assume finite fourth moments and nonsingular Q=E(XX)Q=E(XX^\top), but do not assume that E(u2X)E(u^2\mid X) is constant. Let Ω=E(u2XX)\Omega=E(u^2XX^\top).

What is the asymptotic covariance matrix of n(β^β)\sqrt n(\widehat\beta-\beta^*) under these assumptions?

  1. Q1ΩQ1Q^{-1}\Omega Q^{-1}, because the projection condition alone does not imply homoskedasticity. (correct answer)
  2. E(u2)Q1E(u^2)Q^{-1}, because least squares residuals are orthogonal to the regressors.
  3. Q1Q^{-1}, because minimizing expected squared error normalizes the residual variance.
  4. Ω1QΩ1\Omega^{-1}Q\Omega^{-1}, because the score covariance must be inverted on both sides.
Explanation: When you see a question about OLS asymptotics without a homoskedasticity assumption, your immediate instinct should be: "sandwich estimator." The classical OLS variance formula only simplifies when E(u2X)E(u^2 \mid X) is constant — without that, you need the full robust (sandwich) form. Here's the core derivation. The OLS estimator satisfies n(β^β)=(1nXiXi)11nXiui\sqrt{n}(\hat{\beta} - \beta^*) = \left(\frac{1}{n}\sum X_i X_i^\top\right)^{-1} \frac{1}{\sqrt{n}}\sum X_i u_i. By the law of large numbers, 1nXiXipQ\frac{1}{n}\sum X_i X_i^\top \xrightarrow{p} Q. By the CLT, 1nXiuidN(0,Ω)\frac{1}{\sqrt{n}}\sum X_i u_i \xrightarrow{d} N(0, \Omega) where Ω=E(u2XX)\Omega = E(u^2 X X^\top). Applying the delta method (Slutsky), the asymptotic covariance is Q1ΩQ1Q^{-1}\Omega Q^{-1} — the sandwich form. This is answer A, and it's correct precisely because heteroskedasticity breaks the simplification that would otherwise collapse the sandwich. Answer B is wrong because E(u2)Q1E(u^2)Q^{-1} is only valid under homoskedasticity, where Ω=E(u2)Q\Omega = E(u^2)Q. Orthogonality of residuals and regressors is a first-order condition, not a variance result. Answer C is wrong because minimizing expected squared error determines where β\beta^* is, not the shape of the variance matrix — nothing "normalizes" residual variance to produce Q1Q^{-1} alone. Answer D inverts the roles of QQ and Ω\Omega, which is a structural error; the outer matrices must be Q1Q^{-1}, not Ω1\Omega^{-1}. Study tip: Memorize the sandwich as Q1ΩQ1Q^{-1}\Omega Q^{-1}. Anytime an exam omits homoskedasticity, that's your formula — the simpler forms are just special cases.

Question 4

Let X1,,XnX_1,\ldots,X_n be independent observations from Uniform(0,θ)\operatorname{Uniform}(0,\theta), with estimator Mn=maxiXiM_n=\max_i X_i. An ordinary nonparametric bootstrap sample is drawn with replacement from the observed data, and its maximum is denoted by MnM_n^*.

Which statement best diagnoses whether the ordinary bootstrap consistently estimates the limiting distribution of the scaled estimation error?

  1. It is consistent because both MnM_n and MnM_n^* converge to θ\theta at the same deterministic rate.
  2. It is consistent after studentization because bootstrap resampling removes the boundary of the parameter space.
  3. It is inconsistent only when the observed sample maximum occurs more than once in the original sample.
  4. It is inconsistent because the bootstrap error has an asymptotic atom at zero absent from the target limit. (correct answer)
Explanation: When a question asks about bootstrap consistency for an estimator near a boundary, your first instinct should be to examine the support of the bootstrap resample relative to the original parameter space — specifically, what the maximum of a resample can and cannot see. Here, the key fact is that the scaled error n(θMn)n(\theta - M_n) converges in distribution to an Exponential(1/θ)\text{Exponential}(1/\theta) random variable — a continuous limit with no point mass anywhere. Now consider what happens when you bootstrap. The bootstrap resample is drawn from the empirical distribution, so MnMnM_n^* \leq M_n always. This means n(MnMn)0n(M_n - M_n^*) \geq 0 always, and crucially, whenever any observation equals MnM_n is resampled, Mn=MnM_n^* = M_n exactly, giving a contribution at zero. Because each resample includes MnM_n with probability 1(11/n)n1e1>01 - (1 - 1/n)^n \to 1 - e^{-1} > 0, the bootstrap error distribution has a genuine atom at zero with positive probability in the limit. The true limiting distribution is continuous and has no such atom. This is exactly what D describes — the distributions differ structurally, so the bootstrap is inconsistent. A is wrong because convergence rates matching does not imply distributional consistency; the shapes of the limiting distributions must also match. B is wrong because studentization does not resolve a boundary problem — the issue is that θ\theta is the boundary of the support, and resampling cannot exceed MnM_n. C is wrong because the atom at zero arises generically from resampling, not only when ties exist in the original sample. The takeaway: whenever the MLE sits at a boundary of the parameter support (like a uniform maximum), bootstrap resamples are confined below that boundary, injecting a probability atom that the true limiting distribution lacks — a reliable signal of bootstrap inconsistency.

Question 5

A normal-data analysis compares model M0M_0, in which the mean is fixed at zero, with model M1M_1, in which the mean is unknown. The analyst uses the improper prior π0(σ)1/σ\pi_0(\sigma)\propto1/\sigma under M0M_0 and π1(μ,σ)1/σ\pi_1(\mu,\sigma)\propto1/\sigma under M1M_1. For the observed sample, both posterior distributions are proper.

Can the analyst use the two resulting marginal likelihoods to obtain a uniquely defined Bayes factor?

  1. Yes, because posterior propriety uniquely determines the missing normalizing constants in both priors.
  2. No, because arbitrary prior constants enter the two marginal likelihoods and need not cancel across models. (correct answer)
  3. Yes, because the common symbolic factor 1/σ1/\sigma cancels even though the parameter spaces differ.
  4. No, because an improper prior necessarily makes each posterior distribution improper as well.
Explanation: Whenever you see a question involving Bayes factors with improper priors across different models, your first instinct should be to ask: are the arbitrary normalizing constants truly shared, or do they float independently? The Bayes factor is defined as the ratio of marginal likelihoods, BF10=m1(x)/m0(x)BF_{10} = m_1(\mathbf{x})/m_0(\mathbf{x}), where each marginal likelihood integrates the likelihood against its prior. With improper priors, each prior is only defined up to an arbitrary positive constant: π0(σ)=c0/σ\pi_0(\sigma) = c_0/\sigma and π1(μ,σ)=c1/σ\pi_1(\mu,\sigma) = c_1/\sigma. These constants propagate directly into their respective marginal likelihoods, giving m0(x)c0(finite integral)m_0(\mathbf{x}) \propto c_0 \cdot (\text{finite integral}) and m1(x)c1(finite integral)m_1(\mathbf{x}) \propto c_1 \cdot (\text{finite integral}). The ratio BF10BF_{10} then contains the factor c1/c0c_1/c_0, which is completely arbitrary. Since M0M_0 and M1M_1 are separate models with separately specified priors, there is no reason these constants must agree or cancel. The Bayes factor is therefore not uniquely defined — confirming that B is correct. Choice A is wrong because posterior propriety tells you only that the integral is finite; it does not pin down the value of the normalizing constant for an improper prior — those are entirely separate issues. Choice C is the most tempting distractor: the symbolic form 1/σ1/\sigma looks the same in both priors, but the constants multiplying them are independently arbitrary and live in different parameter spaces, so nothing forces cancellation across models. Choice D states a false general rule; an improper prior can absolutely yield a proper posterior (as the passage explicitly confirms). A useful rule of thumb: improper priors are "safe" within a single model for posterior inference, but they are never safe for cross-model comparison via Bayes factors unless the constants are demonstrably tied together.

Question 6

Suppose n(θ^n0)N(0,σ2)\sqrt n(\widehat\theta_n-0)\Rightarrow N(0,\sigma^2) with σ2>0\sigma^2>0. The parameter of interest is g(θ)=θ2g(\theta)=\theta^2.

Which conclusion correctly accounts for the fact that g(0)=0g'(0)=0?

  1. ng(θ^n)N(0,4σ2)\sqrt n\,g(\widehat\theta_n)\Rightarrow N(0,4\sigma^2) by the ordinary delta method.
  2. ng(θ^n)N(0,σ4)n\,g(\widehat\theta_n)\Rightarrow N(0,\sigma^4) because the convergence rate is squared.
  3. ng(θ^n)σ2χ12n\,g(\widehat\theta_n)\Rightarrow\sigma^2\chi_1^2 by applying the continuous mapping theorem to the squared limit. (correct answer)
  4. No rescaling can produce a nondegenerate limit because the first derivative of gg vanishes.
Explanation: When the delta method's first derivative vanishes at the true parameter value, you need a second-order (quadratic) delta method instead. The ordinary delta method says n(g(θ^n)g(θ))N(0,[g(θ)]2σ2)\sqrt{n}(g(\hat\theta_n) - g(\theta)) \Rightarrow N(0, [g'(\theta)]^2\sigma^2), but when g(0)=0g'(0) = 0, this collapses to a degenerate point mass at zero — useless. Instead, expand to second order: g(θ^n)=g(0)+12g(0)θ^n2+g(\hat\theta_n) = g(0) + \frac{1}{2}g''(0)\hat\theta_n^2 + \cdots, giving ng(θ^n)12g(0)(nθ^n)2n\,g(\hat\theta_n) \approx \frac{1}{2}g''(0)\cdot(\sqrt{n}\,\hat\theta_n)^2. Here g(θ)=θ2g(\theta)=\theta^2, so g(0)=2g''(0)=2, and nθ^nN(0,σ2)\sqrt{n}\,\hat\theta_n \Rightarrow N(0,\sigma^2). By the continuous mapping theorem, ng(θ^n)σ2χ12n\,g(\hat\theta_n) \Rightarrow \sigma^2 \chi_1^2, since the square of a standard normal times σ2\sigma^2 is exactly σ2χ12\sigma^2\chi_1^2. That confirms C is correct. Choice A fails because it blindly applies the ordinary delta method with g(0)=20=0g'(0)=2\cdot 0=0, producing a degenerate limit — and even if you substituted the wrong derivative, 4σ24\sigma^2 is not the correct variance. Choice B gets the rescaling rate right (nn instead of n\sqrt{n}) but claims a normal limit, ignoring that squaring a normal yields a chi-squared distribution, not another normal. Choice D is the most tempting distractor — it correctly identifies that the first derivative vanishes, but wrongly concludes no nondegenerate limit exists; the second-order expansion rescues the situation. The key study pattern: whenever you see g(θ0)=0g'(\theta_0)=0, immediately switch to the second-order delta method, expect a rate of nn (not n\sqrt{n}), and look for a chi-squared limit rather than a normal one.

Question 7

Let Z1Z_1 and Z2Z_2 be independent standard normal random variables. For each nn, define Xn=Z1/nX_n=Z_1/\sqrt n and Yn=Z2/nY_n=Z_2/\sqrt n.

Although both XnX_n and YnY_n converge in probability to zero, what is the correct conclusion about Xn/YnX_n/Y_n?

  1. Its distribution is standard Cauchy for every nn, so it does not converge in probability to a constant. (correct answer)
  2. It converges in probability to one because numerator and denominator have the same limiting distribution.
  3. It converges in probability to zero because the numerator converges in probability to zero.
  4. It diverges in probability because the denominator converges in probability to zero at rate n1/2n^{-1/2}.
Explanation: When a question involves ratios of random variables that are each converging to zero, you cannot simply apply limit rules as if they were deterministic sequences. The key question is: what does the ratio's distribution look like at every finite nn? Notice that Xn/Yn=(Z1/n)/(Z2/n)=Z1/Z2X_n/Y_n = (Z_1/\sqrt{n})/(Z_2/\sqrt{n}) = Z_1/Z_2. The n\sqrt{n} terms cancel exactly, leaving the ratio of two independent standard normals — which is the definition of a standard Cauchy distribution. This holds for every nn, not just in the limit. Since the distribution of Xn/YnX_n/Y_n is exactly Cauchy for all nn, it cannot converge in probability to any constant. A Cauchy random variable has such heavy tails that it doesn't even have a finite mean, and its spread never collapses. Answer A is correct. Answer B is wrong because sharing the same limiting distribution (both converge to 0) tells you nothing about the behavior of their ratio. Convergence in probability is not preserved under division without additional conditions. Answer C commits an even more tempting error: you cannot conclude Xn/Yn0X_n/Y_n \to 0 just because the numerator Xn0X_n \to 0, because the denominator Yn0Y_n \to 0 at the same rate, and the rates cancel. Answer D is also incorrect — the ratio doesn't diverge; it stays Cauchy-distributed at every nn. As a study rule: when you see a ratio where both numerator and denominator shrink at the same rate, always check whether they cancel before concluding anything about convergence. Rates matter more than limits alone.

Question 8

In a completely randomized experiment, exactly mm of NN units are assigned to treatment. Let WiW_i equal 1 if unit ii is treated and 0 otherwise, and define p=m/Np=m/N.

For two distinct units ii and jj, which expression gives Cov(Wi,Wj)\operatorname{Cov}(W_i,W_j) under the randomization distribution?

  1. 00, because both assignment indicators have the same marginal treatment probability.
  2. p(1p)N1-\dfrac{p(1-p)}{N-1}, because conditioning on exactly mm treatments induces negative dependence. (correct answer)
  3. p(1p)-p(1-p), because treatment of one unit excludes treatment of every other unit.
  4. p(1p)N1\dfrac{p(1-p)}{N-1}, because fixing the treated sample size induces positive dependence.
Explanation: Whenever you encounter randomization distributions in experimental design, the key insight is that assignment indicators are not independent — they are constrained by the fixed total iWi=m\sum_i W_i = m. This constraint is what drives the covariance calculation. To derive Cov(Wi,Wj)\operatorname{Cov}(W_i, W_j), use the identity Var ⁣(iWi)=iVar(Wi)+ijCov(Wi,Wj)\operatorname{Var}\!\left(\sum_i W_i\right) = \sum_i \operatorname{Var}(W_i) + \sum_{i \neq j} \operatorname{Cov}(W_i, W_j). Since iWi=m\sum_i W_i = m is fixed, its variance is exactly 0. Each WiW_i is Bernoulli with mean pp, so Var(Wi)=p(1p)\operatorname{Var}(W_i) = p(1-p). Substituting: 0=Np(1p)+N(N1)Cov(Wi,Wj)0 = N \cdot p(1-p) + N(N-1)\operatorname{Cov}(W_i, W_j). Solving gives Cov(Wi,Wj)=p(1p)N1\operatorname{Cov}(W_i, W_j) = -\dfrac{p(1-p)}{N-1}, confirming answer B. The negative sign is intuitive: knowing unit ii is treated slightly reduces the probability that unit jj is also treated, since only mm slots are available. A is wrong because it confuses identical marginal distributions with independence. Equal marginal probabilities do not imply zero covariance when a global constraint links the variables. C overstates the dependence. The expression p(1p)-p(1-p) would require each treated unit to exclude exactly one other, but treatment of unit ii reduces the probability for all remaining N1N-1 units, spread evenly — hence the N1N-1 denominator matters. D gets the sign backwards. The fixed-total constraint forces negative, not positive, dependence between any two indicators. As a study tip: whenever assignments are made under a fixed-total constraint, always expect negative pairwise covariance and use the zero-variance trick to compute it exactly.

Question 9

Suppose X1,,XnX_1,\ldots,X_n are independent observations from Uniform(0,θ)\operatorname{Uniform}(0,\theta). At the true parameter value, differentiating the log-likelihood with respect to θ\theta away from the boundary gives the score n/θ-n/\theta almost surely.

Why does the usual identity Eθ[logL(θ)/θ]=0E_\theta[\partial\log L(\theta)/\partial\theta]=0 fail in this model?

  1. The sample maximum is sufficient, so the full-sample score cannot have expectation zero.
  2. The support depends on θ\theta, and differentiating the likelihood integral introduces a boundary contribution. (correct answer)
  3. The score has infinite variance, so its finite nonzero expectation is not constrained by regular likelihood theory.
  4. The maximum likelihood estimator lies on a boundary, which makes every likelihood derivative undefined almost surely.
Explanation: Whenever you encounter a question about regularity conditions for maximum likelihood theory, your first instinct should be to check whether the support of the distribution depends on the parameter. This is the central issue here. The standard identity Eθ ⁣[logLθ]=0E_\theta\!\left[\frac{\partial \log L}{\partial \theta}\right] = 0 is derived by differentiating under the integral sign: θf(x;θ)dx=θf(x;θ)dx=0.\frac{\partial}{\partial \theta}\int f(x;\theta)\,dx = \int \frac{\partial}{\partial \theta} f(x;\theta)\,dx = 0. This interchange is only valid when the limits of integration are fixed (i.e., do not depend on θ\theta). For Uniform(0,θ)\text{Uniform}(0,\theta), the density is f(x;θ)=θ110xθf(x;\theta) = \theta^{-1}\mathbf{1}_{0 \le x \le \theta}, and the upper integration limit is θ\theta itself. Differentiating the integral properly requires a boundary term via Leibniz's rule, which is nonzero. That boundary contribution is precisely why the score n/θ-n/\theta has a nonzero expectation — the regularity condition breaks down, not the score itself. This confirms B is correct. Choice A is a red herring. Sufficiency of the maximum is true, but it has nothing to do with whether the score's expectation equals zero — these are unrelated properties. Choice C confuses the issue: the score here actually has finite variance, and infinite variance would be a separate regularity failure. The problem is not about variance at all. Choice D is partly true but misidentifies the consequence. The MLE does lie at a boundary, but likelihood derivatives are still well-defined almost everywhere — the issue is whether the derivative-expectation interchange is valid. Study tip: Whenever you see a uniform or truncated distribution with a parameter-dependent bound, immediately flag it as violating the Leibniz interchange condition — this single check resolves most regularity-condition questions on graduate exams.

Question 10

Let UUniform(0,1)U\sim\operatorname{Uniform}(0,1), and define Xn=n1{U1/n}X_n=n\mathbf{1}\{U\leq 1/n\} for every positive integer nn. All variables are defined on the same probability space.

Which statement about the convergence of XnX_n is correct?

  1. Xn0X_n\to 0 in probability but not almost surely, and EXn0E|X_n|\to 0.
  2. Xn0X_n\to 0 in L1L^1 but not almost surely, while E(Xn)=1E(X_n)=1.
  3. Xn0X_n\to 0 almost surely and in distribution, but not in L1L^1. (correct answer)
  4. Xn0X_n\to 0 almost surely and in L1L^1, so its expectations converge to zero.
Explanation: When you see a question mixing almost sure convergence, L1L^1 convergence, and convergence in probability, your first move should be to analyze each mode independently — they don't always agree, and this sequence is a classic counterexample designed to expose exactly those gaps. Start with almost sure convergence. Fix any ω(0,1)\omega \in (0,1). For large enough nn, we have U(ω)>1/nU(\omega) > 1/n, so Xn(ω)=0X_n(\omega) = 0 for all sufficiently large nn. This holds for every ω(0,1)\omega \in (0,1) (a set of probability 1), so Xn0X_n \to 0 almost surely. Since almost sure convergence implies convergence in distribution, Xn0X_n \to 0 in distribution as well. Now check L1L^1 convergence. Compute E[Xn]=nP(U1/n)=n(1/n)=1E[X_n] = n \cdot P(U \leq 1/n) = n \cdot (1/n) = 1 for every nn. Since the expectations never approach 0, EXn0=1↛0E|X_n - 0| = 1 \not\to 0, so XnX_n does not converge to 0 in L1L^1. This makes C the correct answer — almost sure convergence does not imply L1L^1 convergence without uniform integrability, and this sequence fails that condition. A is wrong because Xn0X_n \to 0 does hold almost surely (stronger than in probability), and EXn=1↛0E|X_n| = 1 \not\to 0. B claims L1L^1 convergence, which fails since EXn=1E|X_n| = 1 stays constant. D incorrectly asserts L1L^1 convergence; almost sure convergence alone does not force E[Xn]0E[X_n] \to 0. Study tip: Always check whether a sequence is uniformly integrable before concluding that almost sure convergence implies L1L^1 convergence — the two modes are independent without it.