Statistics Graduate Level Quiz: Convergence Mode Relationships
10 questions · exam conditions
0:00
Convergence Mode RelationshipsQuestion 1 of 10

Let X,X1,X2,X,X_1,X_2,\ldots be mutually independent random variables, each having the standard normal distribution.

Which statement correctly illustrates the distinction between convergence in distribution and convergence in probability?

XnpXX_n\xrightarrow{p}X because XnX_n and XX have identical distributions and equal expectations.
XndXX_n\xrightarrow{d}X and E[Xn]E[X]E[X_n]\to E[X], but XnX_n does not converge to XX in probability.
XnL1XX_n\xrightarrow{L^1}X because the common normal distribution makes the sequence uniformly integrable.
XnX_n fails to converge in distribution to XX because independence prevents pathwise agreement.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Convergence Mode Relationships

Practice Convergence Mode Relationships in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Convergence Mode Relationships, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Let X,X1,X2,X,X_1,X_2,\ldots be mutually independent random variables, each having the standard normal distribution.

Which statement correctly illustrates the distinction between convergence in distribution and convergence in probability?

  1. XnpXX_n\xrightarrow{p}X because XnX_n and XX have identical distributions and equal expectations.
  2. XndXX_n\xrightarrow{d}X and E[Xn]E[X]E[X_n]\to E[X], but XnX_n does not converge to XX in probability. (correct answer)
  3. XnL1XX_n\xrightarrow{L^1}X because the common normal distribution makes the sequence uniformly integrable.
  4. XnX_n fails to converge in distribution to XX because independence prevents pathwise agreement.
Explanation: When comparing convergence modes, the key question is always: what exactly must be small? Convergence in distribution only requires that CDFs agree in the limit — it says nothing about whether XnX_n and XX are close as random variables on the same probability space. Convergence in probability requires P(XnX>ε)0P(|X_n - X| > \varepsilon) \to 0, which is a statement about their joint behavior. Here, every XnN(0,1)X_n \sim N(0,1) and XN(0,1)X \sim N(0,1), so trivially XndXX_n \xrightarrow{d} X — the CDFs are identical at every step, not just in the limit. Expectations also agree: E[Xn]=E[X]=0E[X_n] = E[X] = 0. But because XnX_n and XX are mutually independent, XnXN(0,2)X_n - X \sim N(0,2) for every nn, so P(XnX>ε)=P(N(0,2)>ε)P(|X_n - X| > \varepsilon) = P(|N(0,2)| > \varepsilon), which is a fixed positive constant — never tending to zero. Thus Xn̸pXX_n \not\xrightarrow{p} X, making B the correct statement. A is wrong because identical distributions and equal expectations are insufficient for convergence in probability; you need the joint distribution of (Xn,X)(X_n, X) to concentrate near zero difference, which independence destroys. C is wrong because L1L^1 convergence implies convergence in probability, and we just showed that fails here. Uniform integrability is a red herring — it helps with moment convergence, not pathwise closeness. D is wrong because independence has no bearing on distributional convergence; the marginal CDFs still match perfectly. Study tip: Whenever XnX_n and XX are i.i.d., you always get XndXX_n \xrightarrow{d} X for free, but independence makes XnXX_n - X spread out, killing convergence in probability. This is the canonical counterexample to the converse of "p\xrightarrow{p} implies d\xrightarrow{d}."

Question 2

For random variables XnX_n and XX on a common probability space, suppose that for every ε>0\varepsilon>0 and every positive integer nn, P(XnX>ε)1/(n2ε2)P(|X_n-X|>\varepsilon)\leq 1/(n^2\varepsilon^2).

Which convergence conclusion follows from this bound without any independence assumption?

  1. XnpXX_n\xrightarrow{p}X, but almost-sure convergence requires independence of the events {XnX>ε}\{|X_n-X|>\varepsilon\} and cannot be deduced from this bound alone.
  2. XnL1XX_n\xrightarrow{L^1}X, but the bound is insufficient to establish almost-sure convergence.
  3. XnL2XX_n\xrightarrow{L^2}X, because the ε2\varepsilon^{-2} factor in the probability bound implies control over second moments.
  4. Xna.s.XX_n\xrightarrow{a.s.}X, and hence convergence in probability and in distribution also hold. (correct answer)
Explanation: When a question gives you an explicit probability bound, your first instinct should be to check whether the Borel-Cantelli lemma applies — this is the bridge between probabilistic bounds and almost-sure convergence. Here, you're given P(XnX>ε)1n2ε2P(|X_n - X| > \varepsilon) \leq \frac{1}{n^2 \varepsilon^2}. For any fixed ε>0\varepsilon > 0, sum these bounds over all nn: n=1P(XnX>ε)n=11n2ε2=π26ε2<.\sum_{n=1}^{\infty} P(|X_n - X| > \varepsilon) \leq \sum_{n=1}^{\infty} \frac{1}{n^2 \varepsilon^2} = \frac{\pi^2}{6\varepsilon^2} < \infty. Since the sum of probabilities is finite, the first Borel-Cantelli lemma guarantees — with no independence assumption — that P(XnX>ε i.o.)=0P(|X_n - X| > \varepsilon \text{ i.o.}) = 0. Because this holds for every \varepsilon > 0$, you conclude X_n \xrightarrow{a.s.} X$$, making D correct. Almost-sure convergence then implies convergence in probability and in distribution for free. A is wrong precisely because it invokes independence as a requirement. Borel-Cantelli's first lemma requires only summability of the probabilities — independence is needed only for the second lemma (the converse direction). B is wrong because the bound controls tail probabilities, not L1L^1 norms; you cannot extract moment information from this alone. C is similarly wrong — the ε2\varepsilon^{-2} factor comes from a Chebyshev-style setup but does not imply L2L^2 convergence; that would require bounding E[XnX2]E[|X_n - X|^2] directly. A reliable strategy: whenever tail probabilities decay like n2n^{-2} or faster, immediately check summability and invoke Borel-Cantelli. This pattern appears frequently on graduate probability exams.

Question 3

A sequence of estimators is known to be Cauchy in probability: for every ε>0\varepsilon>0, P(TnTm>ε)0P(|T_n-T_m|>\varepsilon)\to 0 as m,nm,n\to\infty.

Which conclusion is valid on the underlying probability space?

  1. There exists a finite random variable TT such that Tna.s.TT_n\xrightarrow{a.s.}T, with no additional assumptions.
  2. A probability limit exists only if {Tn}\{T_n\} is additionally uniformly integrable or almost surely bounded.
  3. There exists a distributional limit for TnT_n, but no random variable on the same space need be a probability limit.
  4. There exists a finite random variable TT such that TnpTT_n\xrightarrow{p}T, and the limit is unique almost surely. (correct answer)
Explanation: Whenever you see a question about Cauchy sequences in probability, think of the analogy to metric space completeness — the key question is whether the probability space is "complete enough" to guarantee a limit exists. The space of random variables under convergence in probability is indeed complete in this sense. Formally, if {Tn}\{T_n\} is Cauchy in probability, you can extract a subsequence TnkT_{n_k} that converges almost surely to some measurable random variable TT. From that almost-sure convergence along the subsequence, you can show the full sequence satisfies TnpTT_n \xrightarrow{p} T. Moreover, probability limits are unique almost surely: if TnpTT_n \xrightarrow{p} T and TnpTT_n \xrightarrow{p} T', then P(TT)=0P(T \neq T') = 0. This confirms D is correct. A is wrong because the Cauchy-in-probability condition does not guarantee almost-sure convergence of the full sequence — only convergence in probability. Almost-sure convergence is strictly stronger and requires additional structure or arguments beyond the Cauchy criterion alone. B is wrong because it imposes unnecessary conditions. Uniform integrability and almost-sure boundedness are relevant for L1L^1 or LL^\infty convergence, not for the existence of a probability limit. No such extra assumption is needed here. C is wrong in a subtle but important way: it confuses weak convergence (distributional limits) with convergence in probability. A Cauchy-in-probability sequence guarantees a limit on the same probability space, not merely a distributional limit. As a study tip: remember that "Cauchy in probability \Rightarrow convergence in probability" is the probabilistic analogue of completeness in R\mathbb{R}, and always distinguish between almost-sure, in-probability, and distributional convergence — they are a favorite source of traps on graduate exams.

Question 4

An estimator and an auxiliary statistic are computed from the same data, so no independence assumption is available.

Suppose E[(Tnθ)2]0E[(T_n-\theta)^2]\to 0 and ZndN(0,1)Z_n\xrightarrow{d}N(0,1). Let Wn=(Tnθ)ZnW_n=(T_n-\theta)Z_n. Which conclusion is guaranteed?

  1. Wna.s.0W_n\xrightarrow{a.s.}0, because mean-square convergence can be combined with weak convergence.
  2. Wnp0W_n\xrightarrow{p}0, although convergence in mean or almost surely need not hold. (correct answer)
  3. WnL10W_n\xrightarrow{L^1}0, because convergence in distribution makes the sequence {Zn}\{Z_n\} uniformly integrable.
  4. WnL20W_n\xrightarrow{L^2}0, even without independence or higher-moment assumptions on ZnZ_n.
Explanation: When you see a product of two random sequences—one converging to zero, one remaining stochastically bounded—your first instinct should be to reach for convergence in probability tools, specifically Slutsky-type reasoning. Here, E[(Tnθ)2]0E[(T_n - \theta)^2] \to 0 means TnpθT_n \xrightarrow{p} \theta, so (Tnθ)p0(T_n - \theta) \xrightarrow{p} 0. Since ZndN(0,1)Z_n \xrightarrow{d} N(0,1), the sequence {Zn}\{Z_n\} is bounded in probability (stochastically bounded, i.e., Op(1)O_p(1)): for any ε>0\varepsilon > 0, you can find MM such that P(Zn>M)<εP(|Z_n| > M) < \varepsilon for all large nn. A sequence converging in probability to zero, multiplied by an Op(1)O_p(1) sequence, converges in probability to zero. Thus Wn=(Tnθ)Znp0W_n = (T_n - \theta)Z_n \xrightarrow{p} 0, confirming B. A is wrong because mean-square convergence of TnT_n to θ\theta does not imply almost-sure convergence of WnW_n. Almost-sure convergence requires path-by-path control that neither assumption provides, and the joint behavior without independence is uncontrolled. C is wrong because ZndN(0,1)Z_n \xrightarrow{d} N(0,1) does not imply uniform integrability of {Zn}\{Z_n\}. Uniform integrability requires control over tail expectations, which weak convergence alone cannot guarantee. D is wrong because L2L^2 convergence of WnW_n would require E[Wn2]=E[(Tnθ)2Zn2]0E[W_n^2] = E[(T_n-\theta)^2 Z_n^2] \to 0, which needs E[Zn2]E[Z_n^2] to be uniformly bounded—an assumption not given here. Study tip: Memorize that op(1)Op(1)=op(1)o_p(1) \cdot O_p(1) = o_p(1). This "little-oh times big-oh" rule is your go-to for products of degenerating and stochastically bounded sequences, and it appears constantly in asymptotic statistics.

Question 5

Suppose AnpAA_n\xrightarrow{p}A and BndBB_n\xrightarrow{d}B. In general, these marginal statements do not imply An+BndA+BA_n+B_n\xrightarrow{d}A+B. Which additional condition is sufficient to make the displayed conclusion valid?

  1. The random variables AA and BB both have finite variances, with no restriction on their joint dependence.
  2. The sequence {An}\{A_n\} is bounded in probability, while BB has a continuous distribution.
  3. The limit AA is almost surely equal to a fixed constant, so Slutsky's theorem applies. (correct answer)
  4. The expectations of AnA_n and BnB_n converge separately to those of AA and BB.
Explanation: When you see a question about combining convergence in probability with convergence in distribution, your first instinct should be to recall Slutsky's theorem: if XndXX_n \xrightarrow{d} X and YnpcY_n \xrightarrow{p} c for some constant cc, then Xn+YndX+cX_n + Y_n \xrightarrow{d} X + c. The key ingredient is that one limit must be a degenerate constant, not a random variable. This is precisely what makes C correct. If AcA \equiv c almost surely, then AnpcA_n \xrightarrow{p} c, and Slutsky's theorem directly gives An+Bndc+B=A+BA_n + B_n \xrightarrow{d} c + B = A + B. The constant collapses the joint distribution problem — you no longer need to worry about the dependence structure between AnA_n and BnB_n. A is wrong because finite variances say nothing about joint behavior. Convergence in distribution is a statement about marginal laws, not moments, and knowing Var(A)\text{Var}(A) and Var(B)\text{Var}(B) are finite tells you nothing about how An+BnA_n + B_n behaves jointly. B is a tempting distractor: boundedness in probability (tightness) controls tail behavior but does not pin down the limit of An+BnA_n + B_n without specifying what AnA_n actually converges to. Continuity of BB's distribution is irrelevant here. D is wrong because convergence of expectations (first moments) is far weaker than convergence in distribution — you can construct counterexamples easily. As a study tip: whenever you see a mix of p\xrightarrow{p} and d\xrightarrow{d}, immediately ask yourself "Is one limit a constant?" — that's the Slutsky trigger.

Question 6

Let A1,A2,A_1,A_2,\ldots be independent events satisfying P(An)=1/nP(A_n)=1/n, and define Xn=1AnX_n=\mathbf{1}_{A_n}.

Which description of the convergence of XnX_n to zero is correct?

  1. Xn0X_n\to 0 in probability and in every LrL^r with r>0r>0, but not almost surely. (correct answer)
  2. Xn0X_n\to 0 almost surely and in probability, but not in any LrL^r with r1r\geq 1.
  3. Xn0X_n\to 0 only in distribution, because the divergent series of event probabilities prevents probability convergence.
  4. Xn0X_n\to 0 in probability and almost surely, but only in LrL^r when 0<r<10<r<1.
Explanation: When analyzing convergence of indicator random variables, you need to check four modes simultaneously: almost sure (a.s.), in probability, in LrL^r, and in distribution. Each has a distinct criterion, and they don't always align. Almost sure convergence requires asking whether Xn(ω)0X_n(\omega) \to 0 for almost every sample point. By the second Borel-Cantelli lemma, since the AnA_n are independent and P(An)=1/n=\sum P(A_n) = \sum 1/n = \infty, infinitely many AnA_n occur with probability 1. This means Xn=1X_n = 1 infinitely often almost surely, so Xn↛0X_n \not\to 0 a.s. In probability, you need P(Xn>ϵ)0P(|X_n| > \epsilon) \to 0. Since P(Xn=1)=1/n0P(X_n = 1) = 1/n \to 0, convergence in probability holds trivially. In LrL^r, note that Xnrr=E[Xnr]=P(An)=1/n0\|X_n\|_r^r = E[|X_n|^r] = P(A_n) = 1/n \to 0 for every r>0r > 0. So Xn0X_n \to 0 in LrL^r for all positive rr. This confirms A is correct. Choice B is wrong because XnX_n does not converge a.s. (Borel-Cantelli rules this out), and it does converge in every LrL^r. Choice C confuses Borel-Cantelli's conclusion: divergence of P(An)\sum P(A_n) kills a.s. convergence, not convergence in probability. Choice D incorrectly restricts LrL^r convergence to r<1r < 1, ignoring that E[Xnr]=1/n0E[X_n^r] = 1/n \to 0 for all r>0r > 0. Study tip: Memorize both Borel-Cantelli lemmas as a pair — the first (convergent series) gives a.s. convergence to zero; the second (divergent series + independence) gives a.s. non-convergence. They're a frequent source of traps on graduate probability exams.

Question 7

A sequence of simulation-based estimators is defined on a common probability space.

Suppose XnpXX_n\xrightarrow{p}X and the family {(XnX)2:n1}\{(X_n-X)^2:n\geq 1\} is uniformly integrable. Which conclusion is strongest among those listed?

  1. XndXX_n\xrightarrow{d}X, but neither first- nor second-mean convergence can be concluded.
  2. XnL1XX_n\xrightarrow{L^1}X, but convergence in L2L^2 is not ensured by uniform integrability.
  3. XnL2XX_n\xrightarrow{L^2}X, although almost-sure convergence of the full sequence need not hold. (correct answer)
  4. Xna.s.XX_n\xrightarrow{a.s.}X and XnL2XX_n\xrightarrow{L^2}X, because uniform integrability strengthens probability convergence.
Explanation: When you see uniform integrability (UI) paired with convergence in probability, your instinct should be to connect them through the Vitali Convergence Theorem: if XnpXX_n \xrightarrow{p} X and {Yn}\{Y_n\} is UI, then E[Yn]E[Y]E[Y_n] \to E[Y]. Here, the UI family is {(XnX)2}\{(X_n - X)^2\}, which converges in probability to (XX)2=0(X - X)^2 = 0. By Vitali, E[(XnX)2]0E[(X_n - X)^2] \to 0, which is precisely the definition of XnL2XX_n \xrightarrow{L^2} X. Since L2L^2 convergence implies L1L^1 convergence (by Jensen's or Cauchy-Schwarz), and L1L^1 implies convergence in distribution, answer C is the strongest valid conclusion. Answer A is too weak — you can conclude far more than mere convergence in distribution when UI of the squared differences is given. A undersells what UI buys you. Answer B correctly identifies L1L^1 convergence but claims L2L^2 is unattainable. This is the key trap: UI of (XnX)2(X_n - X)^2 specifically licenses L2L^2 convergence, not just L1L^1. B mistakes which family is UI. Answer D overclaims. Almost-sure convergence is a strictly stronger mode than convergence in probability, and uniform integrability provides no mechanism to upgrade p\xrightarrow{p} to a.s.\xrightarrow{a.s.} — counterexamples like the "sliding bump" sequence confirm this. Study tip: Always track which family is declared UI. UI of XnXr|X_n - X|^r paired with p\xrightarrow{p} gives LrL^r convergence via Vitali — match the exponent to the moment.

Question 8

Suppose that every subsequence {Xnk}\{X_{n_k}\} contains a further subsequence {Xnkj}\{X_{n_{k_j}}\} such that Xnkja.s.XX_{n_{k_j}}\xrightarrow{a.s.}X. What is the strongest conclusion that is necessarily valid?

  1. The original sequence converges to XX almost surely, since every subsequence has an almost-surely convergent refinement.
  2. The original sequence converges to XX in probability, but almost-sure or L1L^1 convergence need not follow. (correct answer)
  3. The original sequence converges to XX in distribution only, because the exceptional null sets may depend on the subsequence.
  4. The original sequence is Cauchy in L1L^1, although it need not converge to XX in probability.
Explanation: When you see a condition like "every subsequence contains a further almost-surely convergent subsequence," you should immediately recognize this as the subsequence characterization of convergence in probability. This is a fundamental theorem in probability theory worth knowing cold. The key result is: XnpXX_n \xrightarrow{p} X if and only if every subsequence {Xnk}\{X_{n_k}\} contains a further subsequence {Xnkj}\{X_{n_{k_j}}\} such that Xnkja.s.XX_{n_{k_j}} \xrightarrow{a.s.} X. The condition in the problem is precisely this characterization, so the strongest valid conclusion is convergence in probability — making B correct. Why can't you conclude almost-sure convergence (choice A)? Because the null sets on which convergence fails may differ across subsequences. You cannot take a union over uncountably many subsequences and preserve measure zero. A classic counterexample is the "typewriter sequence" on [0,1][0,1]: it converges in probability to 0 but fails to converge almost surely at any point. Choice C is too weak. Convergence in probability is strictly stronger than convergence in distribution, so stopping at distributional convergence undersells what the hypothesis actually gives you. Choice D is off-target entirely. The hypothesis says nothing about L1L^1 structure; you need uniform integrability to connect almost-sure convergence to L1L^1 convergence. Being Cauchy in L1L^1 is neither implied nor relevant here. Study tip: Memorize the subsequence criterion as a two-way bridge — it's both a way to prove convergence in probability (find the sub-subsequence) and a way to disprove it (exhibit a subsequence with no such refinement).

Question 9

Let UU be uniformly distributed on (0,1)(0,1), and define Xn=n1{U1/n}X_n=n\mathbf{1}_{\{U\leq 1/n\}}.

Which statement correctly characterizes the convergence of XnX_n to zero?

  1. Xn0X_n\to 0 almost surely and in probability, but not in L1L^1. (correct answer)
  2. Xn0X_n\to 0 in L1L^1 and in probability, but not almost surely.
  3. Xn0X_n\to 0 almost surely and in L1L^1, but not in probability.
  4. Xn0X_n\to 0 in probability but not in distribution, because its expectation remains equal to one for all nn.
Explanation: Whenever you see a question about modes of convergence, your instinct should be to check each mode independently — they don't imply one another in general, and this sequence is a classic illustration of why. Start with almost sure convergence. Fix a realization u(0,1)u \in (0,1). For large enough nn, specifically once n>1/un > 1/u, we have u>1/nu > 1/n, so Xn(u)=0X_n(u) = 0. Since every u(0,1)u \in (0,1) eventually satisfies this, Xn0X_n \to 0 almost surely. So A's claim of a.s. convergence is correct. Now check L1L^1 convergence. Compute E[Xn]=nP(U1/n)=n(1/n)=1E[X_n] = n \cdot P(U \leq 1/n) = n \cdot (1/n) = 1 for all nn. Since E[Xn0]=1↛0E[|X_n - 0|] = 1 \not\to 0, there is no L1L^1 convergence. This confirms A's claim that L1L^1 convergence fails. For convergence in probability: P(Xn>ε)=P(U1/n)=1/n0P(|X_n| > \varepsilon) = P(U \leq 1/n) = 1/n \to 0 for any ε<n\varepsilon < n. So Xn0X_n \to 0 in probability — consistent with A. Now address the wrong answers. B is wrong because it claims L1L^1 convergence, which fails as shown above. C is doubly wrong: it denies convergence in probability (which does hold) and falsely grants L1L^1 convergence. D is wrong on two counts: XnX_n does converge in probability, and the reasoning about expectation is a non-sequitur — having constant expectation doesn't prevent convergence in distribution. The takeaway: a.s. convergence does not imply L1L^1 convergence, and vice versa. This sequence — where probability mass collapses onto a shrinking set but with compensating magnitude — is the canonical counterexample to memorize.

Question 10

A sequence of random variables satisfies XndcX_n\xrightarrow{d}c, where cc is a finite constant. Which statement most accurately describes what follows without additional assumptions?

  1. Xna.s.cX_n\xrightarrow{a.s.}c, but convergence in L1L^1 may fail when the sequence is not uniformly integrable.
  2. XnpcX_n\xrightarrow{p}c, but almost-sure and L1L^1 convergence are not guaranteed. (correct answer)
  3. XnL1cX_n\xrightarrow{L^1}c, but convergence in probability can fail because the variables may be dependent.
  4. Only XndcX_n\xrightarrow{d}c follows; convergence in probability requires all variables to share one distribution.
Explanation: When you see convergence in distribution to a constant, that's a special case worth memorizing because it's strictly stronger than typical distributional convergence. The key theorem: if XndcX_n \xrightarrow{d} c where cc is a constant (a degenerate distribution), then XnpcX_n \xrightarrow{p} c. Here's the intuition — for any ϵ>0\epsilon > 0, the CDF of the limit is F(x)=1xcF(x) = \mathbf{1}_{x \geq c}, so P(Xnc>ϵ)=P(Xn>c+ϵ)+P(Xncϵ)0+0=0P(|X_n - c| > \epsilon) = P(X_n > c+\epsilon) + P(X_n \leq c-\epsilon) \to 0 + 0 = 0 by pointwise convergence of CDFs at continuity points. This makes B correct: convergence in probability is guaranteed, but almost-sure and L1L^1 convergence require additional structure. A is wrong because convergence in distribution — even to a constant — does not imply almost-sure convergence. You can construct sequences that converge in probability (and hence in distribution) but fail to converge almost surely (e.g., the "typewriter sequence"). C is wrong on two counts: L1L^1 convergence is not guaranteed without uniform integrability, and convergence in probability is not blocked by dependence — the argument above makes no independence assumption. D is wrong because it invents a false requirement. Convergence in probability from XndcX_n \xrightarrow{d} c holds regardless of whether the variables share a common distribution. Study tip: Memorize this hierarchy exception — convergence in distribution to a constant upgrades to convergence in probability, but stops there. L1L^1 and a.s. convergence still need uniform integrability or monotonicity arguments, respectively.