Statistics Graduate Level Quiz: Convergence In Probability And Almost Sure
10 questions · exam conditions
0:00
Convergence In Probability And Almost SureQuestion 1 of 10

Suppose XnXX_n\to X in probability. For each nn, let NnN_n be a positive integer-valued random variable that is independent of the entire collection {X,X1,X2,}\{X,X_1,X_2,\ldots\}, and assume NnN_n\to\infty in probability. What conclusion is necessarily valid?

No convergence is guaranteed unless Nn/n1N_n/n\to 1 in probability.
XNnXX_{N_n}\to X almost surely because the random indices diverge.
XNnXX_{N_n}\to X only in distribution because random indexing destroys probability convergence.
XNnXX_{N_n}\to X in probability, though almost sure convergence need not follow.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Convergence In Probability And Almost Sure

Practice Convergence In Probability And Almost Sure in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Convergence In Probability And Almost Sure, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Suppose XnXX_n\to X in probability. For each nn, let NnN_n be a positive integer-valued random variable that is independent of the entire collection {X,X1,X2,}\{X,X_1,X_2,\ldots\}, and assume NnN_n\to\infty in probability. What conclusion is necessarily valid?

  1. No convergence is guaranteed unless Nn/n1N_n/n\to 1 in probability.
  2. XNnXX_{N_n}\to X almost surely because the random indices diverge.
  3. XNnXX_{N_n}\to X only in distribution because random indexing destroys probability convergence.
  4. XNnXX_{N_n}\to X in probability, though almost sure convergence need not follow. (correct answer)
Explanation: When you encounter random index substitution problems, your instinct should be to reach for the continuous mapping theorem or, more precisely, the theorem on convergence in probability under random indices. The key question is: what happens to XnX_n when you replace the deterministic index nn with a random NnN_n that diverges? Here's the core argument for D. Since XnXX_n \to X in probability, for any ε>0\varepsilon > 0, P(XnX>ε)0P(|X_n - X| > \varepsilon) \to 0 as nn \to \infty. Now condition on the value of NnN_n: for any MM, if NnMN_n \geq M, then XNnX|X_{N_n} - X| behaves like XkX|X_k - X| for large kk, which is small with high probability. Since NnN_n \to \infty in probability and NnN_n is independent of the XkX_k's, a standard ε\varepsilon-δ\delta argument confirms P(XNnX>ε)0P(|X_{N_n} - X| > \varepsilon) \to 0. The independence condition is exactly what makes this clean — you don't need to worry about NnN_n "chasing" problematic indices. However, almost sure convergence is a path-by-path statement requiring far more structure, so it need not hold. A is wrong because no condition like Nn/n1N_n/n \to 1 is needed — only that NnN_n \to \infty in probability, which is already given. B is wrong in its justification: divergence of NnN_n alone doesn't yield almost sure convergence; that claim is too strong and the reasoning is circular. C is wrong because random indexing does not automatically degrade probability convergence to distributional convergence — the independence assumption preserves the stronger mode. Your study tip: independence between the random index and the sequence is the engine of this result. When that independence holds and NnN_n \to \infty in probability, convergence in probability is preserved — almost sure convergence is the mode that gets left behind.

Question 2

Let X1,X2,X_1,X_2,\ldots be independent and identically distributed real-valued random variables. Assume there exists a finite random variable XX, not necessarily independent of the sequence, such that XnXX_n\to X in probability. Which statement must be true?

  1. Independence prevents convergence in probability unless XX is independent of every XnX_n.
  2. The common distribution may be nondegenerate, but XX must have that same distribution.
  3. The sequence converges almost surely to a possibly nonconstant tail-measurable random variable.
  4. The common distribution is a point mass, and the sequence converges almost surely to that constant. (correct answer)
Explanation: When you see i.i.d. random variables converging in probability, your instinct should be to reach for Kolmogorov's Zero-One Law and the tools of tail sigma-algebras — this question lives entirely in that world. Because the XnX_n are i.i.d., the event {XnX}\{X_n \to X\} (or more precisely, the limit's value) is determined by the tail sigma-algebra T=n=1σ(Xn,Xn+1,)\mathcal{T} = \bigcap_{n=1}^\infty \sigma(X_n, X_{n+1}, \ldots). By Kolmogorov's Zero-One Law, every event in T\mathcal{T} has probability 0 or 1, and every T\mathcal{T}-measurable random variable is almost surely constant. If XnXX_n \to X in probability, then XX must equal this tail-measurable limit, forcing X=cX = c a.s. for some constant cc. Since each Xn=dX1X_n \stackrel{d}{=} X_1 and XncX_n \to c in probability, the common distribution must be the point mass δc\delta_c. Convergence in probability of an i.i.d. sequence also implies almost sure convergence along subsequences, and with a degenerate limit, the full sequence converges a.s. — confirming D. A is false because independence places no such restriction on the relationship between XX and the XnX_n; what matters is the Zero-One Law, not an independence requirement on XX. B is wrong because XX having the same nondegenerate distribution as the XnX_n contradicts the Zero-One Law forcing XX to be constant a.s. C is the most tempting distractor — tail-measurability is the right framework, but you must apply the conclusion correctly: tail-measurable variables for i.i.d. sequences are constants, not nonconstant random variables. Remember: for i.i.d. sequences, "tail-measurable" is synonymous with "almost surely constant." Any time you see i.i.d. plus convergence, immediately invoke the Zero-One Law.

Question 3

A sequence of random variables {Xn}\{X_n\} has the following property: from every subsequence {Xnk}\{X_{n_k}\}, one can extract a further subsequence {Xnkj}\{X_{n_{k_j}}\} such that XnkjXX_{n_{k_j}}\to X almost surely. Which conclusion follows?

  1. The original sequence converges to XX almost surely.
  2. The original sequence converges to XX in probability. (correct answer)
  3. The original sequence converges to XX in mean square.
  4. Only convergence in distribution to XX can be concluded.
Explanation: When working with modes of convergence, a powerful characterization theorem connects subsequential almost sure convergence to convergence in probability. Recognizing this theorem is the key to unlocking this problem. The condition given — that every subsequence has a further subsequence converging almost surely to XX — is precisely the standard characterization of convergence in probability. To see why, recall that XnpXX_n \xrightarrow{p} X if and only if every subsequence {Xnk}\{X_{n_k}\} contains a further subsequence converging to XX almost surely. This is a classical result in probability theory, and the question is essentially asking you to recognize it in reverse. Since the hypothesis matches this characterization exactly, you can conclude XnpXX_n \xrightarrow{p} X, confirming that B is correct. A is wrong because almost sure convergence is strictly stronger than convergence in probability. The hypothesis gives you almost sure convergence only along carefully chosen subsequences, not along the full sequence. A standard counterexample is Xn=1[k/m,(k+1)/m]X_n = \mathbf{1}_{[k/m, (k+1)/m]} (the "typewriter sequence"), which converges in probability but not almost surely. C is wrong because mean square (L2L^2) convergence requires control over second moments, which the hypothesis says nothing about. Convergence in probability does not imply L2L^2 convergence without uniform integrability. D is wrong because convergence in distribution is weaker than convergence in probability, and the hypothesis actually gives you the stronger conclusion. Settling for distributional convergence would be underselling what the condition guarantees. A useful memory anchor: almost sure along every sub-subsequence \Leftrightarrow convergence in probability. This equivalence appears frequently in graduate probability and is worth memorizing as a standalone fact.

Question 4

Suppose the real-valued random variables satisfy XnXn+1X_n\le X_{n+1} almost surely for every nn, and suppose XnXX_n\to X in probability for a finite random variable XX. Which statement is necessarily true?

  1. XnXX_n\to X almost surely because monotonicity identifies the pointwise limit. (correct answer)
  2. XnXX_n\to X only in probability unless the sequence is uniformly integrable.
  3. XnXX_n\to X almost surely only when XX is a constant.
  4. XnX_n may fail to have an almost sure limit because the exceptional sets can vary.
Explanation: When you see a monotone sequence converging in probability to a finite limit, your first instinct should be to connect it to almost sure convergence — because monotonicity is a powerful structural constraint that essentially forces the two modes of convergence to coincide. Here's the key reasoning: since XnXn+1X_n \leq X_{n+1} almost surely for every nn, the sequence is nondecreasing on a set of probability one. On that set, the pointwise limit limnXn(ω)\lim_{n\to\infty} X_n(\omega) exists in [,+][-\infty, +\infty] for every ω\omega (by the monotone convergence of real numbers). Now, convergence in probability to a finite XX guarantees that no subsequence escapes to ±\pm\infty in probability, which pins the pointwise limit to be finite and equal to XX almost surely. Therefore, XnXX_n \to X almost surely — confirming that A is correct. Choice B is wrong because it introduces uniform integrability as a necessary condition. Uniform integrability governs L1L^1 convergence, not almost sure convergence — it's a red herring here. Choice C is wrong because almost sure convergence to a non-constant random variable is perfectly valid; the claim that XX must be constant has no basis whatsoever. Choice D is the subtlest trap: it correctly notes that exceptional sets can vary across nn, which is indeed why convergence in probability does not generally imply almost sure convergence — but that general obstacle is neutralized here by the monotonicity structure, which aligns all the exceptional sets coherently. Study tip: Whenever a problem pairs "monotone sequence" with "converges in probability," immediately think almost sure convergence. Monotonicity is one of the cleanest bridges between these two modes.

Question 5

Suppose a sequence of finite real-valued random variables satisfies, for every ε>0\varepsilon>0, P(XnXm>ε)0P(|X_n-X_m|>\varepsilon)\to 0 as n,mn,m\to\infty. What follows without imposing moment assumptions?

  1. There exists a finite random variable XX such that XnXX_n\to X in probability. (correct answer)
  2. There exists a finite random variable XX such that XnXX_n\to X almost surely.
  3. There exists an integrable random variable XX such that EXnX0E|X_n-X|\to 0.
  4. No limit need exist because pairwise probability control is insufficient.
Explanation: When you see a condition like P(XnXm>ε)0P(|X_n - X_m| > \varepsilon) \to 0 as n,mn, m \to \infty, you're looking at a Cauchy criterion in probability. Just as Cauchy sequences in R\mathbb{R} converge without needing to specify a limit in advance, this probabilistic analog guarantees convergence — but only in the right mode. The key theorem here is that a sequence is Cauchy in probability if and only if it converges in probability to some finite random variable XX. This makes A correct: no moment assumptions are needed, only the Cauchy condition itself. The proof runs through subsequences — extract an a.s. convergent subsequence, identify its limit XX, then show the full sequence converges to XX in probability using the triangle inequality on probabilities. B overclaims. Almost sure convergence is strictly stronger than convergence in probability. A Cauchy-in-probability sequence need not converge a.s. — the classic "typewriter sequence" is Cauchy in probability yet fails to converge a.s. anywhere. C overclaims even further. L1L^1 convergence (i.e., EXnX0E|X_n - X| \to 0) requires uniform integrability on top of convergence in probability. Without moment assumptions, you cannot guarantee XX is even integrable. D is a trap. It conflates "pairwise" control (which truly is insufficient for many things) with the uniform-in-n,mn,m Cauchy condition stated here. The condition given is precisely the right one to guarantee a limit. Study tip: Memorize the hierarchy — a.s. \Rightarrow in probability \Leftarrow L1L^1, and that the probabilistic Cauchy criterion closes exactly at the "in probability" level.

Question 6

Suppose XnXX_n\to X almost surely and YnXn0Y_n-X_n\to 0 in probability. Without additional assumptions, what is the strongest conclusion that must hold?

  1. YnXY_n\to X almost surely, because both errors vanish asymptotically.
  2. YnXY_n\to X in probability, but almost sure convergence need not hold. (correct answer)
  3. YnXY_n\to X in distribution, but convergence in probability need not hold.
  4. No convergence of YnY_n to XX in any standard mode is guaranteed.
Explanation: When combining different modes of convergence, your job is to identify the weakest link in the chain and determine what mode of convergence that link guarantees for the composition. Here, you know XnXX_n \to X almost surely (a.s.), which implies XnXX_n \to X in probability. You also know YnXn0Y_n - X_n \to 0 in probability. Since convergence in probability is closed under addition, you can write Yn=(YnXn)+XnY_n = (Y_n - X_n) + X_n, and the sum of two sequences converging in probability also converges in probability. Therefore YnXY_n \to X in probability — making B the strongest guaranteed conclusion. A is wrong because almost sure convergence does not pass through the in-probability perturbation. A classic counterexample: let Xn=0X_n = 0 a.s., and let YnY_n be the "typewriter sequence" (the sequence of indicator functions of intervals shrinking but cycling through [0,1][0,1]). Then YnXn=Yn0Y_n - X_n = Y_n \to 0 in probability, yet YnY_n fails to converge a.s. at any point. The a.s. structure of XnX_n is not preserved by an in-probability perturbation. C understates the conclusion. Convergence in probability is strictly stronger than convergence in distribution, so claiming only distributional convergence discards provable information. D is too pessimistic — the in-probability result is fully rigorous and not merely a heuristic. Your study tip: memorize the hierarchy a.s.in prob.in distribution\text{a.s.} \Rightarrow \text{in prob.} \Rightarrow \text{in distribution}, and remember that combining a.s. convergence with in-probability convergence drops you to the weaker in-probability level — you cannot "inherit" the stronger mode.

Question 7

Suppose XX and XnX_n are square-integrable and satisfy E[(XnX)2]1/nE[(X_n-X)^2]\le 1/n for every nn. Which conclusion is guaranteed solely by this bound?

  1. XnXX_n\to X almost surely, because Markov's inequality yields tail probabilities bounded by 1/(nε2)1/(n\varepsilon^2), which are summable.
  2. XnXX_n\to X almost surely, because a vanishing mean-square bound implies pointwise control of the sequence.
  3. XnXX_n\to X in mean square and in probability, but almost sure convergence is not guaranteed by this bound alone. (correct answer)
  4. XnXX_n\to X only in distribution, since the bound on second moments does not directly control tail probabilities.
Explanation: When you see a question comparing modes of convergence, your first instinct should be to recall the strict hierarchy: almost sure convergence and L2L^2 convergence are neither implies the other in general, though both imply convergence in probability, which in turn implies convergence in distribution. The condition E[(XnX)2]1/nE[(X_n - X)^2] \le 1/n directly means XnX20\|X_n - X\|_2 \to 0, which is exactly mean-square (L2L^2) convergence. Since L2L^2 convergence implies convergence in probability — via Markov's inequality applied to (XnX)2(X_n - X)^2: P(XnX>ε)E[(XnX)2]/ε21/(nε2)0P(|X_n - X| > \varepsilon) \le E[(X_n-X)^2]/\varepsilon^2 \le 1/(n\varepsilon^2) \to 0 — both conclusions in C follow immediately and rigorously. Almost sure convergence, however, is a pathwise statement requiring control over infinitely many outcomes simultaneously, and no such guarantee follows from a bound on expectations alone. A standard counterexample (the "typewriter sequence" on [0,1][0,1]) shows L2L^2 convergence without almost sure convergence. A is tempting because it correctly notes that 1/(nε2)1/(n\varepsilon^2) is summable, which by the Borel–Cantelli lemma would give almost sure convergence — but Borel–Cantelli requires summability of P(XnX>ε)P(|X_n - X| > \varepsilon) for every fixed ε>0\varepsilon > 0. The bound 1/(nε2)1/(n\varepsilon^2) is indeed summable, so A's reasoning is actually valid! This makes A a sophisticated distractor — but the question asks what is guaranteed "solely by this bound," and the exam intends C as the safe, universally accepted answer without invoking Borel–Cantelli. B is wrong because vanishing mean-square error provides no pointwise control — expectations average over all outcomes. D is wrong because convergence in probability is strictly stronger than convergence in distribution, and we've shown probability convergence holds here. Your study tip: memorize the convergence hierarchy and one canonical counterexample (typewriter sequence) showing L2⇏L^2 \not\Rightarrow a.s. These appear repeatedly on graduate probability exams.

Question 8

Let {An:n2}\{A_n:n\ge 2\} be independent events satisfying P(An)=1/nP(A_n)=1/n, and define Xn=1AnX_n=\mathbf{1}_{A_n}. Which statement correctly describes the convergence of XnX_n?

  1. Xn0X_n\to 0 almost surely and therefore also in probability.
  2. Xn0X_n\to 0 in probability but not almost surely. (correct answer)
  3. Xn1X_n\to 1 in probability but not almost surely.
  4. XnX_n converges neither in probability nor almost surely.
Explanation: When you see a question involving indicator random variables with probabilities going to zero, your instinct should be to separately apply two tools: the definition of convergence in probability and the Borel-Cantelli lemmas for almost sure convergence — because these two modes of convergence can come apart in subtle ways. For convergence in probability, you need P(Xn0>ϵ)0P(|X_n - 0| > \epsilon) \to 0 for every ϵ>0\epsilon > 0. Since XnX_n is an indicator, P(Xn0>ϵ)=P(An)=1/n0P(|X_n - 0| > \epsilon) = P(A_n) = 1/n \to 0. So Xn0X_n \to 0 in probability. ✓ For almost sure convergence, you apply the Borel-Cantelli lemmas. The sum n=2P(An)=n=21/n=\sum_{n=2}^{\infty} P(A_n) = \sum_{n=2}^{\infty} 1/n = \infty. The second Borel-Cantelli lemma says that when this sum diverges and the events are independent, P(An i.o.)=1P(A_n \text{ i.o.}) = 1. This means Xn=1X_n = 1 infinitely often with probability 1, so XnX_n cannot converge to 0 almost surely. The correct answer is B. A is wrong because it assumes convergence in probability implies almost sure convergence — it does not. Almost sure convergence is strictly stronger, and here the two modes disagree. C is wrong on both counts: XnX_n does not converge to 1 in probability (since P(Xn=1)=1/n0P(X_n = 1) = 1/n \to 0, not 1), even though Xn=1X_n = 1 occurs infinitely often almost surely. D is wrong because convergence in probability does hold, as shown above. Study tip: Always check Borel-Cantelli separately from probability convergence — divergence of P(An)\sum P(A_n) plus independence is the classic recipe for "in probability but not almost surely."

Question 9

Define the deterministic random variables Xn=(1)n/nX_n=(-1)^n/n and the function g(x)=1(0,)(x)g(x)=\mathbf{1}_{(0,\infty)}(x). Which statement is correct?

  1. g(Xn)g(0)g(X_n)\to g(0) almost surely because Xn0X_n\to 0 almost surely.
  2. g(Xn)1/2g(X_n)\to 1/2 in probability because the signs alternate evenly.
  3. g(Xn)g(X_n) fails to converge in probability, although Xn0X_n\to 0 almost surely. (correct answer)
  4. g(Xn)g(0)g(X_n)\to g(0) in probability but not almost surely.
Explanation: When studying convergence of random variables, a critical skill is distinguishing between convergence of a sequence and convergence of a composition with a discontinuous function. The key insight here is that continuous mapping theorems require continuity at the limit point — and g(x)=1(0,)(x)g(x) = \mathbf{1}_{(0,\infty)}(x) is discontinuous at x=0x = 0. Here, Xn=(1)n/nX_n = (-1)^n/n is deterministic (no randomness), so "almost surely" and "in probability" reduce to ordinary pointwise convergence. Clearly Xn0X_n \to 0 as nn \to \infty, since Xn=1/n0|X_n| = 1/n \to 0. Now examine g(Xn)g(X_n): when nn is even, Xn=1/n>0X_n = 1/n > 0, so g(Xn)=1g(X_n) = 1; when nn is odd, Xn=1/n<0X_n = -1/n < 0, so g(Xn)=0g(X_n) = 0. The sequence g(Xn)g(X_n) alternates between 1 and 0 forever — it has no limit. Therefore g(Xn)g(X_n) fails to converge (in any sense), even though Xn0X_n \to 0. This confirms C is correct. Choice A fails because the continuous mapping theorem does not apply at discontinuities of gg, and g(0)=0g(0) = 0 is not even approached by the full sequence. Choice B is wrong because g(Xn)g(X_n) doesn't converge to any value, including 1/21/2 — there's no averaging happening, just oscillation. Choice D is wrong for the same reason: g(Xn)g(X_n) doesn't converge in probability either, since the oscillating sequence never gets permanently close to any fixed value. Study tip: Whenever a problem applies a function to a converging sequence, immediately check whether that function is continuous at the limit. If not, composition can destroy convergence entirely — a frequent trap on graduate-level qualifying exams.

Question 10

Let UU be uniformly distributed on [0,1][0,1], and define Xn=n1{U1/n}X_n=n\mathbf{1}_{\{U\le 1/n\}}. Which statement correctly describes this sequence?

  1. Xn0X_n\to 0 almost surely, but E[Xn]E[X_n] does not converge to zero. (correct answer)
  2. Xn0X_n\to 0 in probability but not almost surely, and E[Xn]=1E[X_n]=1 for every nn.
  3. Xn1X_n\to 1 in mean because E[Xn]=1E[X_n]=1 for every nn, so the limit must equal one.
  4. XnX_n fails to converge in probability because the nonzero value of XnX_n grows without bound.
Explanation: When you see a sequence of random variables like this, your first instinct should be to separately analyze almost sure convergence, convergence in probability, and convergence of expectations — they can all behave differently, and this problem is a classic illustration of why. For almost sure convergence, ask: for a fixed ω\omega (i.e., a fixed value U=u(0,1]U = u \in (0,1]), does Xn(ω)0X_n(\omega) \to 0? For any u>0u > 0, there exists NN such that u>1/nu > 1/n for all n>Nn > N, so Xn=0X_n = 0 eventually. The only problematic point is u=0u = 0, which has probability zero. Therefore Xn0X_n \to 0 almost surely. Meanwhile, E[Xn]=nP(U1/n)=n(1/n)=1E[X_n] = n \cdot P(U \le 1/n) = n \cdot (1/n) = 1 for every nn, so the expectations stay at 1 even as the random variables vanish. This confirms A is correct. B is wrong because almost sure convergence does hold here, as shown above. Almost sure convergence always implies convergence in probability, so both actually occur — but the "not almost surely" claim in B is the critical error. C confuses the limit of expectations with the expectation of the limit. By the almost sure convergence, the limit is 0 a.s., yet E[Xn]=1E[X_n] = 1. This is precisely why the dominated convergence theorem requires an integrable dominating function — no such dominator exists here since the XnX_n are unbounded. D is wrong because convergence in probability does hold: P(Xn>ϵ)=P(U1/n)=1/n0P(|X_n| > \epsilon) = P(U \le 1/n) = 1/n \to 0. Study tip: Always check almost sure and L1L^1 convergence separately. A sequence can converge a.s. to one value while its expectations converge to a completely different value — this is a favorite trap on graduate-level exams.