STATISTICS GRADUATE LEVEL • CONVERGENCE & LIMIT THEOREMS

Convergence in Lp

Understanding how sequences of random variables converge through the lens of p-th moment integrability.

Historical Context & Motivation

The notion of convergence in Lp arose from a deep interplay between the emerging fields of real analysis, measure theory, and probability during the late nineteenth and early twentieth centuries. As mathematicians began to formalize the theory of integration beyond the classical Riemann framework, they required rigorous definitions of what it meant for a sequence of functions—or random variables—to approach a limit. The classical pointwise and uniform notions of convergence proved insufficient for many problems in Fourier analysis and probability theory, and the concept of convergence measured through integral norms emerged as a powerful and natural alternative. The Lp spaces, named in honor of Henri Lebesgue, provided the function-space architecture within which this mode of convergence could be precisely stated and studied.

1902
Lebesgue Integral
Henri Lebesgue publishes his doctoral thesis introducing the Lebesgue integral, laying the measure-theoretic foundation upon which Lp spaces would later be constructed.
1907
Riesz–Fischer Theorem
Frigyes Riesz and Ernst Fischer independently prove the Riesz–Fischer theorem, establishing the completeness of L² and linking Fourier series convergence to convergence in the L² norm.
1910
General Lp Spaces
Riesz extends the framework to general Lp spaces for p ≥ 1, establishing Hölder's and Minkowski's inequalities as the fundamental tools for these Banach spaces.
1933
Kolmogorov's Axioms
Andrey Kolmogorov axiomatizes probability theory using measure theory, making Lp convergence of random variables a central concept in rigorous probability and statistics.
1940s–60s
Modern Probability Theory
The development of martingale theory and ergodic theory by Doob, Kakutani, and others relies heavily on Lp convergence results, cementing its role in advanced stochastic analysis.

The central question that Lp convergence addresses is the following: given a sequence of random variables or measurable functions {Xn}, in what sense can we say that the entire 'shape' of the distribution—captured through its p-th moment—converges to that of a limiting random variable X? Unlike pointwise or almost sure convergence, Lp convergence provides a global, integral-based measure of closeness that accounts for the magnitude of deviations weighted by their probability. This makes it an indispensable tool for proving limit theorems, establishing rates of convergence, and characterizing the behavior of estimators in mathematical statistics.

Core Principles & Definitions

Before diving into the formal machinery, it is essential to understand the conceptual pillars that support convergence in Lp. At its heart, this mode of convergence asks whether the expected value of the p-th power of the absolute difference between Xn and X vanishes as n → ∞. The parameter p controls how severely large deviations are penalized: larger values of p impose stricter requirements on the tails of the distribution. The following foundational ideas underpin the entire theory.

1

The Lp Space

For p ≥ 1, the space Lp(Ω, F, P) consists of all random variables X such that E[|X|p] < ∞. Equipped with the norm ‖X‖p = (E[|X|p])1/p, this is a complete normed vector space (Banach space).
2

Lp Convergence Definition

A sequence {Xn} converges to X in Lp if ‖Xn − X‖p → 0 as n → ∞, i.e., E[|Xn − X|p] → 0.
3

Moment Control

Lp convergence implies that the p-th absolute moments of Xn converge to those of X. In particular, L¹ convergence guarantees convergence of means, and L² convergence guarantees convergence of both means and variances.
4

Hierarchy of Modes

Lp convergence implies Lr convergence for all 1 ≤ r ≤ p (by Jensen's inequality). It also implies convergence in probability, but not vice versa. Almost sure convergence does not automatically imply Lp convergence without additional integrability conditions.
5

Uniform Integrability Bridge

The critical link between convergence in probability and Lp convergence is uniform integrability. A sequence converges in Lp if and only if it converges in probability and {|Xn|p} is uniformly integrable.
KEY TAKEAWAY
Think of Lp convergence like evaluating student exam performance across an entire class, not just checking whether each individual student passes. Pointwise (almost sure) convergence asks whether each student's score approaches the target; Lp convergence asks whether the overall weighted average discrepancy across the whole class vanishes. The exponent p controls how harshly you penalize the worst-performing students: a higher p means even a few large outlier deviations will prevent convergence, much like an instructor who weights the largest errors most heavily when computing a summary grade.

Visual Explanation

The following diagram illustrates the hierarchy of convergence modes in probability theory and where Lp convergence sits within that hierarchy. Each arrow indicates an implication: if mode A implies mode B, an arrow points from A to B. The conditions under which reverse implications hold are annotated along the edges. Understanding this hierarchy is essential, as it clarifies when Lp convergence can be deduced from other modes and what additional hypotheses (such as uniform integrability or domination) are needed.

The hierarchy of convergence modes in probability. Solid arrows denote unconditional implications; dashed arrows indicate implications that hold only under additional conditions such as uniform integrability. Lp convergence (top, violet) implies both Lr convergence for r ≤ p and convergence in probability.

Observe that Lp convergence occupies a strong position in the hierarchy: it implies convergence in probability (and hence convergence in distribution), and it also implies Lr convergence whenever r ≤ p. The reverse implication from convergence in probability back to Lp requires the additional condition of uniform integrability of the sequence {|Xn − X|p}. Almost sure convergence, while a strong pointwise condition, does not by itself imply Lp convergence—one needs domination or uniform integrability to bridge the gap. This interplay among modes of convergence is one of the most elegant and practically important aspects of measure-theoretic probability.

Mathematical Framework

We now formalize the definition and key properties of Lp convergence, along with the essential inequalities and theorems that govern its behavior. Let (Ω, F, P) be a probability space and let p ∈ [1, ∞).

DEFINITION OF Lp CONVERGENCE
X_n →^{Lp} X ⟺ lim_{n→∞} E[|X_n − X|^p] = 0
Here Xn and X are random variables in Lp, meaning E[|Xn|p] < ∞ and E[|X|p] < ∞. Equivalently, ‖Xn − X‖p = (E[|Xn − X|p])1/p → 0.
MARKOV'S INEQUALITY (Lp ⟹ CONVERGENCE IN PROBABILITY)
P(|X_n − X| > ε) ≤ E[|X_n − X|^p] / ε^p → 0
For any ε > 0 and p ≥ 1, Markov's inequality shows that Lp convergence implies convergence in probability. This is the fundamental mechanism linking integral-based and distributional notions of convergence.
LYAPUNOV'S INEQUALITY (Lp ⟹ Lr FOR r ≤ p)
‖X‖_r = (E[|X|^r])^{1/r} ≤ (E[|X|^p])^{1/p} = ‖X‖_p for 1 ≤ r ≤ p
A consequence of Jensen's inequality applied to the convex function t ↦ tp/r. This establishes that Lp ⊆ Lr whenever r ≤ p, and in particular, convergence in Lp implies convergence in Lr.
VITALI CONVERGENCE THEOREM
X_n →^{Lp} X ⟺ X_n →^P X and {|X_n|^p} is uniformly integrable
This is the definitive characterization: Lp convergence is equivalent to convergence in probability plus uniform integrability of the p-th powers. Recall that a family {Yα} is uniformly integrable if supα E[|Yα| 𝟙(|Yα| > M)] → 0 as M → ∞.
📐 Dominated Convergence as a Special Case
The Dominated Convergence Theorem (DCT) provides a sufficient condition for Lp convergence: if Xn → X a.s. (or in probability) and |Xn| ≤ Y a.s. for some Y ∈ Lp, then Xn → X in Lp. Domination by a single integrable function automatically ensures uniform integrability, so the DCT is a corollary of the Vitali convergence theorem.

Detailed Breakdown: Relationships Among Convergence Modes

One of the most important skills in advanced probability is knowing exactly when one mode of convergence implies another and when it does not. The following diagram provides a visual summary of the key counterexamples and sufficient conditions, focusing on the role that the parameter p and the underlying probability space play in determining the strength of Lp convergence.

This diagram shows how the parameter p affects the strength of Lp convergence, along with two classic counterexamples demonstrating that L¹ convergence does not imply L² convergence and that almost sure convergence does not imply L¹ convergence without domination. The green summary box at the bottom collects the key implication chains.

The counterexamples above are instructive. In the left box, the sequence Xn = n · 𝟙(U ≤ 1/n²) has spikes of height n on a set of probability 1/n², which is enough for the first moment to vanish but the second moment to remain constant. In the right box, Xn = n² · 𝟙(U ≤ 1/n) converges to zero almost surely because ∑ P(Xn ≠ 0) = ∑ 1/n = ∞ but the individual probabilities shrink, and by Borel–Cantelli's second moment method one can verify a.s. convergence. However, the expectation E[Xn] = n → ∞, so L¹ convergence fails catastrophically. These examples cement the lesson that tail behavior and uniform integrability are the gatekeepers of Lp convergence.

Worked Example

Let us work through a complete example that illustrates how to verify Lp convergence using both direct computation and the Vitali convergence theorem.

Verifying L² Convergence
1
Step 1 — Define the SequenceLet U ~ Uniform(0, 1) and define Xn = √n · 𝟙(U ≤ 1/n) for n ≥ 1. We want to determine whether Xn → 0 in L². Note that Xn takes the value √n with probability 1/n and 0 otherwise.
2
Step 2 — Compute E[|Xₙ|²]E[|Xn|²] = E[(√n)² · 𝟙(U ≤ 1/n)] = n · P(U ≤ 1/n) = n · (1/n) = 1. Since the limit random variable is X = 0, we need E[|Xn − 0|²] = E[|Xn|²] = 1 for all n.
E[|Xn|²] = 1 ≠ 0, so Xn does NOT converge to 0 in L².
3
Step 3 — Check Convergence in ProbabilityFor any ε > 0, P(|Xn| > ε) = P(U ≤ 1/n) = 1/n → 0 as n → ∞. So Xn → 0 in probability. This confirms that convergence in probability does not imply L² convergence.
Xn → 0 in probability ✓
4
Step 4 — Diagnose the Failure via Uniform IntegrabilityBy the Vitali theorem, the failure of L² convergence despite convergence in probability must be due to a failure of uniform integrability of {|Xn|²}. Indeed, E[|Xn|² · 𝟙(|Xn|² > M)] = n · (1/n) · 𝟙(n > M) = 𝟙(n > M), which equals 1 for all n > M. Hence supn E[|Xn|² · 𝟙(|Xn|² > M)] = 1 for all M, confirming that {|Xn|²} is NOT uniformly integrable.
Uniform integrability fails ⟹ Vitali theorem confirms L² convergence does not hold.
5
Step 5 — Modify for L¹ ConvergenceNow check L¹: E[|Xn|] = √n · (1/n) = 1/√n → 0. Since E[|Xn − 0|¹] → 0, we have Xn → 0 in L¹. This example beautifully demonstrates how the same sequence can converge in one Lp space but fail in another—the spikes are tall enough (√n) that their squares (n) are not integrable uniformly, but their first powers (√n · 1/n = 1/√n) decay.
Xn → 0 in L¹ but NOT in L². Convergence mode depends critically on p.

Strengths, Limitations & Comparisons

Each mode of convergence has its own advantages and drawbacks depending on the context. The following table compares Lp convergence with the other principal modes, highlighting the situations in which each excels and the pitfalls to watch for.

Comparison of convergence modes in probability theory
Mode of ConvergenceStrengthsLimitations
Lp ConvergenceControls p-th moments; implies convergence of means (p=1) and variances (p=2); metrizable via ‖·‖p; easily upgradable to stronger p via domination.Requires all variables to belong to Lp; does not imply almost sure convergence; sensitive to tail behavior; harder to verify directly than convergence in probability.
Almost Sure ConvergenceStrongest pathwise statement; intuitive (sample paths converge); implies convergence in probability.Does not imply Lp convergence without domination or U.I.; not metrizable in general; difficult to establish for dependent sequences.
Convergence in ProbabilityWeaker and easier to verify; sufficient for many statistical applications; metrizable via Ky Fan metric.Does not control moments; does not imply a.s. or Lp convergence; conclusions about expectations require additional arguments.
Convergence in DistributionWeakest mode; sufficient for CLT-type results; does not require variables on the same probability space.Says nothing about moments or pointwise behavior; cannot conclude E[Xn] → E[X] in general.
KEY TAKEAWAY
Lp convergence is the mode of choice in statistics whenever you need guarantees about the behavior of moments and expectations. For instance, when proving that the mean squared error of an estimator converges to zero—establishing L² consistency—you are implicitly using L² convergence. In applications like portfolio optimization in mathematical finance, where controlling the variance (second moment) of returns is paramount, Lp convergence with p = 2 provides exactly the right analytical framework. Think of Lp convergence as a quality assurance standard in manufacturing: it doesn't just check that individual products are approximately correct (pointwise), but ensures that the aggregate deviation metric across the entire production line meets tolerance.

Connections to Advanced Theory

Lp convergence is not merely a foundational concept—it serves as a launching point for several deep areas of modern probability and statistics. The following table summarizes how Lp convergence connects to more advanced topics that students encounter in graduate-level coursework and research.

Connections from Lp convergence to advanced topics in probability and statistics
Lp Convergence ConceptAdvanced ExtensionKey Connection
L² convergence of estimatorsMean Squared Error consistencyAn estimator θ̂n is L²-consistent if E[(θ̂n − θ)²] → 0, equivalently bias² + variance → 0.
Uniform integrabilityMartingale convergence (Doob)An L¹-bounded martingale converges a.s.; with U.I., it also converges in L¹, and Mn = E[M | Fn].
Lp norm as distanceWasserstein distancesThe Wasserstein-p distance Wp(μ, ν) is the Lp optimal transport cost; convergence in Wp is equivalent to convergence in distribution plus convergence of p-th moments.
Completeness of LpRiesz–Fischer and Fourier analysisThe completeness of L² guarantees that Fourier series of L² functions converge in the L² norm, forming the mathematical backbone of signal processing and spectral analysis.
Lp rates of convergenceNonparametric estimation theoryMinimax rates for density estimation and regression are typically stated in L² (MISE); the rate n−2s/(2s+1) for Sobolev-s smooth densities is a hallmark result.

As you progress deeper into measure-theoretic probability, you will find that Lp convergence appears repeatedly in the proofs of major theorems. The Vitali convergence theorem will reappear in the study of martingales, where uniform integrability characterizes exactly when a martingale converges not only almost surely but also in L¹. In statistical learning theory, L² convergence rates quantify how quickly an estimator's risk decreases with sample size, and these rates are central to determining optimal bandwidth selection for kernel estimators, optimal penalties in regularization, and minimax lower bounds. The interplay between Lp convergence, uniform integrability, and tightness continues to be a fertile area connecting probability, functional analysis, and optimal transport.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain why Lp convergence for p = 2 implies L¹ convergence, but the converse is not generally true. What role does the exponent p play in controlling the sensitivity to large deviations?
PROBLEM 2BASIC CALCULATION
Let U ~ Uniform(0,1) and define Xn = n1/3 · 𝟙(U ≤ 1/n). Determine whether Xn → 0 in L¹ and in L³.
PROBLEM 3INTERMEDIATE
Suppose Xn → X in probability and |Xn| ≤ Y almost surely where E[Y²] < ∞. Prove that Xn → X in L².
PROBLEM 4APPLIED
In a statistical estimation context, let θ̂n = (1/n)∑Xi where X₁, X₂, ... are i.i.d. with mean μ and variance σ². Show that θ̂n → μ in L². What is the rate of convergence in the L² norm?
PROBLEM 5CRITICAL THINKING
Construct a sequence of random variables {Xn} that converges to 0 in Lp for every p ∈ [1, ∞) but does NOT converge to 0 almost surely. Prove both claims.

Summary

Convergence in Lp is a mode of convergence defined by the condition E[|Xn − X|p] → 0, which measures closeness through the p-th moment of the absolute difference. It occupies a strong position in the hierarchy of convergence modes: it implies Lr convergence for r ≤ p (via Lyapunov's inequality) and convergence in probability (via Markov's inequality), but it does not imply almost sure convergence without additional conditions.

The Vitali convergence theorem provides the definitive characterization: Lp convergence is equivalent to convergence in probability combined with uniform integrability of the p-th powers. The Dominated Convergence Theorem is a powerful sufficient condition, since domination by a single Lp function automatically ensures U.I. In statistical practice, L² convergence is especially important because it controls the mean squared error of estimators, providing the theoretical foundation for consistency and rate-of-convergence results throughout parametric and nonparametric inference.

Varsity Tutors • Statistics Graduate Level • Convergence in Lp