Statistics Graduate Level Quiz: Jensens Inequality
10 questions · exam conditions
0:00
Jensens InequalityQuestion 1 of 10

Let TT be an integrable estimator of a parameter θ\theta, and let SS be a statistic. Define the Rao–Blackwellized estimator T=E[TS]T^*=\mathbb{E}[T\mid S]. Under absolute-error loss, which statement is guaranteed without assuming that SS is sufficient or that TT is unbiased?

ETθETθ\mathbb{E}|T^*-\theta|\leq \mathbb{E}|T-\theta|, by the convexity of xxθx\mapsto|x-\theta| and Jensen's inequality applied conditionally.
ETθ<ETθ\mathbb{E}|T^*-\theta|<\mathbb{E}|T-\theta| whenever Var(TS)>0\operatorname{Var}(T\mid S)>0 with positive probability, because absolute value is strictly convex away from zero.
ETθETθ\mathbb{E}|T^*-\theta|\geq \mathbb{E}|T-\theta|, because conditioning on SS removes variation around θ\theta and inflates the risk.
No risk ordering is guaranteed unless SS is sufficient, because the Rao–Blackwell theorem requires sufficiency to ensure the conditional expectation does not depend on θ\theta.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Jensens Inequality

Practice Jensens Inequality in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Jensens Inequality, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Let TT be an integrable estimator of a parameter θ\theta, and let SS be a statistic. Define the Rao–Blackwellized estimator T=E[TS]T^*=\mathbb{E}[T\mid S]. Under absolute-error loss, which statement is guaranteed without assuming that SS is sufficient or that TT is unbiased?

  1. ETθETθ\mathbb{E}|T^*-\theta|\leq \mathbb{E}|T-\theta|, by the convexity of xxθx\mapsto|x-\theta| and Jensen's inequality applied conditionally. (correct answer)
  2. ETθ<ETθ\mathbb{E}|T^*-\theta|<\mathbb{E}|T-\theta| whenever Var(TS)>0\operatorname{Var}(T\mid S)>0 with positive probability, because absolute value is strictly convex away from zero.
  3. ETθETθ\mathbb{E}|T^*-\theta|\geq \mathbb{E}|T-\theta|, because conditioning on SS removes variation around θ\theta and inflates the risk.
  4. No risk ordering is guaranteed unless SS is sufficient, because the Rao–Blackwell theorem requires sufficiency to ensure the conditional expectation does not depend on θ\theta.
Explanation: When you see a question about Rao–Blackwellization, your first instinct should be to reach for Jensen's inequality — not the sufficiency condition. Many students conflate the classical Rao–Blackwell theorem (which requires sufficiency to guarantee the conditional expectation is a valid, parameter-free estimator) with the purely probabilistic risk comparison, which holds more generally. Here's the core argument for A: For any fixed θ\theta, the function xxθx \mapsto |x - \theta| is convex. By the conditional form of Jensen's inequality, E[TS]θE[TθS]|\mathbb{E}[T \mid S] - \theta| \leq \mathbb{E}[|T - \theta| \mid S] almost surely. Taking expectations of both sides gives ETθETθ\mathbb{E}|T^* - \theta| \leq \mathbb{E}|T - \theta|. This requires only integrability of TT — no sufficiency, no unbiasedness. A is therefore always guaranteed. B is tempting but wrong. Absolute value xθ|x - \theta| is not strictly convex everywhere — it has a "kink" at x=θx = \theta and is linear on each side. Strict convexity arguments that yield strict inequality (like those valid for x2x^2) fail here; you cannot guarantee strict improvement under absolute-error loss even when Var(TS)>0\operatorname{Var}(T \mid S) > 0. C is simply backwards. Conditioning smooths out variation via Jensen's, which weakly reduces risk, not inflates it. D conflates two separate issues. Sufficiency ensures TT^* doesn't depend on θ\theta (making it a legitimate estimator), but the risk inequality follows from Jensen alone and needs no sufficiency assumption. Study tip: Always separate the probabilistic question (does Jensen apply?) from the statistical question (is the result a valid estimator?). Sufficiency matters for the latter, not the former.

Question 2

Let X1,,XnX_1,\ldots,X_n be positive, identically distributed random variables with finite common mean μ\mu. They are not assumed independent. Define X=n1i=1nXi\overline{X}=n^{-1}\sum_{i=1}^n X_i and assume ElogX<\mathbb{E}|\log \overline{X}|<\infty. Which is the sharp Jensen conclusion?

  1. E[logX]logμ\mathbb{E}[\log \overline{X}]\leq \log\mu, with equality exactly when X=μ\overline{X}=\mu almost surely. (correct answer)
  2. E[logX]logμ\mathbb{E}[\log \overline{X}]\leq \log\mu, with equality exactly when every Xi=μX_i=\mu almost surely.
  3. E[logX]logμ\mathbb{E}[\log \overline{X}]\geq \log\mu, with equality exactly when X=μ\overline{X}=\mu almost surely.
  4. E[logX]logμ\mathbb{E}[\log \overline{X}]\leq \log\mu, with equality exactly when the XiX_i are mutually independent.
Explanation: Whenever you see a question invoking Jensen's inequality, your first move should be to identify the function's concavity and what random variable it's being applied to. Here, f(x)=logxf(x) = \log x is concave on the positive reals. Jensen's inequality for a concave function states that E[f(Y)]f(E[Y])\mathbb{E}[f(Y)] \leq f(\mathbb{E}[Y]) for any integrable random variable YY. Apply this directly to Y=XY = \overline{X}: since E[X]=μ\mathbb{E}[\overline{X}] = \mu, you immediately get E[logX]logE[X]=logμ\mathbb{E}[\log \overline{X}] \leq \log \mathbb{E}[\overline{X}] = \log \mu. Notice that the entire argument hinges only on X\overline{X} itself — no assumptions about independence or the joint structure of the XiX_i are needed. The equality condition for Jensen with a strictly concave function is that the argument must be almost surely constant — that is, X=c\overline{X} = c a.s. for some constant, and since E[X]=μ\mathbb{E}[\overline{X}] = \mu, that constant must be μ\mu. This confirms A is correct. B is wrong because it over-specifies the equality condition. Requiring each Xi=μX_i = \mu a.s. is sufficient but not necessary — you only need X=μ\overline{X} = \mu a.s., which can happen in non-trivial ways (e.g., dependent XiX_i that average to a constant without each being constant). C reverses the inequality direction, confusing concavity with convexity. D is a classic distractor that injects independence into a result that requires no such assumption — Jensen's inequality is purely a statement about the distribution of a single random variable. Your study tip: always ask "what random variable am I applying Jensen to?" The equality condition belongs to that variable being a.s. constant — not to any underlying structure that produces it.

Question 3

Suppose XX is nondegenerate, E[X]=0\mathbb{E}[X]=0, and its moment-generating function MX(t)=E[etX]M_X(t)=\mathbb{E}[e^{tX}] is finite for every real tt. Let KX(t)=logMX(t)K_X(t)=\log M_X(t). Which statement follows directly from the strict Jensen's inequality applied to xetxx\mapsto e^{tx}?

  1. KX(t)>0K_X(t)>0 for every t0t\neq 0, regardless of the sign of tt. (correct answer)
  2. KX(t)>0K_X(t)>0 for t>0t>0 and KX(t)<0K_X(t)<0 for t<0t<0, reflecting the asymmetry of the exponential.
  3. KX(t)0K_X(t)\geq 0 for every tt, but equality may occur at some nonzero tt for suitable distributions.
  4. KX(t)tE[X]+t2Var(X)/2K_X(t)\leq t\mathbb{E}[X]+t^2\operatorname{Var}(X)/2 for every real tt, as a direct consequence of strict convexity.
Explanation: When you see a question involving the cumulant generating function KX(t)=logMX(t)K_X(t) = \log M_X(t) and Jensen's inequality, your instinct should be to apply strict convexity of xetxx \mapsto e^{tx} (for t0t \neq 0) to a nondegenerate random variable. Here's the core reasoning: since etxe^{tx} is strictly convex and XX is nondegenerate (not a point mass), strict Jensen's inequality gives E[etX]>etE[X]\mathbb{E}[e^{tX}] > e^{t\mathbb{E}[X]}. Because E[X]=0\mathbb{E}[X] = 0, this becomes MX(t)>e0=1M_X(t) > e^0 = 1 for every t0t \neq 0. Taking logarithms (a strictly increasing function) preserves the inequality: KX(t)=logMX(t)>log1=0K_X(t) = \log M_X(t) > \log 1 = 0. This holds for all t0t \neq 0, confirming A is correct. The key insight is that strict convexity plus nondegeneracy forces MX(t)>1M_X(t) > 1 symmetrically — both positive and negative tt yield the same conclusion. B is wrong because it claims asymmetry. The argument above is symmetric in tt: for negative tt, etXe^{tX} is still strictly convex, so MX(t)>1M_X(t) > 1 still holds, giving KX(t)>0K_X(t) > 0. C is wrong because it weakens the conclusion to 0\geq 0 and suggests equality might occur at nonzero tt. But strict Jensen's inequality (nondegeneracy is essential here!) guarantees strict inequality. D is a distractor involving a Gaussian-style upper bound on KX(t)K_X(t). This would be an upper bound direction and doesn't follow from Jensen applied here — it conflates Jensen with a Taylor/moment bound. Study tip: Memorize the chain — nondegeneracy + strict convexity → strict Jensen → MX(t)>1M_X(t) > 1KX(t)>0K_X(t) > 0. This chain is symmetric in tt, so never assume sign-dependence without cause.

Question 4

Let WW, XX, and YY be random variables on the same probability space, with 0W10\leq W\leq 1 almost surely. Let ϕ\phi be convex, and assume all displayed expectations are finite. Which inequality is valid even when WW is dependent on both XX and YY?

  1. Eϕ(WX+(1W)Y)E[W]Eϕ(X)+(1E[W])Eϕ(Y)\mathbb{E}\phi(WX+(1-W)Y)\leq\mathbb{E}[W]\mathbb{E}\phi(X)+(1-\mathbb{E}[W])\mathbb{E}\phi(Y)
  2. Eϕ(WX+(1W)Y)E[Wϕ(X)+(1W)ϕ(Y)]\mathbb{E}\phi(WX+(1-W)Y)\leq\mathbb{E}[W\phi(X)+(1-W)\phi(Y)] (correct answer)
  3. Eϕ(WX+(1W)Y)E[Wϕ(X)+(1W)ϕ(Y)]\mathbb{E}\phi(WX+(1-W)Y)\geq\mathbb{E}[W\phi(X)+(1-W)\phi(Y)]
  4. ϕ(E[WX+(1W)Y])E[Wϕ(X)+(1W)ϕ(Y)]\phi(\mathbb{E}[WX+(1-W)Y])\geq\mathbb{E}[W\phi(X)+(1-W)\phi(Y)]
Explanation: When you see a question mixing convexity with a random weight, your first instinct should be Jensen's inequality — but the key is identifying which version applies and whether independence is required. The core tool here is the conditional Jensen's inequality: for convex ϕ\phi and any random weight W[0,1]W \in [0,1], the convexity of ϕ\phi guarantees that pointwise, for every realization of (W,X,Y)(W, X, Y), ϕ(WX+(1W)Y)Wϕ(X)+(1W)ϕ(Y).\phi(WX + (1-W)Y) \leq W\phi(X) + (1-W)\phi(Y). This holds because convexity is a deterministic, sample-path property — it doesn't care about the joint distribution of WW, XX, and YY. Taking expectations of both sides of a pointwise inequality is always valid (by monotonicity of expectation), so Eϕ(WX+(1W)Y)E[Wϕ(X)+(1W)ϕ(Y)],\mathbb{E}\phi(WX+(1-W)Y) \leq \mathbb{E}[W\phi(X)+(1-W)\phi(Y)], which is exactly B — and it requires no independence assumption whatsoever. A is wrong because factoring E[Wϕ(X)]\mathbb{E}[W\phi(X)] into E[W]E[ϕ(X)]\mathbb{E}[W]\cdot\mathbb{E}[\phi(X)] requires WXW \perp X (and similarly WYW \perp Y), which fails when dependence is present. C reverses the inequality in B, contradicting Jensen's inequality for convex ϕ\phi. D applies Jensen in the wrong direction: for convex ϕ\phi, Jensen gives ϕ(E[Z])E[ϕ(Z)]\phi(\mathbb{E}[Z]) \leq \mathbb{E}[\phi(Z)], so ϕ(E[])\phi(\mathbb{E}[\cdot]) is the smaller side, not the larger one. Study tip: Always distinguish between pointwise inequalities (which survive taking expectations freely) and results that require independence to factor joint expectations. Dependence only breaks factorizations, not pointwise convexity arguments.

Question 5

Let probability measure PP be absolutely continuous with respect to QQ, with likelihood ratio L=dP/dQL=dP/dQ. For a statistic SS, let PSP_S and QSQ_S denote the induced distributions of SS. Assume the relevant divergences are finite. Which statement correctly applies conditional Jensen's inequality to xxlogxx\mapsto x\log x?

  1. D(PSQS)D(PQ)D(P_S\|Q_S)\leq D(P\|Q), with equality exactly when SS is independent of LL under QQ.
  2. D(PSQS)D(PQ)D(P_S\|Q_S)\geq D(P\|Q), with equality exactly when SS is sufficient under QQ alone.
  3. D(PSQS)=D(PQ)D(P_S\|Q_S)=D(P\|Q) for every statistic SS because likelihood ratios preserve expectation.
  4. D(PSQS)D(PQ)D(P_S\|Q_S)\leq D(P\|Q), with equality exactly when LL is measurable with respect to SS up to null sets. (correct answer)
Explanation: Whenever you see KL divergence and a statistic SS, think about the data-processing inequality: processing data can only destroy information, never create it. The rigorous derivation uses conditional Jensen's inequality applied to the convex function ϕ(x)=xlogx\phi(x) = x\log x. Here's the core argument. We have D(PQ)=EQ[LlogL]D(P\|Q) = E_Q[L \log L]. The induced likelihood ratio is LS=dPS/dQS=EQ[LS]L_S = dP_S/dQ_S = E_Q[L \mid S] (the conditional expectation of LL given SS, under QQ). Since ϕ(x)=xlogx\phi(x) = x\log x is convex, Jensen's inequality applied conditionally gives: EQ[ϕ(L)S]ϕ(EQ[LS])=ϕ(LS)E_Q[\phi(L) \mid S] \geq \phi(E_Q[L \mid S]) = \phi(L_S) Taking expectations over SS: D(PQ)=EQ[LlogL]EQS[LSlogLS]=D(PSQS)D(P\|Q) = E_Q[L\log L] \geq E_{Q_S}[L_S \log L_S] = D(P_S\|Q_S), confirming D(PSQS)D(PQ)D(P_S\|Q_S) \leq D(P\|Q). Equality in Jensen's holds if and only if LL is already σ(S)\sigma(S)-measurable (i.e., L=EQ[LS]L = E_Q[L\mid S] a.s.), meaning LL is measurable with respect to SS. This is exactly the definition of SS being sufficient for PP versus QQ. So D is correct. A gets the inequality direction right but misidentifies the equality condition. Independence of SS and LL under QQ would collapse LSL_S to a constant, a much stronger and incorrect condition. B flips the inequality entirely — divergence cannot increase under a statistic, period. C claims equality always holds, which contradicts the strict inequality whenever LL is not SS-measurable. As a study tip: whenever equality in Jensen's is discussed, always ask "when is the random variable already constant given the conditioning?" — that directly identifies sufficiency conditions.

Question 6

Let GH\mathcal{G}\subseteq\mathcal{H} be two sub-σ\sigma-fields, let XX be integrable, and let ϕ\phi be convex with all relevant expectations finite. Which ordering of expected convex transforms is guaranteed?

  1. Eϕ(E[XH])Eϕ(E[XG])ϕ(E[X])\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{H}])\leq\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{G}])\leq\phi(\mathbb{E}[X]), because less information yields a smaller expected convex transform.
  2. Eϕ(E[XG])Eϕ(E[XH])Eϕ(X)\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{G}])\geq\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{H}])\geq\mathbb{E}\phi(X), because conditioning on a coarser field inflates the convex transform.
  3. Eϕ(E[XG])Eϕ(E[XH])Eϕ(X)\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{G}])\leq\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{H}])\leq\mathbb{E}\phi(X), by two applications of Jensen's inequality and the tower property. (correct answer)
  4. ϕ(E[X])Eϕ(E[XG])Eϕ(E[XH])\phi(\mathbb{E}[X])\leq\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{G}])\leq\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{H}]), but no guarantee that either term is bounded above by Eϕ(X)\mathbb{E}\phi(X).
Explanation: Whenever you see conditional expectations composed with a convex function, your instinct should be to reach for Jensen's inequality combined with the tower property (also called the law of iterated expectations). Here's the core reasoning for why C is correct. Since GH\mathcal{G}\subseteq\mathcal{H}, the tower property tells you E[XG]=E[E[XH]G]\mathbb{E}[X\mid\mathcal{G}] = \mathbb{E}[\mathbb{E}[X\mid\mathcal{H}]\mid\mathcal{G}]. Now apply Jensen's inequality to the convex function ϕ\phi: conditioning on G\mathcal{G} and then applying ϕ\phi gives E[ϕ(E[XH])G]ϕ(E[E[XH]G])=ϕ(E[XG])\mathbb{E}[\phi(\mathbb{E}[X\mid\mathcal{H}])\mid\mathcal{G}] \geq \phi(\mathbb{E}[\mathbb{E}[X\mid\mathcal{H}]\mid\mathcal{G}]) = \phi(\mathbb{E}[X\mid\mathcal{G}]). Taking full expectations yields Eϕ(E[XH])Eϕ(E[XG])\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{H}])\geq\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{G}]). A second application of Jensen directly gives Eϕ(X)Eϕ(E[XH])\mathbb{E}\phi(X)\geq\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{H}]). Chaining these produces the inequality in C: more information → larger expected convex transform, bounded above by Eϕ(X)\mathbb{E}\phi(X). A reverses the correct direction by claiming less information yields a smaller expected convex transform — exactly backwards. B correctly identifies that Eϕ(E[XG])Eϕ(E[XH])\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{G}])\leq\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{H}]), but then wrongly claims both exceed Eϕ(X)\mathbb{E}\phi(X), flipping Jensen's inequality entirely. D gets the lower bound ϕ(E[X])Eϕ(E[XG])\phi(\mathbb{E}[X])\leq\mathbb{E}\phi(\mathbb{E}[X\mid\mathcal{G}]) right but wrongly claims no upper bound by Eϕ(X)\mathbb{E}\phi(X) exists — Jensen guarantees exactly that bound. Study tip: Memorize the mantra — conditioning contracts, ϕ\phi expands — meaning finer σ\sigma-fields push Eϕ(E[X])\mathbb{E}\phi(\mathbb{E}[X\mid\cdot]) upward toward Eϕ(X)\mathbb{E}\phi(X). Tower property + Jensen, applied twice, is the engine behind nearly every such ordering problem.

Question 7

A random variable XX is supported on [2,1][-2,-1], has mean E[X]=3/2\mathbb{E}[X]=-3/2, and has positive variance. What does Jensen's inequality imply about its third moment?

  1. E[X3]278\mathbb{E}[X^3]\geq -\frac{27}{8} because odd powers preserve the Jensen direction.
  2. E[X3]>278\mathbb{E}[X^3]> -\frac{27}{8} because the cubic function is strictly convex on the support.
  3. E[X3]<278\mathbb{E}[X^3]< -\frac{27}{8} because the cubic function is strictly concave on the support. (correct answer)
  4. No comparison with 27/8-27/8 is possible without knowing the full distribution of XX.
Explanation: Whenever you see Jensen's inequality applied to a specific support, your first task is to determine the local convexity or concavity of the function on that support — not globally. The cubic function f(x)=x3f(x) = x^3 has second derivative f(x)=6xf''(x) = 6x. On the support [2,1][-2, -1], we have x<0x < 0, so f(x)=6x<0f''(x) = 6x < 0 throughout. This means ff is strictly concave on [2,1][-2, -1]. Jensen's inequality for a strictly concave function states that E[f(X)]<f(E[X])\mathbb{E}[f(X)] < f(\mathbb{E}[X]) whenever XX has positive variance. Applying this: E[X3]<(E[X])3=(32)3=278\mathbb{E}[X^3] < \left(\mathbb{E}[X]\right)^3 = \left(-\frac{3}{2}\right)^3 = -\frac{27}{8} This confirms C is correct. Choice A is wrong on two counts: it claims the inequality goes the other way (\geq), and the phrase "odd powers preserve the Jensen direction" is meaningless — direction is determined by local concavity/convexity, not the parity of the exponent. Choice B makes the opposite convexity error. x3x^3 is strictly convex for x>0x > 0, but on [2,1][-2,-1] it is strictly concave. Confusing global behavior with local behavior is a classic trap. Choice D is tempting but wrong. Jensen's inequality requires only knowledge of E[X]\mathbb{E}[X], the sign of variance, and the local curvature — all of which are given. No further distributional information is needed. Study tip: Always evaluate f(x)f''(x) over the actual support before applying Jensen's — the sign of the second derivative on the support determines everything.

Question 8

Two predictive models assign strictly positive densities p1p_1 and p2p_2 to an outcome YY whose true distribution is QQ. For a fixed 0<λ<10<\lambda<1, define the mixture density pλ=λp1+(1λ)p2p_\lambda=\lambda p_1+(1-\lambda)p_2. Assume all expected log losses are finite. Which comparison follows from Jensen's inequality?

  1. EQ[logpλ(Y)]=λEQ[logp1(Y)]+(1λ)EQ[logp2(Y)]\mathbb{E}_Q[-\log p_\lambda(Y)]=\lambda\mathbb{E}_Q[-\log p_1(Y)]+(1-\lambda)\mathbb{E}_Q[-\log p_2(Y)]
  2. EQ[logpλ(Y)]λEQ[logp1(Y)]+(1λ)EQ[logp2(Y)]\mathbb{E}_Q[-\log p_\lambda(Y)]\geq\lambda\mathbb{E}_Q[-\log p_1(Y)]+(1-\lambda)\mathbb{E}_Q[-\log p_2(Y)]
  3. EQ[logpλ(Y)]log ⁣(λEQ[p1(Y)]+(1λ)EQ[p2(Y)])\mathbb{E}_Q[-\log p_\lambda(Y)]\leq-\log\!\left(\lambda\mathbb{E}_Q[p_1(Y)]+(1-\lambda)\mathbb{E}_Q[p_2(Y)]\right)
  4. EQ[logpλ(Y)]λEQ[logp1(Y)]+(1λ)EQ[logp2(Y)]\mathbb{E}_Q[-\log p_\lambda(Y)]\leq\lambda\mathbb{E}_Q[-\log p_1(Y)]+(1-\lambda)\mathbb{E}_Q[-\log p_2(Y)] (correct answer)
Explanation: Whenever you encounter a question involving a log of a mixture (or any concave function applied to a convex combination), Jensen's inequality is your primary tool. Recall that ff is concave if f(λa+(1λ)b)λf(a)+(1λ)f(b)f(\lambda a + (1-\lambda)b) \geq \lambda f(a) + (1-\lambda)f(b). Since log\log is concave, applying it to a mixture satisfies this direction. Here's the key move. Because pλ=λp1+(1λ)p2p_\lambda = \lambda p_1 + (1-\lambda)p_2, Jensen's inequality for the concave function log\log gives: log(λp1(y)+(1λ)p2(y))λlogp1(y)+(1λ)logp2(y).\log(\lambda p_1(y) + (1-\lambda)p_2(y)) \geq \lambda \log p_1(y) + (1-\lambda)\log p_2(y). Multiplying through by 1-1 flips the inequality (since log-\log is convex): logpλ(y)λlogp1(y)(1λ)logp2(y).-\log p_\lambda(y) \leq -\lambda \log p_1(y) - (1-\lambda)\log p_2(y). Taking expectations under QQ preserves the direction, yielding exactly D: the expected log loss of the mixture is no greater than the convex combination of the individual log losses. This is the key log-pooling result — mixing predictions never hurts in terms of worst-case log loss. A claims equality, which only holds when log\log is linear — it never is, so A is wrong. B reverses the inequality, contradicting Jensen's direction for a concave function. C applies Jensen in the wrong place — it bounds EQ[logpλ]\mathbb{E}_Q[-\log p_\lambda] using log-\log of an expectation of pp, which conflates two separate applications of Jensen and yields a weaker, less relevant bound. As a study habit: always identify whether your function is convex or concave first, then determine which direction Jensen pulls the inequality — getting that direction right is the entire question.

Question 9

A positive random variable XX is supported on [1,3][1,3] and satisfies E[X]=2\mathbb{E}[X]=2 and Var(X)=1/4\operatorname{Var}(X)=1/4. Using a variance-refined form of Jensen's inequality and the curvature of f(x)=logxf(x)=-\log x, which lower bound is valid?

  1. E[logX]log2+172\mathbb{E}[-\log X]\geq -\log 2+\frac{1}{72} (correct answer)
  2. E[logX]log2+136\mathbb{E}[-\log X]\geq -\log 2+\frac{1}{36}
  3. E[logX]log2\mathbb{E}[-\log X]\geq -\log 2, but no variance-dependent improvement follows.
  4. E[logX]log2172\mathbb{E}[-\log X]\geq -\log 2-\frac{1}{72}
Explanation: When you see a question combining Jensen's inequality with variance information, think "variance-refined Jensen's bound." The standard Jensen's inequality gives E[f(X)]f(E[X])\mathbb{E}[f(X)] \geq f(\mathbb{E}[X]) for convex ff, but a tighter version incorporates curvature and variance. The variance-refined lower bound states: for a convex function ff, E[f(X)]f(μ)+12f(ξ)Var(X)\mathbb{E}[f(X)] \geq f(\mu) + \frac{1}{2}f''(\xi)\operatorname{Var}(X) where ξ\xi lies in the support. For f(x)=logxf(x) = -\log x, compute f(x)=1/x2f''(x) = 1/x^2. To find a valid lower bound, you want the minimum curvature over the support [1,3][1,3], which occurs at the largest xx: f(3)=1/9f''(3) = 1/9. Using μ=2\mu = 2 and Var(X)=1/4\operatorname{Var}(X) = 1/4: E[logX]log2+121914=log2+172\mathbb{E}[-\log X] \geq -\log 2 + \frac{1}{2} \cdot \frac{1}{9} \cdot \frac{1}{4} = -\log 2 + \frac{1}{72} This confirms answer A. Answer B uses f(2)=1/4f''(2) = 1/4 (evaluating curvature at the mean rather than at the support endpoint), giving 121414=132\frac{1}{2} \cdot \frac{1}{4} \cdot \frac{1}{4} = \frac{1}{32}, which is not 136\frac{1}{36} — and more importantly, using the mean's curvature doesn't guarantee a valid lower bound without additional argument. Answer C ignores that convexity plus variance information always yields a strict improvement when Var(X)>0\operatorname{Var}(X) > 0. Answer D is negative, but since logx-\log x is convex, the variance correction must push the bound upward, not downward. Your strategy: for variance-refined Jensen bounds, always use the minimum value of ff'' over the support to ensure the inequality direction is preserved — this is where most errors occur.

Question 10

An integrable random variable XX satisfies E[X]=1\mathbb{E}[X]=-1 and EX=1\mathbb{E}|X|=1. Which conclusion is forced by the equality case of Jensen's inequality for xxx\mapsto |x|?

  1. X=1X=-1 almost surely, so the distribution must be degenerate.
  2. X0X\leq 0 almost surely, although its variance need not be zero. (correct answer)
  3. P(X<0)=P(X>0)\mathbb{P}(X<0)=\mathbb{P}(X>0), with possible mass at zero.
  4. Var(X)=0\operatorname{Var}(X)=0, but the almost-sure value is otherwise undetermined.
Explanation: When you see conditions on both E[X]\mathbb{E}[X] and EX\mathbb{E}|X|, immediately think about Jensen's inequality applied to the convex function φ(x)=x\varphi(x) = |x|. Jensen's inequality states E[X]EX|\mathbb{E}[X]| \leq \mathbb{E}|X|, with equality holding if and only if XX is almost surely confined to a region where φ\varphi is linear — that is, where x|x| does not "bend." Here, E[X]=1=1=EX|\mathbb{E}[X]| = |-1| = 1 = \mathbb{E}|X|, so equality holds exactly. The function x|x| is linear on (,0](-\infty, 0] (where it equals x-x) and on [0,)[0, \infty) (where it equals xx). Equality in Jensen's is achieved when XX stays within one of these linear branches almost surely — meaning X0X \leq 0 a.s. or X0X \geq 0 a.s. Since E[X]=1<0\mathbb{E}[X] = -1 < 0, the branch X0X \geq 0 a.s. is ruled out (it would force a non-negative mean). Therefore, X0X \leq 0 almost surely. This is answer B. Notice that XX could still be spread across many negative values, so its variance is not forced to zero. Answer A is too strong: X=1X = -1 a.s. would additionally require Var(X)=0\operatorname{Var}(X) = 0, which is not implied. Answer C is a fabrication — nothing about the equality condition requires symmetry around zero or equal probabilities on each side. Answer D correctly identifies zero variance as sufficient but not necessary; the equality condition constrains the sign of XX, not its spread. Your study tip: memorize that equality in Jensen's for x|x| forces the random variable onto a single linear branch of the absolute value — this pins down the sign, not the exact value.