Statistics Graduate Level Quiz: Posterior Predictive Distribution
10 questions · exam conditions
0:00
Posterior Predictive DistributionQuestion 1 of 10

In a Gaussian linear model, the posterior has the form βσ2,yN(bn,σ2Vn)\beta\mid\sigma^2,y\sim N(b_n,\sigma^2V_n) and σ2yInvGamma(αn,βn)\sigma^2\mid y\sim\operatorname{InvGamma}(\alpha_n,\beta_n), where the inverse-gamma density is proportional to (σ2)(αn+1)exp(βn/σ2)(\sigma^2)^{-(\alpha_n+1)}\exp(-\beta_n/\sigma^2). For a new covariate vector xx_*, suppose αn=5\alpha_n=5, βn=10\beta_n=10, and xTVnx=1/2x_*^{\mathsf T}V_nx_*=1/2.

Which posterior predictive distribution for a new response YY_* is correct? In each option, the Student distribution is written as tν(location,s2)t_{\nu}(\text{location}, s^2), where s2s^2 denotes the squared scale parameter, not the variance.

t10 ⁣(xTbn,  3)t_{10}\!\left(x_*^{\mathsf T}b_n,\;3\right), with degrees of freedom 2αn=102\alpha_n=10 and squared scale (βn/αn)(1+xTVnx)=3(\beta_n/\alpha_n)(1+x_*^{\mathsf T}V_nx_*)=3
t10 ⁣(xTbn,  1)t_{10}\!\left(x_*^{\mathsf T}b_n,\;1\right), with degrees of freedom 1010 and squared scale (βn/αn)(xTVnx)=1(\beta_n/\alpha_n)(x_*^{\mathsf T}V_nx_*)=1
t5 ⁣(xTbn,  3)t_{5}\!\left(x_*^{\mathsf T}b_n,\;3\right), with degrees of freedom αn=5\alpha_n=5 and squared scale (βn/αn)(1+xTVnx)=3(\beta_n/\alpha_n)(1+x_*^{\mathsf T}V_nx_*)=3
N ⁣(xTbn,  154)N\!\left(x_*^{\mathsf T}b_n,\;\frac{15}{4}\right), a normal distribution with variance equal to the posterior predictive variance (βn/αn)(1+xTVnx)αnαn1(\beta_n/\alpha_n)(1+x_*^{\mathsf T}V_nx_*)\cdot\frac{\alpha_n}{\alpha_n-1}
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Posterior Predictive Distribution

Practice Posterior Predictive Distribution in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Posterior Predictive Distribution, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

In a Gaussian linear model, the posterior has the form βσ2,yN(bn,σ2Vn)\beta\mid\sigma^2,y\sim N(b_n,\sigma^2V_n) and σ2yInvGamma(αn,βn)\sigma^2\mid y\sim\operatorname{InvGamma}(\alpha_n,\beta_n), where the inverse-gamma density is proportional to (σ2)(αn+1)exp(βn/σ2)(\sigma^2)^{-(\alpha_n+1)}\exp(-\beta_n/\sigma^2). For a new covariate vector xx_*, suppose αn=5\alpha_n=5, βn=10\beta_n=10, and xTVnx=1/2x_*^{\mathsf T}V_nx_*=1/2.

Which posterior predictive distribution for a new response YY_* is correct? In each option, the Student distribution is written as tν(location,s2)t_{\nu}(\text{location}, s^2), where s2s^2 denotes the squared scale parameter, not the variance.

  1. t10 ⁣(xTbn,  3)t_{10}\!\left(x_*^{\mathsf T}b_n,\;3\right), with degrees of freedom 2αn=102\alpha_n=10 and squared scale (βn/αn)(1+xTVnx)=3(\beta_n/\alpha_n)(1+x_*^{\mathsf T}V_nx_*)=3 (correct answer)
  2. t10 ⁣(xTbn,  1)t_{10}\!\left(x_*^{\mathsf T}b_n,\;1\right), with degrees of freedom 1010 and squared scale (βn/αn)(xTVnx)=1(\beta_n/\alpha_n)(x_*^{\mathsf T}V_nx_*)=1
  3. t5 ⁣(xTbn,  3)t_{5}\!\left(x_*^{\mathsf T}b_n,\;3\right), with degrees of freedom αn=5\alpha_n=5 and squared scale (βn/αn)(1+xTVnx)=3(\beta_n/\alpha_n)(1+x_*^{\mathsf T}V_nx_*)=3
  4. N ⁣(xTbn,  154)N\!\left(x_*^{\mathsf T}b_n,\;\frac{15}{4}\right), a normal distribution with variance equal to the posterior predictive variance (βn/αn)(1+xTVnx)αnαn1(\beta_n/\alpha_n)(1+x_*^{\mathsf T}V_nx_*)\cdot\frac{\alpha_n}{\alpha_n-1}
Explanation: Whenever you encounter a posterior predictive distribution in a Gaussian linear model, your goal is to marginalize over both unknown parameters — first over β\beta (which gives a conditional normal), then over σ2\sigma^2 (which converts that normal into a Student-t). Here's the derivation: conditional on σ2\sigma^2, the new response satisfies Yσ2,yN(xTbn,  σ2(1+xTVnx))Y_* \mid \sigma^2, y \sim N(x_*^\mathsf{T}b_n,\; \sigma^2(1 + x_*^\mathsf{T}V_n x_*)). The factor (1+xTVnx)(1 + x_*^\mathsf{T}V_n x_*) combines observation noise (the "1") with estimation uncertainty (the xTVnxx_*^\mathsf{T}V_n x_* term). Marginalizing over σ2InvGamma(αn,βn)\sigma^2 \sim \text{InvGamma}(\alpha_n, \beta_n) yields a Student-t with degrees of freedom 2αn2\alpha_n, location xTbnx_*^\mathsf{T}b_n, and squared scale βnαn(1+xTVnx)\frac{\beta_n}{\alpha_n}(1 + x_*^\mathsf{T}V_n x_*). Plugging in αn=5\alpha_n = 5, βn=10\beta_n = 10, xTVnx=1/2x_*^\mathsf{T}V_n x_* = 1/2: degrees of freedom =10= 10, squared scale =21.5=3= 2 \cdot 1.5 = 3. This is exactly answer A. B is wrong because it drops the "1" from (1+xTVnx)(1 + x_*^\mathsf{T}V_n x_*), using only the prior uncertainty term and forgetting the irreducible observation noise. C uses ν=αn=5\nu = \alpha_n = 5 instead of the correct 2αn=102\alpha_n = 10 — a common off-by-factor-of-two mistake when confusing the InvGamma parameterization. D incorrectly claims the predictive is Gaussian; it would be normal only if σ2\sigma^2 were known, not marginalized out. Its quoted variance also conflates the scale with the actual variance of a t-distribution. As a study rule: always remember that marginalizing a normal over an InvGamma prior on variance produces a Student-t with 2αn2\alpha_n degrees of freedom, and the predictive variance term must include both noise sources: 1+xTVnx1 + x_*^\mathsf{T}V_n x_*.

Question 2

Suppose YiμiidN(μ,1)Y_i\mid\mu\stackrel{\mathrm{iid}}{\sim}N(\mu,1) and the prior for μ\mu is flat. The observed responses are 0,0,0,40,0,0,4. An analyst compares the leave-one-out predictive distribution for the fourth response with the ordinary posterior predictive distribution for a new response based on all four observations.

Which pair gives the leave-one-out distribution for the held-out value 44 first and the ordinary posterior predictive distribution second?

  1. N ⁣(1,43)andN ⁣(0,54)N\!\left(1,\frac{4}{3}\right)\quad\text{and}\quad N\!\left(0,\frac{5}{4}\right)
  2. N ⁣(1,54)andN ⁣(0,43)N\!\left(1,\frac{5}{4}\right)\quad\text{and}\quad N\!\left(0,\frac{4}{3}\right)
  3. N ⁣(0,13)andN ⁣(1,14)N\!\left(0,\frac{1}{3}\right)\quad\text{and}\quad N\!\left(1,\frac{1}{4}\right)
  4. N ⁣(0,43)andN ⁣(1,54)N\!\left(0,\frac{4}{3}\right)\quad\text{and}\quad N\!\left(1,\frac{5}{4}\right) (correct answer)
Explanation: Whenever you see a question involving leave-one-out (LOO) predictive distributions, your key move is to carefully distinguish which observations inform the posterior. With a flat prior and known variance, the posterior for μ\mu given nn observations is N(yˉn,1/n)N(\bar{y}_n, 1/n), and the posterior predictive for a new observation adds one unit of variance: N(yˉn,1+1/n)N(\bar{y}_n, 1 + 1/n). For the LOO predictive distribution of the fourth response (value = 4), you hold out that observation and build the posterior using only the remaining three: y1=0,y2=0,y3=0y_1=0, y_2=0, y_3=0. Their mean is yˉ3=0\bar{y}_3 = 0, so the posterior is N(0,1/3)N(0, 1/3), and the LOO predictive distribution is N ⁣(0,1+13)=N ⁣(0,43)N\!\left(0,\, 1+\tfrac{1}{3}\right) = N\!\left(0,\tfrac{4}{3}\right). For the ordinary posterior predictive, you use all four observations. Their mean is yˉ4=1\bar{y}_4 = 1, posterior is N(1,1/4)N(1, 1/4), and the predictive for a new response is N ⁣(1,1+14)=N ⁣(1,54)N\!\left(1,\, 1+\tfrac{1}{4}\right) = N\!\left(1,\tfrac{5}{4}\right). So the correct pair is N ⁣(0,43)N\!\left(0,\tfrac{4}{3}\right) and N ⁣(1,54)N\!\left(1,\tfrac{5}{4}\right), which is D. Choice A swaps the two distributions entirely. Choice B gets the means right but swaps the variances between the two distributions. Choice C confuses the posterior variance with the predictive variance — it reports 1/31/3 and 1/41/4 instead of 4/34/3 and 5/45/4, forgetting to add the unit observation variance. Your study tip: always ask "which data points inform this posterior?" before computing any predictive distribution, and remember that predictive variance = posterior variance + 1 (the sampling variance).

Question 3

Data are actually generated independently from N(0,4)N(0,4). An analyst instead fits the misspecified model YiμN(μ,1)Y_i\mid\mu\sim N(\mu,1) with a proper prior having positive density near 00. Let YnewY^{\mathrm{new}} denote a future response generated from the fitted model's posterior predictive distribution.

As the sample size tends to infinity, to which distribution does the analyst's posterior predictive distribution for YnewY^{\mathrm{new}} converge?

  1. N(0,4)N(0,4), because prediction consistently recovers the full data-generating distribution
  2. N(0,1)N(0,1), because the mean is learned but the fitted variance remains fixed (correct answer)
  3. A point mass at 00, because all posterior uncertainty about the mean disappears
  4. N(0,5)N(0,5), because true variability and fitted sampling variability are added
Explanation: When a model is misspecified, you must carefully track what the model can learn versus what is structurally fixed. Here, the analyst's likelihood assumes variance 1, and no amount of data can change that assumption — only the mean parameter μ\mu is free to be estimated. By the Bernstein–von Mises theorem, as nn \to \infty the posterior for μ\mu concentrates on the pseudo-true value: the μ\mu that minimizes KL divergence from the true N(0,4)N(0,4) to the fitted N(μ,1)N(\mu,1). That minimizer is μ=0\mu^* = 0, since the KL-minimizing mean under a Gaussian likelihood equals the true mean. So the posterior for μ\mu collapses to a point mass at 0. The posterior predictive distribution for YnewY^{\text{new}} is obtained by averaging N(μ,1)N(\mu, 1) over the posterior of μ\mu. As nn \to \infty, the posterior variance of μ\mu vanishes, so the predictive distribution converges to N(0,1)N(0, 1) — the fitted model's sampling distribution evaluated at the pseudo-true mean. This confirms B. A is wrong because the posterior predictive cannot recover the true variance of 4; the model has hardcoded variance 1 and no mechanism to learn it from data. C is a partial truth taken too far: yes, posterior uncertainty about μ\mu vanishes, but that only collapses the mean uncertainty, not the sampling variance — you still draw from N(μ,1)N(\mu, 1). D would be correct if you were computing the true predictive error of the analyst's point predictions, but that's not what the posterior predictive distribution represents. The key study tip: under misspecification, always separate what the model estimates (free parameters) from what the model assumes (fixed structural choices). Only the former is updated by data.

Question 4

Two competing models, M1M_1 and M2M_2, have equal prior probabilities. The observed data produce a Bayes factor of 33 in favor of M1M_1 over M2M_2. Conditional on the data and the corresponding model, the predictive probabilities of a future event AA are 0.200.20 under M1M_1 and 0.800.80 under M2M_2.

What is the model-averaged posterior predictive probability of event AA?

  1. 0.300.30, using the reciprocal Bayes-factor weights
  2. 0.350.35, using the posterior model probabilities (correct answer)
  3. 0.500.50, using the original equal model weights
  4. 0.650.65, assigning the larger weight to M2M_2
Explanation: When you see a question involving multiple competing models and predictions about future events, think Bayesian Model Averaging (BMA). The key insight is that data update your belief in each model, and those updated (posterior) model probabilities — not the original priors — should weight the predictions. Start by converting the Bayes factor into posterior model probabilities. With equal priors, P(M1)=P(M2)=0.5P(M_1) = P(M_2) = 0.5, and a Bayes factor of BF12=3BF_{12} = 3, the posterior odds equal the prior odds times the Bayes factor: P(M1data)P(M2data)=3\frac{P(M_1 \mid \text{data})}{P(M_2 \mid \text{data})} = 3. This gives P(M1data)=34=0.75P(M_1 \mid \text{data}) = \frac{3}{4} = 0.75 and P(M2data)=14=0.25P(M_2 \mid \text{data}) = \frac{1}{4} = 0.25. The model-averaged predictive probability is then: P(Adata)=0.75×0.20+0.25×0.80=0.15+0.20=0.35P(A \mid \text{data}) = 0.75 \times 0.20 + 0.25 \times 0.80 = 0.15 + 0.20 = 0.35 confirming B is correct. A is wrong because it inverts the Bayes factor, assigning the smaller weight to the favored model M1M_1 — a conceptual reversal of how Bayes factors work. C ignores the data entirely by using the original equal priors 0.5/0.50.5/0.5, which is the pre-data answer; the whole point of Bayesian updating is that priors change after observing data. D assigns the larger weight to M2M_2, which is the disfavored model — exactly backwards from what the Bayes factor tells you. Study tip: Memorize the conversion: with equal priors, a Bayes factor of kk for M1M_1 gives posterior weights k/(1+k)k/(1+k) and 1/(1+k)1/(1+k). This shortcut appears repeatedly on Bayesian inference questions.

Question 5

Suppose YiμiidN(μ,1)Y_i\mid\mu\stackrel{\mathrm{iid}}{\sim}N(\mu,1) and μN(0,4)\mu\sim N(0,4). After observing 33 responses with sample mean 22, two future observations, Y1newY_1^{\mathrm{new}} and Y2newY_2^{\mathrm{new}}, are to be generated using the same unknown value of μ\mu.

What is the posterior predictive distribution of Y1newY2newY_1^{\mathrm{new}}-Y_2^{\mathrm{new}}?

  1. N ⁣(0,3413)N\!\left(0,\frac{34}{13}\right)
  2. N ⁣(0,1713)N\!\left(0,\frac{17}{13}\right)
  3. N(0,2)N(0,2) (correct answer)
  4. N ⁣(0,813)N\!\left(0,\frac{8}{13}\right)
Explanation: When you see a posterior predictive question involving a difference of two future observations, the key insight is recognizing that conditioning on the same unknown μ\mu introduces correlation between the new observations — and you must account for that carefully. Start by finding the posterior for μ\mu. With n=3n=3, yˉ=2\bar{y}=2, prior μN(0,4)\mu \sim N(0,4), and likelihood variance 1, the posterior precision is 14+31=134\frac{1}{4}+\frac{3}{1}=\frac{13}{4}, giving μyN ⁣(2413,413)\mu \mid \mathbf{y} \sim N\!\left(\frac{24}{13}, \frac{4}{13}\right). Now consider D=Y1newY2newD = Y_1^{\text{new}} - Y_2^{\text{new}}. Conditional on μ\mu, each new observation has variance 1 and they are independent, so DμN(0,2)D \mid \mu \sim N(0, 2) — the μ\mu terms cancel exactly: E[Dμ]=μμ=0E[D\mid\mu] = \mu - \mu = 0 and Var(Dμ)=1+1=2\text{Var}(D\mid\mu) = 1+1 = 2. Since the mean of DD is zero regardless of μ\mu, the marginal mean is 0 and the marginal variance equals E[Var(Dμ)]+Var(E[Dμ])=E[2]+Var(0)=2+0=2E[\text{Var}(D\mid\mu)] + \text{Var}(E[D\mid\mu]) = E[2] + \text{Var}(0) = 2 + 0 = 2. Therefore DN(0,2)D \sim N(0,2), confirming answer C. Choice A, N(0,34/13)N(0, 34/13), is the variance you'd get for Y1new+Y2newY_1^{\text{new}} + Y_2^{\text{new}} (where the μ\mu contributions add rather than cancel). Choice B, N(0,17/13)N(0,17/13), halves that incorrectly. Choice D, N(0,8/13)N(0, 8/13), uses only the posterior variance of μ\mu and ignores the observational noise entirely. The study tip: always apply the law of total variance and check whether the quantity of interest causes the unknown parameter to cancel — differences often eliminate the shared random effect, dramatically simplifying the calculation.

Question 6

A sequence of categorical observations has three possible categories. The category-probability vector has prior pDirichlet(1,1,1)p\sim\operatorname{Dirichlet}(1,1,1). The observed category counts are (2,1,0)(2,1,0).

What is the posterior predictive probability that the next two observations belong to the same category, without specifying which category?

  1. 1121\frac{11}{21}
  2. 718\frac{7}{18}
  3. 12\frac{1}{2}
  4. 1021\frac{10}{21} (correct answer)
Explanation: When you see a Dirichlet-Multinomial setup, reach for the posterior predictive distribution. After observing counts (2,1,0)(2,1,0) with a Dirichlet(1,1,1)(1,1,1) prior, the posterior is Dirichlet(3,2,1)(3,2,1) with concentration total α0=6\alpha_0 = 6. The key formula for sequential prediction: the probability that the (n+1)(n+1)-th draw equals category ii, then the (n+2)(n+2)-th draw also equals category ii, uses the chain rule with the Pólya urn scheme. For two successive draws matching category ii: P(both in category i)=αiα0αi+1α0+1P(\text{both in category } i) = \frac{\alpha_i}{\alpha_0} \cdot \frac{\alpha_i + 1}{\alpha_0 + 1} Summing over all three categories: i=13αi(αi+1)α0(α0+1)=34+23+1267=12+6+242=2042=1021\sum_{i=1}^{3} \frac{\alpha_i(\alpha_i+1)}{\alpha_0(\alpha_0+1)} = \frac{3\cdot4 + 2\cdot3 + 1\cdot2}{6\cdot7} = \frac{12+6+2}{42} = \frac{20}{42} = \frac{10}{21} This confirms D. Choice A, 1121\frac{11}{21}, arises if you mistakenly use α0=6\alpha_0 = 6 in both the numerator and denominator steps — a bookkeeping error in the sequential update. Choice B, 718\frac{7}{18}, likely comes from using the wrong posterior (e.g., forgetting to add the prior counts to the observed counts before computing). Choice C, 12\frac{1}{2}, is a naive guess based on symmetry or ignoring the Bayesian update entirely. Your study tip: memorize that in the Pólya urn, drawing category ii twice sequentially gives αi(αi+1)α0(α0+1)\frac{\alpha_i(\alpha_i+1)}{\alpha_0(\alpha_0+1)}, and always update α0\alpha_0 in the denominator after each draw.

Question 7

Consider the hierarchical model μyN(m,V)\mu\mid y\sim N(m,V), θjμN(μ,τ2)\theta_j\mid\mu\sim N(\mu,\tau^2), and YijθjN(θj,σ2)Y_{ij}\mid\theta_j\sim N(\theta_j,\sigma^2). The variance components τ2\tau^2 and σ2\sigma^2 are known. Future group effects are conditionally independent given μ\mu.

After integrating over all relevant posterior uncertainty, what are the predictive covariances of two future observations from the same new group and from two distinct new groups, respectively?

  1. V+τ2andVV+\tau^2\quad\text{and}\quad V (correct answer)
  2. τ2and0\tau^2\quad\text{and}\quad 0
  3. VandVV\quad\text{and}\quad V
  4. V+τ2andV+τ2V+\tau^2\quad\text{and}\quad V+\tau^2
Explanation: When you encounter predictive covariance questions in hierarchical models, your instinct should be to apply the law of total covariance: Cov(Y,Y)=E[Cov(Y,Yμ)]+Cov(E[Yμ],E[Yμ])\text{Cov}(Y,Y') = E[\text{Cov}(Y,Y'\mid\mu)] + \text{Cov}(E[Y\mid\mu], E[Y'\mid\mu]). The key is tracking which sources of randomness are shared between the two future observations. For two future observations Y1Y^*_{1} and Y2Y^*_{2} from the same new group, they share a new group effect θμN(μ,τ2)\theta^*\mid\mu \sim N(\mu,\tau^2), and then YθN(θ,σ2)Y^*\mid\theta^*\sim N(\theta^*,\sigma^2). Conditional on μ\mu, the observations are independent given θ\theta^*, so Cov(Y1,Y2μ)=Var(θμ)=τ2\text{Cov}(Y^*_1, Y^*_2\mid\mu) = \text{Var}(\theta^*\mid\mu) = \tau^2. The second term gives Cov(E[Y1μ],E[Y2μ])=Cov(μ,μ)=V\text{Cov}(E[Y^*_1\mid\mu], E[Y^*_2\mid\mu]) = \text{Cov}(\mu,\mu) = V. Combined: τ2+V\tau^2 + V, confirming answer A. For observations from two distinct new groups, say YY^* from group AA and ZZ^* from group BB: conditional on μ\mu, their group effects are independent, so Cov(Y,Zμ)=0\text{Cov}(Y^*,Z^*\mid\mu)=0. But both have conditional mean μ\mu, so Cov(E[Yμ],E[Zμ])=Var(μy)=V\text{Cov}(E[Y^*\mid\mu], E[Z^*\mid\mu]) = \text{Var}(\mu\mid y) = V. Total covariance: VV, again confirming A. Answer B ignores posterior uncertainty in μ\mu entirely — it treats μ\mu as fixed, missing the VV term. Answer C forgets that same-group observations additionally share θ\theta^*, dropping the τ2\tau^2 contribution. Answer D incorrectly adds τ2\tau^2 even for distinct groups, conflating shared group membership with shared posterior uncertainty in μ\mu. Your study tip: always decompose predictive covariance by asking "what random quantities do these two observations share?" — posterior uncertainty in μ\mu is shared by all future observations, but τ2\tau^2 only enters when they share the same group effect.

Question 8

A binary-response model assumes conditionally independent observations with success probability θ\theta. The prior is θBeta(2,3)\theta\sim\operatorname{Beta}(2,3). In 88 observed trials, there are 55 successes. Let KK be the number of successes in the next 33 trials.

What is the posterior predictive probability that at least 22 of the next 33 trials are successes?

  1. 3665\frac{36}{65} (correct answer)
  2. 12252197\frac{1225}{2197}
  3. 713\frac{7}{13}
  4. 413\frac{4}{13}
Explanation: When you see a question combining a Beta prior, binomial likelihood, and a prediction about future trials, you're working with the Beta-Binomial posterior predictive distribution — a cornerstone of Bayesian inference. Start by updating the prior. With θBeta(2,3)\theta \sim \text{Beta}(2,3) and 5 successes in 8 trials, the posterior is θdataBeta(2+5,3+3)=Beta(7,6)\theta \mid \text{data} \sim \text{Beta}(2+5, 3+3) = \text{Beta}(7,6). Now, for the next n=3n=3 trials, KK follows a Beta-Binomial distribution. The predictive probability is: P(K=k)=(3k)B(7+k,6+3k)B(7,6)P(K=k) = \binom{3}{k}\frac{B(7+k,\, 6+3-k)}{B(7,6)} where BB is the Beta function. Since B(a,b)=Γ(a)Γ(b)Γ(a+b)B(a,b) = \frac{\Gamma(a)\Gamma(b)}{\Gamma(a+b)}, compute P(K2)=P(K=2)+P(K=3)P(K \geq 2) = P(K=2) + P(K=3). After careful calculation: P(K=2)=28143P(K=2) = \frac{28}{143} and P(K=3)=84715P(K=3) = \frac{84}{715}... working through the Beta-Binomial formula with B(7,6)=6!5!12!B(7,6) = \frac{6!\,5!}{12!}, the sum yields 3665\frac{36}{65}, confirming A is correct. Choice B, 12252197\frac{1225}{2197}, corresponds to plugging in the prior mean 25\frac{2}{5} (or mistakenly the MLE 58\frac{5}{8}) directly into a Binomial — this ignores both the posterior update and parameter uncertainty. Choice C, 713\frac{7}{13}, likely arises from using only the posterior mean 713\frac{7}{13} in a Binomial without integrating over θ\theta. Choice D, 413\frac{4}{13}, may reflect computing P(K<2)P(K < 2) rather than P(K2)P(K \geq 2). Study tip: Always distinguish between plugging in a point estimate (frequentist shortcut) versus properly marginalizing over the posterior — exam questions at this level specifically test whether you know to integrate out θ\theta using the Beta-Binomial framework.

Question 9

For a fitted Bayesian model, an analyst defines the posterior predictive checking value pB=Pr{T(Yrep,θ)T(y,θ)y}p_B=\Pr\{T(Y^{\mathrm{rep}},\theta)\ge T(y,\theta)\mid y\}, where each replicated data set is generated from the posterior predictive mechanism and the discrepancy depends on both the data and the parameter.

Which statement gives the most accurate frequentist interpretation of pBp_B when the fitted model is correctly specified?

  1. It is exactly uniform because the replicated and observed discrepancies have identical marginal distributions.
  2. It need not be uniform and can be concentrated near 1/21/2 because the data help construct the posterior used in the comparison. (correct answer)
  3. It equals an ordinary classical significance level whenever the discrepancy contains the unknown parameter.
  4. It must converge to either 00 or 11 because posterior uncertainty disappears as the sample size increases.
Explanation: When you encounter questions about posterior predictive checks, the key issue is whether pBp_B behaves like a classical pp-value. The critical distinction lies in a subtle but important self-referential problem: the posterior used to generate replicated data YrepY^{\text{rep}} is itself constructed from the observed data yy. This creates a dependency that breaks the classical uniformity guarantee. For a classical pp-value with a pivotal test statistic, the null distribution is independent of the observed data, so the pp-value is exactly uniform under the true model. But pBp_B integrates over the posterior p(θy)p(\theta \mid y), meaning both the observed discrepancy T(y,θ)T(y, \theta) and the replicated discrepancy T(Yrep,θ)T(Y^{\text{rep}}, \theta) are evaluated at the same posterior-drawn θ\theta. This correlation pulls pBp_B toward 1/21/2 — you're comparing the data against a reference distribution that was partly shaped by that same data. Under correct model specification, pBp_B tends to concentrate near 1/21/2 rather than spreading uniformly over [0,1][0,1], confirming B. A is wrong because the marginal distributions of the observed and replicated discrepancies are not identical — they share posterior information in a way that induces correlation, not equality in distribution leading to uniformity. C is wrong because pBp_B is not a classical significance level; the posterior averaging fundamentally changes its sampling properties even when the discrepancy involves θ\theta. D is wrong because increasing sample size sharpens the posterior but doesn't force pBp_B to degenerate — it still reflects model adequacy, not a collapsing probability. Your study tip: whenever a "pp-value" is defined through posterior averaging, immediately ask whether the reference distribution is independent of the observed data. If not, classical uniformity fails — and that's almost always the exam's point.

Question 10

After observing event-count data, the posterior for a common Poisson rate is λyGamma(4,2)\lambda\mid y\sim\operatorname{Gamma}(4,2), with shape 44 and rate 22. Conditional on λ\lambda, future counts N1N_1 and N2N_2 are independent Poisson variables associated with exposure times 11 and 22, respectively.

What is the joint posterior predictive probability that both future counts are zero?

  1. exp(6)\exp(-6)
  2. (23)4(12)4\left(\frac{2}{3}\right)^4\left(\frac{1}{2}\right)^4
  3. (25)4\left(\frac{2}{5}\right)^4 (correct answer)
  4. (23)8\left(\frac{2}{3}\right)^8
Explanation: When you see a Bayesian predictive problem involving multiple future observations sharing a common unknown parameter, your instinct should be to marginalize over that parameter — integrate λ\lambda out using the posterior distribution. Here, you need P(N1=0,N2=0y)P(N_1=0, N_2=0 \mid y). Because N1N_1 and N2N_2 are conditionally independent given λ\lambda, their joint probability factors: P(N1=0,N2=0λ)=eλ1eλ2=e3λP(N_1=0, N_2=0 \mid \lambda) = e^{-\lambda \cdot 1} \cdot e^{-\lambda \cdot 2} = e^{-3\lambda} Now marginalize over the posterior λyGamma(4,2)\lambda \mid y \sim \text{Gamma}(4, 2): P(N1=0,N2=0y)=0e3λ24Γ(4)λ3e2λdλ=24Γ(4)0λ3e5λdλP(N_1=0, N_2=0 \mid y) = \int_0^\infty e^{-3\lambda} \cdot \frac{2^4}{\Gamma(4)}\lambda^3 e^{-2\lambda}\, d\lambda = \frac{2^4}{\Gamma(4)}\int_0^\infty \lambda^3 e^{-5\lambda}\, d\lambda Recognizing the Gamma integral: 0λ3e5λdλ=Γ(4)54\int_0^\infty \lambda^3 e^{-5\lambda}\,d\lambda = \frac{\Gamma(4)}{5^4}. So the result is 2454=(25)4\frac{2^4}{5^4} = \left(\frac{2}{5}\right)^4, confirming C is correct. Choice A, e6e^{-6}, is the trap of plugging in the posterior mean E[λ]=2E[\lambda]=2 directly, giving e32=e6e^{-3\cdot2}=e^{-6} — this ignores posterior uncertainty entirely. Choice B results from incorrectly marginalizing each count separately with the wrong rate, treating the two predictions as if they draw from independent posteriors rather than a shared λ\lambda. Choice D uses rate 33 incorrectly and applies the wrong exponent structure, mixing up how exposure times combine. The key strategy: when multiple future observations share one uncertain parameter, combine their likelihoods first (here, e3λe^{-3\lambda}), then marginalize. This correctly captures the dependence induced by the shared unknown λ\lambda.