Statistics Graduate Level Quiz: Exponential Families
10 questions · exam conditions
0:00
Exponential FamiliesQuestion 1 of 10

A normal model has mean μ>0\mu>0 and variance constrained by σ2=μ2\sigma^2=\mu^2. Viewed as a subfamily of the two-parameter normal exponential family with natural parameters η1=μ/σ2\eta_1=\mu/\sigma^2 and η2=1/(2σ2)\eta_2=-1/(2\sigma^2), which statement is correct?

It is a full one-parameter family with η2=η1/2\eta_2=-\eta_1/2 and natural parameter space η1>0\eta_1>0.
It is a curved one-parameter family with η2=η12/2\eta_2=-\eta_1^2/2 and natural parameter space η1>0\eta_1>0.
It is a curved one-parameter family with η2=1/(2η12)\eta_2=-1/(2\eta_1^2) and natural parameter space η1<0\eta_1<0.
It is a full two-parameter family because both η1\eta_1 and η2\eta_2 vary as μ\mu varies.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Exponential Families

Practice Exponential Families in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Exponential Families, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A normal model has mean μ>0\mu>0 and variance constrained by σ2=μ2\sigma^2=\mu^2. Viewed as a subfamily of the two-parameter normal exponential family with natural parameters η1=μ/σ2\eta_1=\mu/\sigma^2 and η2=1/(2σ2)\eta_2=-1/(2\sigma^2), which statement is correct?

  1. It is a full one-parameter family with η2=η1/2\eta_2=-\eta_1/2 and natural parameter space η1>0\eta_1>0.
  2. It is a curved one-parameter family with η2=η12/2\eta_2=-\eta_1^2/2 and natural parameter space η1>0\eta_1>0. (correct answer)
  3. It is a curved one-parameter family with η2=1/(2η12)\eta_2=-1/(2\eta_1^2) and natural parameter space η1<0\eta_1<0.
  4. It is a full two-parameter family because both η1\eta_1 and η2\eta_2 vary as μ\mu varies.
Explanation: When a question asks you to embed a constrained subfamily into an exponential family, your first task is to express the natural parameters purely in terms of the free parameter, then classify the resulting curve in natural parameter space. Here, the two-parameter normal exponential family has natural parameters η1=μ/σ2\eta_1 = \mu/\sigma^2 and η2=1/(2σ2)\eta_2 = -1/(2\sigma^2). The constraint σ2=μ2\sigma^2 = \mu^2 (with μ>0\mu > 0) means σ=μ\sigma = \mu. Substituting: η1=μ/μ2=1/μ\eta_1 = \mu/\mu^2 = 1/\mu and η2=1/(2μ2)\eta_2 = -1/(2\mu^2). Since η1=1/μ\eta_1 = 1/\mu, we get μ=1/η1\mu = 1/\eta_1, so η2=1/(2(1/η1)2)=η12/2\eta_2 = -1/(2(1/\eta_1)^2) = -\eta_1^2/2. Because μ>0\mu > 0, we have η1>0\eta_1 > 0. This relationship η2=η12/2\eta_2 = -\eta_1^2/2 traces a curved (nonlinear) path through the two-dimensional natural parameter space — the hallmark of a curved exponential family. This confirms B is correct. A is wrong because it claims η2=η1/2\eta_2 = -\eta_1/2, a linear relationship, which would make this a full (not curved) subfamily. The algebra simply doesn't support linearity here. C has the right functional form η2=1/(2η12)\eta_2 = -1/(2\eta_1^2), but this expression equals μ2/2-\mu^2/2, not η12/2-\eta_1^2/2. It also incorrectly states η1<0\eta_1 < 0, when μ>0\mu > 0 forces η1>0\eta_1 > 0. D is a conceptual error. Variation of both natural parameters with a single underlying parameter doesn't make a family "full" — fullness requires the natural parameter space to contain an open set in R2\mathbb{R}^2, which a one-dimensional curve never does. Remember: curved = nonlinear constraint in natural parameter space; full = open set in natural parameter space. Always substitute and simplify before classifying.

Question 2

For one observation from the uniform model f(xθ)=θ1I{0<x<θ}f(x\mid\theta)=\theta^{-1}I\{0<x<\theta\} with θ>0\theta>0, a researcher proposes treating logθ-\log\theta as a natural parameter. Which assessment is most accurate?

  1. The model is a regular one-parameter exponential family because its log density contains logθ-\log\theta.
  2. The model is not a regular fixed-support exponential family because the support depends on θ\theta. (correct answer)
  3. The model is a curved exponential family because the indicator is a nonlinear canonical statistic.
  4. The model becomes a regular exponential family after using the parameter η=1/θ\eta=1/\theta.
Explanation: When classifying a statistical model as an exponential family, the most critical requirement is that the support of the distribution must not depend on the parameter. This is the fixed-carrier condition, and it's the first thing you should check before attempting any reparameterization. For the uniform model f(xθ)=θ1I{0<x<θ}f(x\mid\theta) = \theta^{-1}I\{0 < x < \theta\}, the support {x:0<x<θ}\{x : 0 < x < \theta\} moves with θ\theta. This immediately disqualifies the model from being any form of regular exponential family, because the exponential family form h(x)exp{η(θ)T(x)A(θ)}h(x)\exp\{\eta(\theta)T(x) - A(\theta)\} requires h(x)h(x) to be free of θ\theta—but the indicator function I{0<x<θ}I\{0 < x < \theta\} cannot be factored into a θ\theta-free piece times a θ\theta-only piece. That's why B is correct. Answer A is tempting because logf=logθ\log f = -\log\theta does appear in the log density, making it look like the log density has the right shape. But this reasoning ignores the indicator function entirely—the support problem persists regardless of how you write the density's main term. Answer D proposes the reparameterization η=1/θ\eta = 1/\theta. While valid reparameterizations can simplify exponential family representations, they cannot fix a varying-support problem. The support still depends on θ\theta, and no algebraic trick removes that dependence. Answer C misuses the term "curved exponential family." Curved exponential families are subfamilies of regular exponential families with constraints on the natural parameter space—they don't arise from indicator-function complications. Study tip: Before attempting any exponential family classification, always check the support first. If the support depends on θ\theta, the conversation ends there—no reparameterization can rescue it.

Question 3

Independent binary responses satisfy P(Yi=1)=Φ(xiTβ)P(Y_i=1)=\Phi(x_i^{\mathsf T}\beta), where Φ\Phi is the standard normal distribution function. There are n>pn>p observations, and the design matrix has rank pp.

Viewed inside the full exponential family of independent Bernoulli distributions, how should this probit regression model generally be classified?

  1. It is not an exponential family because the probit link is not the Bernoulli canonical link.
  2. It is a full regular family with natural parameter vector XβX\beta.
  3. It is a curved family with coordinates ηi=\logit{Φ(xiTβ)}\eta_i=\logit\{\Phi(x_i^{\mathsf T}\beta)\}. (correct answer)
  4. It is a full family because every component ηi\eta_i ranges over all real numbers.
Explanation: When working with exponential families, a crucial distinction is between the full family (where the natural parameter vector ranges freely over an open set with dimension equal to the sufficient statistic) and a curved exponential family (where the parameter of interest lives on a lower-dimensional submanifold). This question asks you to place probit regression within that taxonomy. The full Bernoulli exponential family has natural parameters ηi=logpi1pi\eta_i = \log\frac{p_i}{1-p_i} for each observation, living freely in Rn\mathbb{R}^n. In probit regression, however, these nn natural parameters are constrained: ηi=logit{Φ(xiTβ)}\eta_i = \text{logit}\{\Phi(x_i^\mathsf{T}\beta)\}, which traces out a pp-dimensional curved submanifold inside Rn\mathbb{R}^n as β\beta varies. Because the model is parameterized by βRp\beta \in \mathbb{R}^p with p<np < n, it sits on a curved, lower-dimensional surface within the full family — making C the correct classification. Answer A is wrong because membership in an exponential family doesn't require the canonical link. Probit regression is still a valid statistical model; the question is about its position within the full Bernoulli family, not whether it "qualifies" as exponential. Answer B fails because XβX\beta is not the natural parameter vector — logit{Φ(xiTβ)}\text{logit}\{\Phi(x_i^\mathsf{T}\beta)\} is, and the model does not fill the full nn-dimensional natural parameter space. Answer D confuses the range of individual components with the dimension of the parameter manifold; even though each ηi\eta_i can range over all reals, the nn components are jointly constrained to a pp-dimensional surface. A useful rule of thumb: whenever a model imposes p<np < n constraints on an nn-dimensional natural parameter, think curved exponential family.

Question 4

An exponential-family representation uses the statistic vector T(X)=(X,2X+1)TT(X)=(X,2X+1)^{\mathsf T} and natural parameter vector η=(η1,η2)T\eta=(\eta_1,\eta_2)^{\mathsf T}.

Which statement correctly identifies the nonidentifiability in this representation?

  1. Two parameter vectors are equivalent when their changes satisfy Δη1+2Δη2=0\Delta\eta_1+2\Delta\eta_2=0. (correct answer)
  2. Two parameter vectors are equivalent when their changes satisfy 2Δη1+Δη2=02\Delta\eta_1+\Delta\eta_2=0.
  3. The representation is identifiable because the two components of T(X)T(X) are numerically different.
  4. Only changes satisfying Δη1=Δη2=0\Delta\eta_1=\Delta\eta_2=0 can produce the same distribution.
Explanation: When a statistic vector has linearly dependent components, the exponential family representation becomes nonidentifiable — multiple parameter vectors produce the identical distribution. Your job is to find which parameter changes leave the inner product ηTT(X)\eta^{\mathsf T} T(X) unchanged for all values of XX. The key quantity is the linear combination in the exponent: ηTT(X)=η1X+η2(2X+1)\eta^{\mathsf T} T(X) = \eta_1 \cdot X + \eta_2 \cdot (2X+1). Expanding gives (η1+2η2)X+η2(\eta_1 + 2\eta_2)X + \eta_2. Two parameter vectors η\eta and η+Δη\eta + \Delta\eta yield the same distribution if and only if their exponents agree for all XX, meaning both the coefficient of XX and the constant term are unchanged. The coefficient of XX changes by Δη1+2Δη2\Delta\eta_1 + 2\Delta\eta_2, so equivalence requires Δη1+2Δη2=0\Delta\eta_1 + 2\Delta\eta_2 = 0. This confirms A is correct. B reverses the coefficients — it would be correct if T(X)=(2X,X+1)TT(X) = (2X, X+1)^{\mathsf T} or some other arrangement, but it misreads which component multiplies which parameter. C is a classic trap: the two components of T(X)T(X) are numerically different for any given XX, but that's irrelevant — what matters is whether they are linearly independent as functions, and (X)(X) and (2X+1)(2X+1) are not (they span only a one-dimensional function space). D would be the answer only if T(X)T(X) had linearly independent components; since it doesn't, a whole line of parameter shifts produces the same distribution. Study tip: Whenever you see an exponential family, immediately check whether the components of T(X)T(X) are linearly independent as functions of XX. If not, set ΔηTT(X)=0\Delta\eta^{\mathsf T} T(X) = 0 and solve — the resulting constraint equation directly reveals the nonidentifiability structure.

Question 5

A minimal exponential family is written as f(xη)=h(x)exp{ηTT(x)A(η)}f(x\mid\eta)=h(x)\exp\{\eta^{\mathsf T}T(x)-A(\eta)\}. A new statistic is defined by T(x)=MT(x)+bT^*(x)=MT(x)+b, where MM is nonsingular and bb is fixed.

Which natural parameter and log-partition function produce an equivalent representation using TT^*?

  1. η=MTη\eta^*=M^{\mathsf T}\eta and A(η)=A(η)ηTM1bA^*(\eta^*)=A(\eta)-\eta^{\mathsf T}M^{-1}b
  2. η=Mη\eta^*=M\eta and A(η)=A(η)ηTbA^*(\eta^*)=A(\eta)-\eta^{\mathsf T}b
  3. η=M1η\eta^*=M^{-1}\eta and A(η)=A(η)+bTbA^*(\eta^*)=A(\eta)+b^{\mathsf T}b
  4. η=MTη\eta^*=M^{-\mathsf T}\eta and A(η)=A(η)+ηTM1bA^*(\eta^*)=A(\eta)+\eta^{\mathsf T}M^{-1}b (correct answer)
Explanation: When reparametrizing an exponential family, your goal is to rewrite ηTT(x)\eta^{\mathsf T}T(x) in terms of the new sufficient statistic T(x)=MT(x)+bT^*(x) = MT(x) + b, preserving the density exactly. This requires finding η\eta^* such that (η)TT(x)(\eta^*)^{\mathsf T}T^*(x) reproduces the original inner product, then adjusting the log-partition function to absorb any leftover constants. Start by solving: you need (η)TT(x)=ηTT(x)(\eta^*)^{\mathsf T}T^*(x) = \eta^{\mathsf T}T(x). Substituting T(x)=MT(x)+bT^*(x) = MT(x) + b, expand the left side as (η)TMT(x)+(η)Tb(\eta^*)^{\mathsf T}MT(x) + (\eta^*)^{\mathsf T}b. For the T(x)T(x) terms to match, you need (η)TM=ηT(\eta^*)^{\mathsf T}M = \eta^{\mathsf T}, which gives η=MTη\eta^* = M^{-\mathsf T}\eta. This confirms D is correct. The residual constant term (η)Tb=ηTM1b(\eta^*)^{\mathsf T}b = \eta^{\mathsf T}M^{-1}b must be absorbed into the new log-partition function: A(η)=A(η)+ηTM1bA^*(\eta^*) = A(\eta) + \eta^{\mathsf T}M^{-1}b, exactly as in D. A uses MTηM^{\mathsf T}\eta instead of MTηM^{-\mathsf T}\eta—a common transpose/inverse mix-up—and gets the sign on the correction term wrong. B applies MM directly to η\eta without transposing or inverting, ignoring the structure of the inner product algebra entirely. C inverts without transposing (M1ηM^{-1}\eta instead of MTηM^{-\mathsf T}\eta) and replaces the correction term with the unrelated quantity bTbb^{\mathsf T}b, which has no basis in the derivation. A useful rule of thumb: when you see a linear transformation of a sufficient statistic, always derive η\eta^* by matching inner products term-by-term, and track every leftover scalar into AA^*. Confusing transpose with inverse is the most common trap here.

Question 6

Independent event counts satisfy YiPoisson(eiλ)Y_i\sim\operatorname{Poisson}(e_i\lambda), where the exposures ei>0e_i>0 are known and λ>0\lambda>0 is common to all observations.

Using η=logλ\eta=\log\lambda, which pair gives the canonical statistic and log-partition function for the joint model, up to terms independent of η\eta?

  1. T=iYiT=\sum_iY_i and A(η)=eηieiA(\eta)=e^\eta\sum_ie_i (correct answer)
  2. T=ieiYiT=\sum_ie_iY_i and A(η)=neηA(\eta)=ne^\eta
  3. T=iYiT=\sum_iY_i and A(η)=neηieiA(\eta)=ne^\eta\sum_ie_i
  4. T=iYi/eiT=\sum_iY_i/e_i and A(η)=eηieiA(\eta)=e^\eta\sum_ie_i
Explanation: When you encounter exponential family questions, your job is to factor the joint log-likelihood into the canonical form (η)=ηT(Y)A(η)+C(Y)\ell(\eta) = \eta \cdot T(\mathbf{Y}) - A(\eta) + C(\mathbf{Y}), identifying the sufficient statistic TT and log-partition function A(η)A(\eta) by inspection. Start with the joint log-likelihood. Since the YiY_i are independent Poisson(eiλ)(e_i\lambda), the log-likelihood is: (λ)=i[Yilog(eiλ)eiλlog(Yi!)]\ell(\lambda) = \sum_i \left[ Y_i \log(e_i\lambda) - e_i\lambda - \log(Y_i!) \right] Substituting λ=eη\lambda = e^\eta: (η)=iYi(logei+η)eηieiilog(Yi!)\ell(\eta) = \sum_i Y_i(\log e_i + \eta) - e^\eta\sum_i e_i - \sum_i\log(Y_i!) Collecting terms involving η\eta: (η)=ηiYiTeηieiA(η)+(terms free of η)\ell(\eta) = \eta\underbrace{\sum_i Y_i}_{T} - \underbrace{e^\eta \sum_i e_i}_{A(\eta)} + \text{(terms free of } \eta\text{)} This confirms Answer A: T=iYiT = \sum_i Y_i and A(η)=eηieiA(\eta) = e^\eta \sum_i e_i. Answer B uses T=ieiYiT = \sum_i e_i Y_i, which would arise if eie_i multiplied YiY_i inside the canonical term — but the log-likelihood shows η\eta pairs with Yi\sum Y_i, not eiYi\sum e_i Y_i. The exposures belong in the log-partition function, not the sufficient statistic. Answer C incorrectly multiplies nn and ei\sum e_i together in A(η)A(\eta), confusing the structural role of the exposures. Answer D divides YiY_i by eie_i, which has no basis in the log-likelihood factorization. A reliable strategy: always write out the full joint log-likelihood first, then isolate every term containing your natural parameter. Whatever multiplies η\eta is your canonical statistic TT; whatever you subtract is A(η)A(\eta).

Question 7

A single multinomial observation has three categories with probabilities p1,p2,p3>0p_1,p_2,p_3>0. Using category 3 as the baseline gives natural parameters η1=log(p1/p3)\eta_1=\log(p_1/p_3) and η2=log(p2/p3)\eta_2=\log(p_2/p_3).

Suppose the model is restricted by p1p3=p22p_1p_3=p_2^2. Which description of the restricted model is correct?

  1. It is a one-parameter full exponential family satisfying η1=2η2\eta_1=2\eta_2. (correct answer)
  2. It is a curved exponential family satisfying η1η2=2\eta_1\eta_2=2.
  3. It is a one-parameter full exponential family satisfying 2η1=η22\eta_1=\eta_2.
  4. It is not an exponential family because the restriction is nonlinear in the probabilities.
Explanation: When a multinomial model is restricted, your first move should always be to translate the constraint into the natural parameter space — that's where exponential family structure becomes transparent. The natural parameters here are η1=log(p1/p3)\eta_1 = \log(p_1/p_3) and η2=log(p2/p3)\eta_2 = \log(p_2/p_3). Now apply the constraint p1p3=p22p_1 p_3 = p_2^2. Taking logarithms of both sides gives logp1+logp3=2logp2\log p_1 + \log p_3 = 2\log p_2. Rewriting in terms of the natural parameters: log(p1/p3)=2log(p2/p3)\log(p_1/p_3) = 2\log(p_2/p_3), which is exactly η1=2η2\eta_1 = 2\eta_2. This is a linear constraint on the natural parameters, so the restricted model is a one-parameter exponential family obtained by substituting η1=2η2\eta_1 = 2\eta_2 and letting η2=θ\eta_2 = \theta vary freely. Because the natural parameter space is an open interval and the parameterization is linear, this is a full exponential family — confirming answer A is correct. Answer B claims η1η2=2\eta_1 \eta_2 = 2, which would be a nonlinear (curved) constraint in the natural parameter space, making it a curved exponential family. But the log transformation turns the multiplicative constraint into an additive one, so this is the wrong algebraic form entirely. Answer C gets the direction backwards: it says 2η1=η22\eta_1 = \eta_2, but the derivation clearly gives η1=2η2\eta_1 = 2\eta_2, not its reciprocal. Answer D is a conceptual trap — nonlinearity in the probabilities does not disqualify a model from being an exponential family; what matters is linearity in the natural parameters after the log transformation. Your takeaway: always convert restrictions to log-odds (natural parameter) form before classifying the family. A constraint that looks curved in probability space can be perfectly linear — and thus "full" — in natural parameter space.

Question 8

For a scalar canonical exponential family, f(xη)=h(x)exp{ηT(x)A(η)}f(x\mid\eta)=h(x)\exp\{\eta T(x)-A(\eta)\}. A conjugate prior has kernel π(η)exp{ξηνA(η)}\pi(\eta)\propto\exp\{\xi\eta-\nu A(\eta)\}. After observing nn independent observations, assume the posterior mode is interior and unique.

Which equation characterizes the posterior mode η^\widehat\eta?

  1. A(η^)=(ξ+iT(Xi))/(ν+n)A''(\widehat\eta)=(\xi+\sum_iT(X_i))/(\nu+n)
  2. A(η^)=(ν+iT(Xi))/(ξ+n)A'(\widehat\eta)=(\nu+\sum_iT(X_i))/(\xi+n)
  3. A(η^)=(ξ+n)/(ν+iT(Xi))A'(\widehat\eta)=(\xi+n)/(\nu+\sum_iT(X_i))
  4. A(η^)=(ξ+iT(Xi))/(ν+n)A'(\widehat\eta)=(\xi+\sum_iT(X_i))/(\nu+n) (correct answer)
Explanation: When working with exponential families and conjugate priors, your instinct should be to write out the log-posterior and differentiate. The posterior combines the conjugate prior kernel with the likelihood, giving: logπ(ηx)ξηνA(η)+ηiT(xi)nA(η)=(ξ+iT(xi))η(ν+n)A(η)\log \pi(\eta \mid \mathbf{x}) \propto \xi\eta - \nu A(\eta) + \eta\sum_i T(x_i) - nA(\eta) = \left(\xi + \sum_i T(x_i)\right)\eta - (\nu + n)A(\eta) Taking the derivative with respect to η\eta and setting it to zero at the mode η^\widehat{\eta}: (ξ+iT(xi))(ν+n)A(η^)=0\left(\xi + \sum_i T(x_i)\right) - (\nu + n)A'(\widehat{\eta}) = 0 Solving directly yields A(η^)=ξ+iT(xi)ν+nA'(\widehat{\eta}) = \dfrac{\xi + \sum_i T(x_i)}{\nu + n}, which is D. This makes intuitive sense: in an exponential family, A(η)=Eη[T(X)]A'(\eta) = \mathbb{E}_\eta[T(X)], so the posterior mode sets the mean-value parameter equal to a precision-weighted average of the prior pseudo-observations and data sufficient statistics. A is wrong because it involves A(η^)A''(\widehat{\eta}), the second derivative (which governs curvature, not the mode condition). Confusing AA' with AA'' is a common careless error. B flips the roles of ξ\xi and ν\nu: ξ\xi counts pseudo-observations on the sufficient statistic side, while ν\nu acts as a pseudo-sample-size scaling A(η)A(\eta), so swapping them mixes up the prior's two hyperparameters. C inverts the entire fraction, which would arise from accidentally dividing by the wrong term. As a study tip: always distinguish the two roles of the conjugate prior hyperparameters — one multiplies η\eta (sufficient statistic side) and one multiplies A(η)A(\eta) (log-normalizer side). Keeping that structure clear prevents both B-type and C-type errors.

Question 9

A gamma random variable has density f(xα,β)=βαxα1eβx/Γ(α)f(x\mid\alpha,\beta)=\beta^\alpha x^{\alpha-1}e^{-\beta x}/\Gamma(\alpha) for x>0x>0, with α>0\alpha>0 and β>0\beta>0. Use natural parameters η1=α1\eta_1=\alpha-1 and η2=β\eta_2=-\beta and canonical statistics T1(X)=logXT_1(X)=\log X and T2(X)=XT_2(X)=X.

What is Cov(logX,X)\operatorname{Cov}(\log X,X) under this model?

  1. α/β\alpha/\beta
  2. 1/β-1/\beta
  3. 1/β1/\beta (correct answer)
  4. 1/(αβ)1/(\alpha\beta)
Explanation: Whenever you see a question about covariances of sufficient statistics in an exponential family, reach for the information matrix identity: the covariance matrix of the canonical sufficient statistics equals the matrix of second-order partial derivatives of the log-partition function A(η)A(\boldsymbol{\eta}) with respect to the natural parameters. For the gamma family with natural parameters η1=α1\eta_1 = \alpha - 1 and η2=β\eta_2 = -\beta, the log-partition function is: A(η1,η2)=logΓ(η1+1)(η1+1)log(η2)A(\eta_1, \eta_2) = \log\Gamma(\eta_1+1) - (\eta_1+1)\log(-\eta_2) The cross-covariance is given by the mixed partial derivative: Cov(T1,T2)=2Aη1η2\operatorname{Cov}(T_1, T_2) = \frac{\partial^2 A}{\partial \eta_1 \,\partial \eta_2} Taking A/η1=ψ(η1+1)log(η2)\partial A/\partial \eta_1 = \psi(\eta_1+1) - \log(-\eta_2), where ψ\psi is the digamma function, and then differentiating with respect to η2\eta_2: 2Aη1η2=1η2=1β\frac{\partial^2 A}{\partial \eta_1\,\partial \eta_2} = -\frac{1}{\eta_2} = \frac{1}{\beta} So Cov(logX,X)=1/β\operatorname{Cov}(\log X, X) = 1/\beta, confirming C. Choice A (α/β\alpha/\beta) corresponds to Var(X)=α/β2\operatorname{Var}(X) = \alpha/\beta^2 confused with the covariance — it mixes up variance of XX and scales incorrectly. Choice B (1/β-1/\beta) is a sign error: students often forget that η2=β<0\eta_2 = -\beta < 0, so 1/η2=+1/β-1/\eta_2 = +1/\beta, not 1/β-1/\beta. Choice D (1/(αβ)1/(\alpha\beta)) has no grounding in the derivatives and likely reflects a dimensional-analysis guess. Study tip: Memorize that in any exponential family, Cov(Ti,Tj)=2A/ηiηj\operatorname{Cov}(T_i, T_j) = \partial^2 A/\partial\eta_i\,\partial\eta_j. This single identity replaces lengthy moment calculations and is a recurring tool on graduate-level theory exams.

Question 10

A scalar natural exponential family has log-partition function A(η)=rlog(η)A(\eta)=-r\log(-\eta) for η<0\eta<0, where r>0r>0 is known. Let the mean be parameterized as m=r/ηm=-r/\eta.

What is the Fisher information in one observation when the model is parameterized by mm?

  1. Im(m)=r/m2I_m(m)=r/m^2 (correct answer)
  2. Im(m)=m2/rI_m(m)=m^2/r
  3. Im(m)=r2/m2I_m(m)=r^2/m^2
  4. Im(m)=1/(rm2)I_m(m)=1/(rm^2)
Explanation: When you see Fisher information asked under a reparameterization, your first instinct should be the reparameterization formula: if you change from the natural parameter η\eta to a new parameter mm, the Fisher information transforms as Im(m)=Iη(η)(dηdm)2I_m(m) = I_\eta(\eta)\left(\frac{d\eta}{dm}\right)^2. Start by finding Iη(η)I_\eta(\eta). In a natural exponential family, the Fisher information in the natural parameter equals the variance, which equals A(η)A''(\eta). Differentiating A(η)=rlog(η)A(\eta) = -r\log(-\eta) twice: A(η)=r/ηA'(\eta) = -r/\eta and A(η)=r/η2A''(\eta) = r/\eta^2. So Iη(η)=r/η2I_\eta(\eta) = r/\eta^2. Next, find dη/dmd\eta/dm. From m=r/ηm = -r/\eta, solve to get η=r/m\eta = -r/m, so dη/dm=r/m2d\eta/dm = r/m^2. Applying the chain rule formula: Im(m)=rη2(rm2)2I_m(m) = \frac{r}{\eta^2}\cdot\left(\frac{r}{m^2}\right)^2. But substitute η=r/m\eta = -r/m, giving η2=r2/m2\eta^2 = r^2/m^2, so Iη=r/(r2/m2)=m2/rI_\eta = r/(r^2/m^2) = m^2/r. Then Im(m)=m2rr2m4=rm2I_m(m) = \frac{m^2}{r}\cdot\frac{r^2}{m^4} = \frac{r}{m^2}. This confirms answer A. Answer B, m2/rm^2/r, is actually IηI_\eta expressed in terms of mm — it forgets to apply the Jacobian factor. Answer C, r2/m2r^2/m^2, likely arises from squaring r/mr/m without correctly assembling the transformation. Answer D, 1/(rm2)1/(rm^2), flips the role of rr in the numerator and denominator. Study tip: Always track which parameterization you're working in and apply the squared-derivative Jacobian explicitly — confusing IηI_\eta with ImI_m is one of the most common errors on reparameterization problems.