Statistics Graduate Level Quiz: Identifiability And Parameterization
9 questions · exam conditions
0:00
Identifiability And ParameterizationQuestion 1 of 9

A random-intercept model is proposed for one response from each of many clusters: Yi=μ+bi+εiY_i=\mu+b_i+\varepsilon_i, where biN(0,σb2)b_i\sim N(0,\sigma_b^2) and εiN(0,σe2)\varepsilon_i\sim N(0,\sigma_e^2) independently. Each cluster contributes exactly one observation.

Which modification most directly resolves the variance-component nonidentifiability in this model?

Increase the number of clusters indefinitely while retaining one independent response from each cluster.
Center all observed responses at their sample mean before fitting both variance components.
Obtain at least two conditionally independent responses in some clusters, creating observable within-cluster covariance.
Replace the Gaussian assumptions by two distinct symmetric distributions with unknown scale parameters.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Identifiability And Parameterization

Practice Identifiability And Parameterization in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Identifiability And Parameterization, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A random-intercept model is proposed for one response from each of many clusters: Yi=μ+bi+εiY_i=\mu+b_i+\varepsilon_i, where biN(0,σb2)b_i\sim N(0,\sigma_b^2) and εiN(0,σe2)\varepsilon_i\sim N(0,\sigma_e^2) independently. Each cluster contributes exactly one observation.

Which modification most directly resolves the variance-component nonidentifiability in this model?

  1. Increase the number of clusters indefinitely while retaining one independent response from each cluster.
  2. Center all observed responses at their sample mean before fitting both variance components.
  3. Obtain at least two conditionally independent responses in some clusters, creating observable within-cluster covariance. (correct answer)
  4. Replace the Gaussian assumptions by two distinct symmetric distributions with unknown scale parameters.
Explanation: Whenever you see a question about mixed models with random effects, ask yourself: what observed information would let me separate the two sources of variance? That's the heart of identifiability. In this model, Var(Yi)=σb2+σe2\text{Var}(Y_i) = \sigma_b^2 + \sigma_e^2. With only one observation per cluster, you observe a single marginal variance — one equation, two unknowns. There is no way to disentangle σb2\sigma_b^2 from σe2\sigma_e^2 because they enter the likelihood only through their sum. This is a classic nonidentifiability problem. C is correct because obtaining at least two conditionally independent observations within some clusters creates observable within-cluster covariance. For observations jkj \neq k in cluster ii, Cov(Yij,Yik)=σb2\text{Cov}(Y_{ij}, Y_{ik}) = \sigma_b^2, since they share bib_i but have independent ε\varepsilon's. This covariance is estimable from the data, giving you a second equation and allowing both components to be identified separately. More replication within clusters is the structural fix. A is wrong because adding more clusters only improves estimation of the marginal variance σb2+σe2\sigma_b^2 + \sigma_e^2 — you get more data on the same unidentifiable sum, not a way to split it. B is wrong because centering is a location transformation. It has no effect on variance estimation or identifiability, which is purely a property of the variance structure. D is wrong because changing distributional families doesn't resolve the structural confounding — two unknown scales summing into one observed variance remains nonidentifiable regardless of shape. Study tip: Identifiability problems are resolved by structural changes to the data design, not by transformations or larger samples of the same structure. Always ask what the data can and cannot distinguish.

Question 2

Suppose X1,,XnX_1,\ldots,X_n are independent with distribution N(θ2,1)N(\theta^2,1), where initially θR\theta\in\mathbb{R}. For one observed sample, the sample mean is Xˉ=0.49\bar X=0.49.

Which statement best distinguishes the sample-specific likelihood behavior from identifiability of the parameter?

  1. The parameter θ\theta is identifiable because the likelihood has exactly two isolated maximizers, whereas θ2\theta^2 is not identifiable.
  2. The parameter θ\theta is not identifiable, but θ2\theta^2 is identifiable; restricting the parameter space to θ0\theta\geq 0 makes θ\theta identifiable. (correct answer)
  3. The parameter θ\theta is identifiable for this sample because the positive maximizer 0.70.7 can be selected by convention without changing the model.
  4. Neither θ\theta nor θ2\theta^2 is identifiable, because the likelihood is unchanged when θ\theta is replaced by θ-\theta.
Explanation: Whenever you see a question mixing a nonlinear parameter transformation with likelihood analysis, your first instinct should be to separate two distinct concepts: identifiability (a property of the model/parameter space) versus likelihood behavior for a specific sample (a data-dependent phenomenon). A parameter θ\theta is identifiable if distinct values of θ\theta produce distinct distributions. Here, N(θ2,1)N(\theta^2, 1) depends on θ\theta only through θ2\theta^2, so θ=0.7\theta = 0.7 and θ=0.7\theta = -0.7 generate the exact same distribution for every possible dataset. This means θ\theta itself is not identifiable — no amount of data can distinguish +θ+\theta from θ-\theta. However, θ2\theta^2 maps each distribution to a unique value, so θ2\theta^2 is identifiable. Restricting to θ0\theta \geq 0 creates a one-to-one correspondence between parameter values and distributions, restoring identifiability for θ\theta. This is exactly what answer B captures. Answer A is wrong because it inverts the identifiability claims — θ\theta is the non-identifiable parameter, not θ2\theta^2. Saying θ2\theta^2 is not identifiable is precisely backwards. Answer C conflates a convention (choosing the positive root) with actual identifiability. Identifiability is a model-level property, not something rescued by selecting a convenient maximizer after observing data. The model itself remains symmetric in θ\theta. Answer D correctly notes the symmetry θθ\theta \to -\theta but wrongly concludes that θ2\theta^2 is also non-identifiable, which it isn't — θ2\theta^2 uniquely parameterizes the family. Study tip: Always check identifiability at the model level before analyzing the likelihood. If f(x;θ)=f(x;θ)f(x;\theta) = f(x;-\theta) for all xx, the parameter is not identifiable — and restricting the parameter space, not data-driven conventions, is the proper fix.

Question 3

A Gaussian factor model is specified as Y=ΛF+εY=\Lambda F+\varepsilon, where FN(0,Ir)F\sim N(0,I_r), εN(0,Ψ)\varepsilon\sim N(0,\Psi), and Ψ\Psi is known. No structural restrictions are placed on the loading matrix Λ\Lambda.

Assuming the marginal covariance of YY is known exactly, which object is identifiable without further constraints?

  1. Each column of Λ\Lambda is identifiable up to an independent sign change, but arbitrary rotations are distinguishable.
  2. The matrix ΛΛT\Lambda\Lambda^{\mathsf T} is identifiable, but Λ\Lambda is unchanged observationally under right multiplication by an orthogonal matrix. (correct answer)
  3. The matrix ΛTΛ\Lambda^{\mathsf T}\Lambda is identifiable, but the marginal covariance cannot determine the column space of Λ\Lambda.
  4. The loading matrix Λ\Lambda is fully identifiable whenever its columns are linearly independent and Ψ\Psi is known.
Explanation: When working with factor models, your first instinct should be to examine what the observed data actually constrains. The marginal covariance of YY is Σ=ΛΛT+Ψ\Sigma = \Lambda\Lambda^{\mathsf{T}} + \Psi. Since Ψ\Psi is known, knowing Σ\Sigma exactly pins down ΛΛT\Lambda\Lambda^{\mathsf{T}} uniquely — this product is identifiable. However, for any orthogonal matrix QQ (so QQT=IQQ^{\mathsf{T}} = I), substituting Λ=ΛQ\Lambda^* = \Lambda Q gives Λ(Λ)T=ΛQQTΛT=ΛΛT\Lambda^*(\Lambda^*)^{\mathsf{T}} = \Lambda QQ^{\mathsf{T}}\Lambda^{\mathsf{T}} = \Lambda\Lambda^{\mathsf{T}}, leaving the covariance unchanged. This means Λ\Lambda and ΛQ\Lambda Q are observationally equivalent for every orthogonal QQ — you can never distinguish them from the marginal distribution alone. This is exactly what B describes, making it correct. A is wrong because the rotational indeterminacy is far broader than independent sign flips. Sign changes correspond to diagonal orthogonal matrices (diagonal entries ±1\pm 1), a tiny subgroup; the full rotation group O(r)O(r) is the actual source of non-identifiability. C is wrong on two counts: ΛTΛ\Lambda^{\mathsf{T}}\Lambda is generally not identifiable (it changes under arbitrary rotations), and the column space of Λ\Lambda actually is determined once ΛΛT\Lambda\Lambda^{\mathsf{T}} is known (it equals the column space of that product). D is wrong because linear independence of the columns is insufficient to resolve rotational indeterminacy. You need additional structural constraints — such as requiring ΛTΨ1Λ\Lambda^{\mathsf{T}}\Psi^{-1}\Lambda to be diagonal with decreasing entries, or sparsity assumptions — to fully identify Λ\Lambda. The key study takeaway: in any factor model question, immediately write out the covariance decomposition and ask which transformations of Λ\Lambda leave it invariant. That invariance group defines precisely what remains unidentified.

Question 4

A proportional-hazards model is written as h(tx)=h0(t)exp(α+βx)h(t\mid x)=h_0(t)\exp(\alpha+\beta x), where the baseline hazard h0(t)h_0(t) is completely unspecified and positive. Exact event and censoring times are observed under independent censoring.

Which statement correctly characterizes the parameterization?

  1. The intercept α\alpha is not separately identifiable from h0(t)h_0(t), while β\beta and the product eαh0(t)e^\alpha h_0(t) are identifiable under standard conditions. (correct answer)
  2. The intercept α\alpha is identifiable from event times, but β\beta is not identifiable unless the baseline hazard is parametrically specified.
  3. Both α\alpha and h0(t)h_0(t) are identifiable because one is constant in time and the other is a function of time.
  4. Neither α\alpha nor β\beta is identifiable because the partial likelihood eliminates every multiplicative component of the hazard.
Explanation: When you encounter identifiability questions in survival models, your first instinct should be to ask: can two different parameter configurations produce identical likelihoods? If yes, those parameters are not separately identifiable. Here, the hazard is h(tx)=h0(t)exp(α+βx)=[eαh0(t)]exp(βx)h(t\mid x) = h_0(t)\exp(\alpha + \beta x) = [e^\alpha h_0(t)]\exp(\beta x). Notice that α\alpha and h0(t)h_0(t) always appear together as the product eαh0(t)e^\alpha h_0(t). You can shift α\alpha by any constant cc and absorb ece^c into h0(t)h_0(t) without changing the hazard at any time or covariate value — the model is observationally identical. Therefore, α\alpha and h0(t)h_0(t) individually have no unique values the data can pin down. However, their product eαh0(t)e^\alpha h_0(t) is identifiable, and so is β\beta, because β\beta governs ratios of hazards across covariate values, which the partial likelihood recovers without ever estimating h0(t)h_0(t). This makes A correct. B is wrong because it reverses the truth: β\beta is identifiable (via the partial likelihood) without specifying h0(t)h_0(t), and α\alpha is never separately identifiable regardless of what you know about event times. C is wrong because the argument "constant vs. function of time" is irrelevant to identifiability. A constant multiplied into an unspecified function simply redefines that function — the time-varying nature of h0(t)h_0(t) offers no leverage on α\alpha alone. D is wrong because the partial likelihood eliminates h0(t)h_0(t) but preserves β\beta — that is precisely its purpose and power. Study tip: On any survival analysis question, remember that Cox's partial likelihood is specifically designed to make β\beta identifiable while treating h0(t)h_0(t) (absorbing any intercept) as a nuisance.

Question 5

For a multinomial logistic regression with categories 1,,K1,\ldots,K, suppose Pr(Y=kx)=exp(xTβk)/j=1Kexp(xTβj)\Pr(Y=k\mid x)=\exp(x^{\mathsf T}\beta_k)/\sum_{j=1}^K\exp(x^{\mathsf T}\beta_j). The design matrix has full column rank, and all category probabilities are positive.

Which action yields an identifiable parameterization without changing the modeled probability distributions?

  1. Constrain every coefficient vector to have unit norm, because only the direction of each βk\beta_k affects category probabilities.
  2. Set one category's coefficient vector to zero, because adding the same vector to every βk\beta_k leaves all probabilities unchanged. (correct answer)
  3. Require the coefficient vectors to be pairwise orthogonal, because correlated category effects cause the likelihood invariance.
  4. Add a strictly convex ridge penalty, because a unique penalized optimizer proves the original likelihood model is identifiable.
Explanation: Whenever you see a question about identifiability in multinomial logistic regression, your first instinct should be to examine the model's invariances — specifically, what transformations of the parameters leave all predicted probabilities unchanged. The core issue here is that the softmax probability Pr(Y=kx)=exp(xTβk)/jexp(xTβj)\Pr(Y=k\mid x)=\exp(x^\mathsf{T}\beta_k)/\sum_j\exp(x^\mathsf{T}\beta_j) is invariant to adding any common vector c\mathbf{c} to every βk\beta_k. You can verify this directly: replacing each βk\beta_k with βk+c\beta_k+\mathbf{c} multiplies every numerator and denominator by exp(xTc)\exp(x^\mathsf{T}\mathbf{c}), which cancels exactly. This means the likelihood has infinitely many maximizers — a classic non-identifiability. B is correct because fixing one category's coefficient vector to zero (typically the reference category) breaks this shift symmetry by anchoring the parameter space, producing a unique, identifiable parameterization while leaving all modeled probabilities unchanged. A is wrong because unit-norm constraints don't resolve the invariance — the issue is translational (adding a shared vector), not about scale or direction of individual βk\beta_k's. Projecting onto the unit sphere doesn't eliminate the degeneracy. C is wrong because pairwise orthogonality among coefficient vectors is unrelated to the source of non-identifiability. The shift invariance exists regardless of how the βk\beta_k's relate to each other geometrically. D is a subtle trap: a ridge penalty does produce a unique optimizer of the penalized objective, but this doesn't make the original likelihood model identifiable — it merely regularizes estimation. Identifiability is a property of the model, not the optimization procedure. Study tip: Always distinguish between fixing a model's identifiability (structural) and regularizing its estimation (computational). On exam questions about identifiability, ask yourself what symmetry is causing the problem, then find the constraint that breaks exactly that symmetry.

Question 6

In a source population, disease status satisfies the logistic model Pr(D=1X=x)=expit(α+βx)\Pr(D=1\mid X=x)=\operatorname{expit}(\alpha+\beta x). Investigators use case-control sampling, selecting subjects with probabilities depending on disease status but not otherwise on XX. The population disease prevalence and the case and control sampling fractions are not known.

Under this sampling scheme, which conclusion about α\alpha and β\beta is correct?

  1. Both parameters remain identifiable because conditioning on the observed numbers of cases and controls removes only nuisance sampling probabilities.
  2. Only α\alpha remains identifiable because case-control sampling preserves disease prevalence but distorts the exposure odds ratio.
  3. Neither parameter is identifiable because retrospective sampling changes the entire conditional distribution of disease given exposure.
  4. The slope β\beta is identifiable, but the population intercept α\alpha is not identifiable without prevalence or sampling-fraction information. (correct answer)
Explanation: Whenever you encounter a question about logistic regression under case-control sampling, your anchor concept should be the prospective vs. retrospective likelihood distinction and what each sampling scheme preserves. In a case-control study, subjects are sampled based on disease status. The observed data follow a retrospective distribution, not the population distribution. However, a fundamental result due to Prentice and Pyke (1979) shows that the retrospective likelihood for case-control data is proportional to the prospective logistic likelihood — but only up to a shift in the intercept. Specifically, if the true model is Pr(D=1X=x)=expit(α+βx)\Pr(D=1\mid X=x)=\operatorname{expit}(\alpha+\beta x), then what you can estimate from case-control data is expit(α+βx)\operatorname{expit}(\alpha^*+\beta x), where α=α+log ⁣(π1π0)\alpha^* = \alpha + \log\!\left(\frac{\pi_1}{\pi_0}\right) encodes unknown sampling fractions π1,π0\pi_1, \pi_0 for cases and controls. Because these fractions are unknown (and prevalence is unknown), α\alpha is not recoverable — but β\beta, the log-odds-ratio parameter, cancels out of the sampling adjustment entirely and remains identifiable. This confirms D as correct. Choice A is wrong because while conditioning on case-control counts is one framing, it does not recover α\alpha — the intercept is still confounded with unknown sampling probabilities. Choice B has it exactly backwards: case-control sampling distorts the intercept (prevalence-related), not the exposure odds ratio, so β\beta — not α\alpha — is the identifiable parameter. Choice C overstates the damage; the entire slope structure is preserved, so the claim that "neither" is identifiable is false. A useful memory anchor: "case-control kills the intercept, not the slope." On any exam question mixing logistic regression with retrospective sampling, immediately ask yourself which parameter depends on prevalence — that one is lost without additional external information.

Question 7

A zero-inflated Poisson model has probabilities Pr(Y=0)=π+(1π)eλ\Pr(Y=0)=\pi+(1-\pi)e^{-\lambda} and Pr(Y=k)=(1π)eλλk/k!\Pr(Y=k)=(1-\pi)e^{-\lambda}\lambda^k/k! for positive integers kk. The initial parameter space is 0π10\leq\pi\leq1 and λ0\lambda\geq0.

Which restriction is sufficient to remove the boundary nonidentifiability while retaining ordinary Poisson models as special cases?

  1. Require 0<π<10<\pi<1 while allowing λ=0\lambda=0, because a positive inflation probability labels the structural-zero component and ensures distinct distributions.
  2. Require π1/2\pi\leq1/2 and λ0\lambda\geq0, because bounding the mixing weight below one-half prevents equivalent representations of the zero-count distribution.
  3. Require π>0\pi>0 and λ>0\lambda>0, thereby excluding both the ordinary Poisson model (π=0\pi=0) and the degenerate point-mass case.
  4. Require 0π<10\leq\pi<1 and λ>0\lambda>0, so positive-count ratios identify λ\lambda and their total mass relative to zero-count probability identifies π\pi. (correct answer)
Explanation: Identifiability in mixture models is the core concept here. A model is nonidentifiable at a boundary when two different parameter values produce the same distribution. In the zero-inflated Poisson (ZIP), ask yourself: when do distinct (π,λ)(\pi, \lambda) pairs collapse into the same probability distribution? The key insight is that when λ>0\lambda > 0, the positive-count probabilities Pr(Y=k)=(1π)eλλk/k!\Pr(Y=k) = (1-\pi)e^{-\lambda}\lambda^k/k! for k1k \geq 1 uniquely determine λ\lambda through their ratios (Pr(Y=k+1)/Pr(Y=k)=λ/(k+1)\Pr(Y=k+1)/\Pr(Y=k) = \lambda/(k+1)), and once λ\lambda is pinned down, the total mass on positive counts identifies (1π)(1-\pi), which identifies π\pi. The ordinary Poisson is recovered at π=0\pi = 0, which is allowed under D. Requiring λ>0\lambda > 0 excludes the degenerate case λ=0\lambda = 0, where every observation is zero regardless of π\pi, making π\pi completely unidentifiable. Requiring π<1\pi < 1 excludes the pure point-mass at zero, where λ\lambda becomes irrelevant. Answer D threads this needle precisely. Answer A fails because allowing λ=0\lambda = 0 reintroduces the boundary where any π\pi produces an all-zero distribution — the very nonidentifiability you want to remove. Answer B is wrong because the threshold π1/2\pi \leq 1/2 is arbitrary and has no probabilistic justification; the identifiability problem stems from boundary values of π\pi, not large ones. Answer C excludes π=0\pi = 0, which means ordinary Poisson models are no longer special cases — violating the explicit requirement stated in the question. When you encounter mixture model identifiability questions, always check two things: which boundary collapses the parameter space, and which restriction fixes that collapse without discarding scientifically meaningful special cases.

Question 8

In each arm of a randomized trial, a binary outcome is measured with nondifferential sensitivity SeSe and specificity SpSp. Let the true event probabilities be p1p_1 and p0p_0 and the observed positive probabilities be q1q_1 and q0q_0. The values of SeSe and SpSp are unknown, but investigators know that Se+Sp>1Se+Sp>1.

What aspect of the true treatment effect is identifiable from q1q_1 and q0q_0 under these assumptions?

  1. The exact risk difference p1p0p_1-p_0 is identifiable because randomization eliminates uncertainty from outcome misclassification.
  2. The risk ratio p1/p0p_1/p_0 is identifiable, although the two arm-specific event probabilities are not separately identifiable.
  3. The sign and nullity of p1p0p_1-p_0 are identifiable, but its magnitude is not identifiable without more misclassification information. (correct answer)
  4. No feature of the treatment contrast is identifiable, because unknown sensitivity can reverse every observed difference.
Explanation: When outcome misclassification is nondifferential, the observed probability in each arm follows the classic reclassification formula: qi=Sepi+(1Sp)(1pi)q_i = Se \cdot p_i + (1 - Sp)(1 - p_i). You can rewrite this as qi=(Se+Sp1)pi+(1Sp)q_i = (Se + Sp - 1)p_i + (1 - Sp). This is a linear transformation of pip_i with slope λ=Se+Sp1\lambda = Se + Sp - 1 and intercept 1Sp1 - Sp. The key insight is that since Se+Sp>1Se + Sp > 1, we know λ>0\lambda > 0 — the slope is strictly positive. This is why C is correct. Taking the difference: q1q0=λ(p1p0)q_1 - q_0 = \lambda(p_1 - p_0). Because λ>0\lambda > 0, the sign of q1q0q_1 - q_0 equals the sign of p1p0p_1 - p_0, and q1q0=0q_1 - q_0 = 0 if and only if p1p0=0p_1 - p_0 = 0. So you can determine the direction and nullity of the true risk difference. However, without knowing λ\lambda exactly, you cannot recover the magnitude — the attenuation factor is unknown. A is wrong because randomization addresses confounding, not measurement error. Misclassification bias persists regardless of how treatment was assigned, and the true risk difference is still shrunk by the unknown factor λ\lambda. B is wrong because the risk ratio is not preserved under this transformation. The additive intercept term (1Sp)(1-Sp) means q1/q0p1/p0q_1/q_0 \neq p_1/p_0 in general, so the ratio is not recoverable. D is wrong because Se+Sp>1Se + Sp > 1 guarantees λ>0\lambda > 0, which preserves the sign. Only if λ\lambda could be negative would reversal be possible. As a study habit, always ask: does a monotone transformation preserve the sign of a difference even when it distorts the magnitude? That distinction frequently separates what is identifiable from what is not.

Question 9

Consider the two-component normal mixture model f(y)=πϕ(y;μ1,1)+(1π)ϕ(y;μ2,1)f(y)=\pi\phi(y;\mu_1,1)+(1-\pi)\phi(y;\mu_2,1), where 0<π<10<\pi<1 and ϕ(;μ,1)\phi(\cdot;\mu,1) denotes a normal density with mean μ\mu and variance 11.

Which statement about imposing the constraint μ1<μ2\mu_1<\mu_2 is most accurate?

  1. It removes label switching when the component means are distinct, but it excludes the coincident-component case in which the mixing weight is not identifiable. (correct answer)
  2. It makes the model globally identifiable even when the component means coincide, because the mixing proportion then distinguishes the components.
  3. It is insufficient to remove label switching unless the additional restriction π<1/2\pi<1/2 is imposed simultaneously.
  4. It changes the set of mixture distributions whenever the means are distinct, because ordered components represent fewer observable densities.
Explanation: Mixture model identifiability questions hinge on two distinct problems: label switching and parameter identifiability. When you see a question like this, separate them carefully — a constraint can solve one while leaving the other untouched. In a two-component normal mixture, "label switching" means the parameterizations (π,μ1,μ2)(\pi, \mu_1, \mu_2) and (1π,μ2,μ1)(1-\pi, \mu_2, \mu_1) produce the exact same density, so the likelihood has symmetric modes. Imposing μ1<μ2\mu_1 < \mu_2 selects one labeling from each symmetric pair, effectively eliminating this redundancy — but only when μ1μ2\mu_1 \neq \mu_2. When the means coincide (μ1=μ2=μ\mu_1 = \mu_2 = \mu), the density collapses to ϕ(y;μ,1)\phi(y;\mu,1) regardless of π\pi, so π\pi becomes completely unidentifiable. The ordering constraint does nothing here because there's no pair to order. This is precisely what answer A captures, making it correct. Answer B is wrong because when the means coincide, π\pi does not distinguish components — the entire mixing structure vanishes from the observable density, leaving π\pi unidentifiable rather than newly identifiable. Answer C is wrong because adding π<1/2\pi < 1/2 is unnecessary; μ1<μ2\mu_1 < \mu_2 alone is sufficient to remove label switching when means are distinct. The additional restriction would only create ambiguity when π=1/2\pi = 1/2. Answer D is wrong because ordering constraints on parameters do not reduce the set of representable densities — every mixture distribution with distinct means is still representable; the constraint merely eliminates duplicate parameterizations. Study tip: Always ask two separate questions about any identifying constraint: Does it remove parameter redundancy (label switching)? And does it handle degenerate boundary cases (component collapse)? These often have different answers.