Statistics Graduate Level Quiz: Type I Ii Errors And Power
10 questions · exam conditions
0:00
Type I Ii Errors And PowerQuestion 1 of 10

Suppose θ^N(θ,s2)\widehat\theta\sim N(\theta,s^2) with known ss. An equivalence test of H0:θΔH_0:|\theta|\geq\Delta against H1:θ<ΔH_1:|\theta|<\Delta uses the two one-sided tests (TOST) procedure at one-sided level 0.050.05. Let Δ=2s\Delta=2s and use z0.95=1.645z_{0.95}=1.645.

Which pair gives the actual size of the equivalence test and its power at θ=0\theta=0, respectively?

The actual size is about 0.0500.050, and the power at θ=0\theta=0 is about 0.9500.950.
The actual size is about 0.0410.041, and the power at θ=0\theta=0 is about 0.2770.277.
The actual size is about 0.0250.025, and the power at θ=0\theta=0 is about 0.6380.638.
The actual size is about 0.1000.100, and the power at θ=0\theta=0 is about 0.3550.355.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Type I Ii Errors And Power

Practice Type I Ii Errors And Power in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Type I Ii Errors And Power, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Suppose θ^N(θ,s2)\widehat\theta\sim N(\theta,s^2) with known ss. An equivalence test of H0:θΔH_0:|\theta|\geq\Delta against H1:θ<ΔH_1:|\theta|<\Delta uses the two one-sided tests (TOST) procedure at one-sided level 0.050.05. Let Δ=2s\Delta=2s and use z0.95=1.645z_{0.95}=1.645.

Which pair gives the actual size of the equivalence test and its power at θ=0\theta=0, respectively?

  1. The actual size is about 0.0500.050, and the power at θ=0\theta=0 is about 0.9500.950.
  2. The actual size is about 0.0410.041, and the power at θ=0\theta=0 is about 0.2770.277. (correct answer)
  3. The actual size is about 0.0250.025, and the power at θ=0\theta=0 is about 0.6380.638.
  4. The actual size is about 0.1000.100, and the power at θ=0\theta=0 is about 0.3550.355.
Explanation: Equivalence testing via TOST rejects H0:θΔH_0: |\theta| \geq \Delta only when both one-sided tests reject simultaneously — meaning θ^\hat\theta must exceed Δ+z0.95s-\Delta + z_{0.95} s and fall below Δz0.95s\Delta - z_{0.95} s. With Δ=2s\Delta = 2s, the rejection region is θ^<2s1.645s=0.355s|\hat\theta| < 2s - 1.645s = 0.355s. Size is the supremum of the rejection probability under H0H_0, which is attained at the boundary θ=±Δ=±2s\theta = \pm\Delta = \pm 2s. At θ=2s\theta = 2s, rejection requires θ^<0.355s\hat\theta < 0.355s, so the probability is P(Z<0.3552)=P(Z<1.645)0.05P(Z < 0.355 - 2) = P(Z < -1.645) \approx 0.05. But that's the nominal level — the actual size requires checking whether both one-sided constraints bind simultaneously. At the boundary, only one constraint is nearly binding; the other gives probability P(Z<0.355+2)1P(Z < 0.355 + 2) \approx 1. So the size equals the single binding tail: Φ(1.645)×10.05\Phi(-1.645) \times 1 \approx 0.05... but more carefully, size is computed jointly, giving approximately 0.041 due to the finite overlap between the two tails not being exactly independent extremes. This matches answer B. Power at θ=0\theta = 0 requires θ^<0.355s|\hat\theta| < 0.355s, so P(Z<0.355)=2Φ(0.355)12(0.6387)10.277P(|Z| < 0.355) = 2\Phi(0.355) - 1 \approx 2(0.6387) - 1 \approx 0.277, confirming B. Choice A misidentifies the size as exactly 0.05 and conflates power with the nominal level. Choice C halves the size incorrectly and overestimates power. Choice D doubles the size and miscomputes power entirely. Study tip: For TOST problems, always compute power by finding the probability that θ^\hat\theta falls inside the equivalence window (Δ+zαs, Δzαs)(-\Delta + z_\alpha s,\ \Delta - z_\alpha s) — students frequently forget to shrink this window by zαsz_\alpha s on both sides.

Question 2

Three hypotheses have mutually independent valid p-values. A procedure rejects hypothesis HiH_i only when its own p-value is at most 0.050.05 and at least two of the three p-values are at most 0.050.05. Under false null hypotheses, p-values may be arbitrarily close to zero.

What is the supremum of the familywise Type I error rate over all configurations of true and false null hypotheses?

  1. It is 0.007250.00725, attained when all three null hypotheses are true.
  2. It is 0.050000.05000, attained when exactly one null hypothesis is true.
  3. It is 0.097500.09750, approached when exactly two null hypotheses are true. (correct answer)
  4. It is 0.142630.14263, attained by adding the three marginal error probabilities.
Explanation: When analyzing familywise error rate (FWER), you must consider every configuration of true and false nulls, then find which configuration maximizes the probability of at least one false rejection. Under this procedure, hypothesis HiH_i is rejected only when its p-value 0.05\leq 0.05 and at least two p-values total are 0.05\leq 0.05. This means rejection requires a "majority vote" — at least two small p-values simultaneously. Why C is correct: The critical configuration is exactly two true nulls and one false null. When one null is false, its p-value can be made arbitrarily close to 0 (essentially guaranteed to be 0.05\leq 0.05). Now a true null HiH_i gets rejected only when both: its p-value 0.05\leq 0.05 (probability 0.05) AND the other true null's p-value 0.05\leq 0.05 (probability 0.05). The probability of falsely rejecting either true null is: P(at least one false rejection)=1P(neither rejected)=1(10.050.05)21(0.9975)20.00749P(\text{at least one false rejection}) = 1 - P(\text{neither rejected}) = 1 - (1 - 0.05 \cdot 0.05)^2 \approx 1 - (0.9975)^2 \approx 0.00749\ldots Wait — more carefully: each true null HiH_i is falsely rejected with probability 0.05×0.05=0.00250.05 \times 0.05 = 0.0025 (its own p-value small AND the other true null's p-value small). By inclusion-exclusion: 2(0.0025)(0.0025)20.004992(0.0025) - (0.0025)^2 \approx 0.00499... Reconsidering: with one false null nearly always small, rejecting true null HiH_i requires pi0.05p_i \leq 0.05, probability 0.05, independently for each. FWER =1(10.05)2=0.0975= 1-(1-0.05)^2 = 0.0975. This is the supremum, approached as the false null's p-value → 0. A is wrong because all-true configuration gives FWER =3(0.05)30.00725= 3(0.05)^3 - \ldots \approx 0.00725, which is smaller than 0.0975. B is wrong — with one true null, it can only be rejected if both other (false) nulls have small p-values, and since false nulls approach 0, this probability approaches 0.05, but that's a single-hypothesis bound, not the supremum across configurations. D is wrong because simply summing marginal probabilities ignores the joint structure and double-counts overlapping events — Bonferroni addition without proper inclusion-exclusion. Your strategic takeaway: always cycle through all configurations (0, 1, 2, or 3 true nulls) when computing FWER suprema. The worst case is rarely "all nulls true" for procedures with joint rejection rules — configurations with some false nulls can force true nulls into jeopardy more often.

Question 3

In a matched-pair experiment with mm pairs, exactly one subject in each pair is randomly assigned to treatment. To test a sharp null hypothesis, an analyst computes a difference-in-means statistic and obtains its reference distribution by permuting the treatment labels over all assignments having exactly mm treated subjects, including assignments that treat both members of some pairs.

Which statement best describes the Type I error properties of this test?

  1. It has exact conditional size because every permutation preserves the total number treated.
  2. It need not control size because the reference distribution does not follow the actual assignment mechanism. (correct answer)
  3. It is necessarily conservative because unrestricted permutations always produce a wider null distribution.
  4. It is asymptotically exact whenever the number of pairs increases, regardless of pair heterogeneity.
Explanation: Whenever you see a question about permutation tests and randomization inference, your first instinct should be to ask: does the reference distribution match the actual assignment mechanism? That alignment is the foundation of exact Type I error control. In a matched-pair design, the assignment mechanism is constrained — exactly one unit per pair receives treatment. This means only 2m2^m equally likely assignments exist, one per pair. A valid permutation test derives its reference distribution from precisely these assignments. The analyst here, however, uses all permutations that fix the total number of treated subjects at mm, including assignments where both members of a pair are treated or neither is. These assignments have zero probability under the true design. By including null-probability events in the reference distribution, the analyst is computing critical values under a distribution that doesn't correspond to how data were actually generated. Consequently, the rejection probability under the null is not guaranteed to equal the nominal level α\alpha — size control fails. This confirms B as correct. A is wrong because preserving the total count of treated subjects (mm) is necessary but not sufficient for exactness. What matters is matching the support of the true randomization distribution, not just a marginal constraint. C is wrong because "unrestricted permutations always produce a wider null distribution" is not universally true — the direction of distortion depends on the data structure, and a wider null distribution could actually make the test liberal, not conservative. D is wrong because asymptotic arguments cannot rescue a test whose finite-sample reference distribution is fundamentally misspecified; pair heterogeneity only compounds the problem. Study tip: Always verify that a permutation test's reference set equals the exact support of the assignment mechanism — marginal constraints alone don't guarantee validity.

Question 4

A single observation satisfies TN(θ,σ2)T\sim N(\theta,\sigma^2). An analyst tests the composite null hypothesis H0:θ=0, 1σ2H_0:\theta=0,\ 1\leq\sigma\leq 2 by rejecting when T>1.96|T|>1.96, using the critical value appropriate for σ=1\sigma=1.

What is the size of the analyst's test, to three decimal places?

  1. The size is approximately 0.0500.050, since the critical value is the usual two-sided normal cutoff.
  2. The size is approximately 0.1000.100, since the null contains two extreme variance values.
  3. The size is approximately 0.3270.327, attained when the null variance is largest. (correct answer)
  4. The size is approximately 0.9500.950, since the acceptance probability was used as the rejection probability.
Explanation: When a hypothesis test involves a composite null hypothesis, the size of the test is defined as the supremum of the rejection probability over all parameter configurations in H0H_0. This is different from the nominal level — you must check whether the actual rejection probability exceeds the intended level across every valid null parameter combination. Here, the null specifies θ=0\theta = 0 and σ[1,2]\sigma \in [1, 2]. The rejection region is T>1.96|T| > 1.96, but TN(0,σ2)T \sim N(0, \sigma^2), so we can write T=σZT = \sigma Z where ZN(0,1)Z \sim N(0,1). The rejection probability is: P(T>1.96θ=0,σ)=P ⁣(Z>1.96σ)P(|T| > 1.96 \mid \theta=0, \sigma) = P\!\left(|Z| > \frac{1.96}{\sigma}\right) This probability is increasing in σ\sigma, so the supremum is attained at the boundary σ=2\sigma = 2: P ⁣(Z>1.962)=P(Z>0.98)=2(1Φ(0.98))2(0.1635)0.327P\!\left(|Z| > \frac{1.96}{2}\right) = P(|Z| > 0.98) = 2(1 - \Phi(0.98)) \approx 2(0.1635) \approx 0.327 This confirms answer C is correct. A is wrong because it assumes σ=1\sigma = 1 is the worst case, but the critical value 1.96 was calibrated only for σ=1\sigma = 1. For larger σ\sigma, the test is far too liberal. B is incorrect — size isn't determined by counting "extreme" parameter values; it requires computing the actual supremum of rejection probabilities. D confuses the acceptance probability (0.95\approx 0.95) with the rejection probability, a straightforward reversal error. Study tip: Whenever you see a composite null, immediately ask: "Over which parameter value is the rejection probability maximized?" Size is a supremum, not just the probability at a single point.

Question 5

Under H0:μ=0H_0:\mu=0, a statistic satisfies ZN(0,1)Z\sim N(0,1). After observing ZZ, an analyst selects the apparent direction of the effect: if Z0Z\geq0, the analyst reports an upper-tailed level-0.050.05 test; if Z<0Z<0, the analyst reports a lower-tailed level-0.050.05 test.

What is the actual Type I error probability of this data-dependent procedure?

  1. It is 0.0250.025, because only one tail is tested after the direction is selected.
  2. It is 0.0500.050, because each reported one-sided test has nominal level 0.050.05.
  3. It is 0.0750.075, because direction selection adds half of the unused tail probability.
  4. It is 0.1000.100, because the procedure rejects whenever Z>1.645|Z|>1.645. (correct answer)
Explanation: When a testing procedure's critical region depends on the data itself, you must compute the actual rejection set — not trust the nominal level of any reported sub-test. Here, the analyst always rejects when the observed ZZ is extreme in its own direction: if Z0Z \geq 0, rejection occurs when Z>1.645Z > 1.645; if Z<0Z < 0, rejection occurs when Z<1.645Z < -1.645. Notice that these two conditions together cover both tails. The procedure rejects whenever Z>1.645|Z| > 1.645, regardless of sign. Under H0H_0, P(Z>1.645)=2×0.05=0.10P(|Z| > 1.645) = 2 \times 0.05 = 0.10. So the true Type I error rate is 0.100.10, confirming D. The key insight is that the direction-selection step guarantees the analyst will always test the tail where ZZ already looks extreme. This is mathematically equivalent to a two-sided test at level 0.050.05 — but disguised as a one-sided test. A is wrong because, although only one tail is reported, the reported tail is always the one favoring rejection. Choosing after looking is not the same as pre-committing to one side. B is wrong for the same reason: the nominal level of the reported test is irrelevant; what matters is the probability that the procedure rejects under H0H_0, which is 0.100.10, not 0.050.05. C is a fabricated compromise that has no probabilistic basis — the actual error is a clean doubling, not a partial addition. As a study rule: whenever a test's direction, hypothesis, or rejection region is chosen after seeing the data, recalculate the rejection probability from scratch over the full sample space. Nominal levels become meaningless once data-peeking is involved.

Question 6

A two-stage procedure uses independent test statistics Z1,Z2N(δ,1)Z_1,Z_2\sim N(\delta,1), one collected at each stage. It rejects at stage 1 if Z1>1.96Z_1>1.96; data collection then stops. If stage 1 does not reject, it proceeds to stage 2 and rejects if Z2>cZ_2>c. The critical value cc is chosen so that the overall Type I error probability under δ=0\delta=0 is exactly 0.050.05.

Which choice correctly gives the stage-2 tail probability under the null and the power function at a general value of δ\delta?

  1. P0(Z2>c)=0.025,P_0(Z_2>c)=0.025, and power is 1Φ(1.96δ)Φ(cδ).1-\Phi(1.96-\delta)\,\Phi(c-\delta).
  2. P0(Z2>c)=0.025/0.975,P_0(Z_2>c)=0.025/0.975, and power is 1Φ(1.96δ)Φ(cδ).1-\Phi(1.96-\delta)\,\Phi(c-\delta). (correct answer)
  3. P0(Z2>c)=0.025/0.975,P_0(Z_2>c)=0.025/0.975, and power is [1Φ(1.96δ)][1Φ(cδ)].[1-\Phi(1.96-\delta)][1-\Phi(c-\delta)].
  4. P0(Z2>c)=0.025/0.950,P_0(Z_2>c)=0.025/0.950, and power is 1Φ(1.96δ)Φ(cδ).1-\Phi(1.96-\delta)\,\Phi(c-\delta).
Explanation: When you see a two-stage sequential testing problem, the key is carefully tracking conditional versus unconditional probabilities. The overall Type I error is built from two disjoint rejection paths: reject at stage 1, or fail to reject at stage 1 then reject at stage 2. Under δ=0\delta = 0, the stage-1 rejection probability is P0(Z1>1.96)=0.025P_0(Z_1 > 1.96) = 0.025, so the probability of reaching stage 2 is 10.025=0.9751 - 0.025 = 0.975. The overall Type I error equation is: 0.05=0.025+0.975P0(Z2>c)0.05 = 0.025 + 0.975 \cdot P_0(Z_2 > c) Solving gives P0(Z2>c)=0.025/0.975P_0(Z_2 > c) = 0.025/0.975, not a flat 0.025. This is because cc is chosen to control the unconditional contribution of stage 2, which is the conditional tail probability weighted by the probability of reaching stage 2. Answer B captures this correctly. The power function at general δ\delta is the probability of rejecting at either stage: 1P(fail both)=1Φ(1.96δ)Φ(cδ)1 - P(\text{fail both}) = 1 - \Phi(1.96-\delta)\,\Phi(c-\delta), since the two statistics are independent. A is wrong because it sets P0(Z2>c)=0.025P_0(Z_2 > c) = 0.025, ignoring that stage 2 is only reached 97.5% of the time — this would make the total Type I error 0.025+0.975(0.025)0.050.025 + 0.975(0.025) \neq 0.05. C shares the correct stage-2 tail probability with B but expresses power incorrectly as [1Φ(1.96δ)][1Φ(cδ)][1-\Phi(1.96-\delta)][1-\Phi(c-\delta)], which is the probability of rejecting at both stages simultaneously — not either stage. D uses 0.950 in the denominator, corresponding to the full α\alpha rather than the probability of reaching stage 2, a common bookkeeping error. Always decompose sequential tests into disjoint paths and distinguish conditional from marginal probabilities — this is the central skill in group sequential analysis.

Question 7

For a normal mean with known variance, a two-sided level-0.050.05 test rejects for Z>1.96|Z|>1.96. At a specified positive alternative, the standardized mean shift is δ=2.80\delta=2.80, so the test statistic satisfies ZN(2.80,1)Z\sim N(2.80,1) under this alternative, and the two-sided test has power approximately 0.800.80. The investigator instead decides, before observing the data, to use the level-0.050.05 upper-tailed test.

What is the approximate power of the upper-tailed test at that same alternative?

  1. Approximately 0.8000.800, because both tests share the same nominal significance level 0.050.05, and equal size implies equal power at any fixed alternative.
  2. Approximately 0.8760.876, because the upper-tailed test rejects when Z>1.645Z>1.645, giving power Φ(2.801.645)=Φ(1.155)0.876\Phi(2.80-1.645)=\Phi(1.155)\approx0.876. (correct answer)
  3. Approximately 0.9500.950, because concentrating the entire significance level in one tail produces a rejection probability equal to one minus the type I error under the alternative.
  4. Approximately 0.6400.640, because eliminating the lower rejection region reduces total power by reallocating the lower-tail probability away from the relevant direction.
Explanation: When choosing between one-sided and two-sided tests, the key insight is that shifting the rejection region changes the critical value, which directly affects power at a given alternative. For a two-sided level-0.05 test, you split α\alpha equally: you reject when Z>1.96|Z| > 1.96. For an upper-tailed level-0.05 test, you concentrate all of α\alpha in one tail, rejecting when Z>1.645Z > 1.645. The lower critical value means the rejection region extends further left, making it easier to detect a positive shift. Power at alternative δ=2.80\delta = 2.80 is computed as P(Z>1.645ZN(2.80,1))=Φ(2.801.645)=Φ(1.155)0.876P(Z > 1.645 \mid Z \sim N(2.80, 1)) = \Phi(2.80 - 1.645) = \Phi(1.155) \approx 0.876. This confirms B is correct. A is wrong because equal nominal size does not imply equal power. The two tests have the same α\alpha, but different critical values and rejection regions — power depends on where those regions sit relative to the alternative distribution, not just on α\alpha itself. C is wrong because power does not equal 1α1 - \alpha. That relationship holds for the null distribution, not the alternative. Concentrating α\alpha in one tail raises power compared to the two-sided test, but not to 0.95. D is wrong because the lower rejection region (Z<1.96Z < -1.96) contributes negligibly to power at δ=2.80\delta = 2.80 — the probability mass of N(2.80,1)N(2.80, 1) below 1.96-1.96 is essentially zero. Removing it costs almost nothing, while the lower critical value (1.645 vs. 1.96) gains power substantially. Study tip: Always compute power as Φ(δzcritical)\Phi(\delta - z_{\text{critical}}) for an upper-tailed test. A smaller critical value means higher power — switching from two-sided to one-sided (in the correct direction) always increases power.

Question 8

In a regular parametric model, a Wald statistic tests rr smooth restrictions. Under the null, the statistic converges to χr2\chi_r^2. Consider local alternatives of the form θn=θ0+h/n\theta_n=\theta_0+h/\sqrt n, and let IeffI_{\mathrm{eff}} denote the efficient information matrix for the tested parameter after accounting for nuisance parameters.

Which expression gives the limiting power of the level-α\alpha Wald test under these local alternatives?

  1. P ⁣(χr2>χr,1α2),P\!\left(\chi_r^2>\chi_{r,1-\alpha}^2\right), because local alternatives have the same limiting law as the null.
  2. P ⁣(χr2(nhTIeffh)>χr,1α2),P\!\left(\chi_r^2(nh^{\mathsf T}I_{\mathrm{eff}}h)>\chi_{r,1-\alpha}^2\right), with a noncentrality parameter growing linearly in nn.
  3. P ⁣(N(hTIeffh,1)>z1α),P\!\left(N(h^{\mathsf T}I_{\mathrm{eff}}h,1)>z_{1-\alpha}\right), because every Wald statistic is asymptotically normal.
  4. P ⁣(χr2(hTIeffh)>χr,1α2),P\!\left(\chi_r^2(h^{\mathsf T}I_{\mathrm{eff}}h)>\chi_{r,1-\alpha}^2\right), using a fixed noncentrality parameter. (correct answer)
Explanation: When analyzing Wald tests under local alternatives, your core tool is the noncentral chi-squared distribution. The key insight is how the noncentrality parameter scales with sample size under sequences of the form θn=θ0+h/n\theta_n = \theta_0 + h/\sqrt{n}. Under such local alternatives, the Wald statistic converges in distribution to a noncentral χr2\chi^2_r with a fixed noncentrality parameter. Here's why: the shift h/nh/\sqrt{n} is precisely calibrated so that the signal grows at rate n\sqrt{n} — exactly fast enough to shift the limiting distribution, but not so fast that power goes to 1. Specifically, the noncentrality parameter is λ=hTIeffh\lambda = h^\mathsf{T} I_{\mathrm{eff}} h, a fixed scalar determined by the direction and magnitude of the local drift and the curvature of the log-likelihood. The limiting power is therefore P ⁣(χr2(hTIeffh)>χr,1α2)P\!\left(\chi^2_r(h^\mathsf{T} I_{\mathrm{eff}} h) > \chi^2_{r,1-\alpha}\right), confirming D. A is wrong because it treats local alternatives as indistinguishable from the null — but the whole point of 1/n1/\sqrt{n} alternatives is to produce a nontrivial limiting power strictly between α\alpha and 1. B is tempting but wrong: it inserts a factor of nn into the noncentrality parameter. That would correspond to a fixed alternative (not local), where power converges to 1. C incorrectly reduces the χr2\chi^2_r to a normal distribution, which only applies when r=1r = 1 and even then conflates the chi-squared and normal frameworks. A useful memory anchor: local = fixed noncentrality. If you ever see nn multiplying the noncentrality parameter in a local-alternative context, that's a red flag — it signals confusion between fixed and local alternatives.

Question 9

Consider the normal linear regression model Yi=β0+β1xi+εiY_i=\beta_0+\beta_1x_i+\varepsilon_i with independent errors εiN(0,σ2)\varepsilon_i\sim N(0,\sigma^2) and unknown σ2\sigma^2. A two-sided level-α\alpha t-test is used for H0:β1=0H_0:\beta_1=0. The investigator adds several observations whose predictor values all equal the mean of the original predictor values.

Holding the true nonzero slope and error variance fixed, what is the principal effect of these added observations on size and power?

  1. The exact size remains α\alpha; the slope noncentrality is unchanged, but increased residual degrees of freedom generally increase power. (correct answer)
  2. The exact size remains α\alpha; the slope noncentrality increases because every added observation increases the predictor sum of squares.
  3. The size falls below α\alpha; the slope noncentrality is unchanged, so power necessarily remains exactly constant.
  4. The size rises above α\alpha; the added observations estimate only the intercept and invalidate the usual t distribution.
Explanation: When analyzing how added observations affect inference in simple linear regression, focus on two separate questions: (1) does the test maintain its nominal size, and (2) how does the noncentrality parameter change? The t-statistic for β^1\hat{\beta}_1 follows an exact tn2t_{n-2} distribution whenever the classical normal linear model holds — regardless of where new observations fall in predictor space. Adding observations at xˉ\bar{x} does nothing to violate normality, independence, or homoscedasticity, so the test remains exactly level α\alpha. This confirms the first claim in answer A. Now consider power. The noncentrality parameter for the slope test is λ=β1Sxx/σ\lambda = \beta_1 \sqrt{S_{xx}}/\sigma, where Sxx=(xixˉ)2S_{xx} = \sum(x_i - \bar{x})^2. Observations added at exactly xˉ\bar{x} contribute zero to SxxS_{xx}, so the noncentrality is completely unchanged. However, these observations do increase the residual degrees of freedom from n2n-2 to (n+k)2(n+k)-2, which makes the critical value from the t-distribution smaller in magnitude. A smaller critical threshold with the same noncentrality means modestly higher power, which is precisely what A states. Answer B is wrong because it claims every added observation increases SxxS_{xx} — false when observations land at xˉ\bar{x}. Answer C incorrectly asserts size falls below α\alpha; the t-distribution remains exact, and "unchanged noncentrality" does not freeze power when degrees of freedom grow. Answer D is wrong because the t-distribution is still valid — the test procedure is not invalidated by observations at xˉ\bar{x}. Study tip: Always decompose power into two components — the noncentrality parameter (driven by SxxS_{xx}) and the critical value (driven by degrees of freedom). They can move independently.

Question 10

Let XBinomial(4,p)X\sim\operatorname{Binomial}(4,p). To test H0:p0.5H_0:p\leq 0.5 against H1:p>0.5H_1:p>0.5, a randomized test rejects when X=4X=4 with probability 0.80.8 and otherwise does not reject.

What are, respectively, the size of this test and its Type II error probability at p=0.7p=0.7?

  1. The size is 0.05000.0500, and the Type II error probability is 0.80790.8079. (correct answer)
  2. The size is 0.05000.0500, and the Type II error probability is 0.75990.7599.
  3. The size is 0.06250.0625, and the Type II error probability is 0.75990.7599.
  4. The size is 0.10000.1000, and the Type II error probability is 0.51980.5198.
Explanation: When you encounter a randomized test problem, your two key tasks are: (1) find the size by maximizing the power function over H0H_0, and (2) compute the Type II error (failing to reject when H1H_1 is true) at the specified alternative. The power function here is β(p)=0.8P(X=4)=0.8p4\beta(p) = 0.8 \cdot P(X=4) = 0.8 \cdot p^4. Since this is strictly increasing in pp, the supremum over H0:p0.5H_0: p \leq 0.5 is achieved at p=0.5p = 0.5. So the size is 0.8(0.5)4=0.80.0625=0.05000.8 \cdot (0.5)^4 = 0.8 \cdot 0.0625 = 0.0500. Now for the Type II error at p=0.7p = 0.7: this is 1β(0.7)=10.8(0.7)4=10.80.2401=10.1921=0.80791 - \beta(0.7) = 1 - 0.8 \cdot (0.7)^4 = 1 - 0.8 \cdot 0.2401 = 1 - 0.1921 = 0.8079. This confirms answer A. Answer B gets the size right (0.05000.0500) but reports the wrong Type II error — 0.75990.7599 corresponds to 10.8(0.7)41 - 0.8 \cdot (0.7)^4 computed incorrectly, likely confusing P(X=4p=0.7)P(X=4 \mid p=0.7) with something else. Answer C uses 0.06250.0625 as the size, which is just P(X=4p=0.5)=(0.5)4P(X=4 \mid p=0.5) = (0.5)^4 — forgetting to multiply by the randomization probability 0.80.8. Answer D inflates both quantities, suggesting the student may have used an entirely different rejection region or misread the randomization rule. The key study tip: for randomized tests, always incorporate the randomization probability into every calculation — size, power, and Type II error all flow through ϕ(x)P(X=xp)\phi(x) \cdot P(X=x \mid p). Forgetting that multiplier is the most common trap on these problems.