Statistics Graduate Level Quiz: Permutation Tests
10 questions · exam conditions
0:00
Permutation TestsQuestion 1 of 10

A matched-pairs experiment contains twelve pairs. Within each pair, one unit was randomly assigned treatment and the other control. Investigators use the mean of the twelve treated-minus-control within-pair differences as their statistic.

Under the sharp null hypothesis of no treatment effect, which resampling scheme correctly reflects the assignment design?

Randomly divide all twenty-four outcomes into two groups of twelve, ignoring the original matching.
Independently reverse the treatment and control labels within each pair with probability 1/21/2.
Bootstrap the twelve observed paired differences and center each resample at the original sample mean.
Apply one common random sign to all twelve paired differences, choosing either sign with probability 1/21/2.
← Back to quizzes

Statistics Graduate Level Quiz

Statistics Graduate Level Quiz: Permutation Tests

Practice Permutation Tests in Statistics Graduate Level with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Permutation Tests, giving you a quick way to practice the rules, question types, and explanations that matter most for Statistics Graduate Level.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A matched-pairs experiment contains twelve pairs. Within each pair, one unit was randomly assigned treatment and the other control. Investigators use the mean of the twelve treated-minus-control within-pair differences as their statistic.

Under the sharp null hypothesis of no treatment effect, which resampling scheme correctly reflects the assignment design?

  1. Randomly divide all twenty-four outcomes into two groups of twelve, ignoring the original matching.
  2. Independently reverse the treatment and control labels within each pair with probability 1/21/2. (correct answer)
  3. Bootstrap the twelve observed paired differences and center each resample at the original sample mean.
  4. Apply one common random sign to all twelve paired differences, choosing either sign with probability 1/21/2.
Explanation: When conducting a randomization (permutation) test, the resampling scheme must mirror the exact random mechanism used in the original assignment — nothing more, nothing less. In a matched-pairs design, randomization happened within each pair independently: a fair coin determined which unit got treatment and which got control. Under the sharp null (no treatment effect whatsoever), the two outcomes within each pair are exchangeable, meaning flipping the labels inside any pair is equally likely and valid. This is precisely what B does: for each of the twelve pairs, independently flip the sign of the paired difference with probability 1/21/2. Over all 212=40962^{12} = 4096 equally likely sign combinations, you reconstruct the full randomization distribution of the test statistic — exactly as the design dictates. A is wrong because pooling all 24 outcomes and reshuffling ignores the pairing structure entirely. This destroys the within-pair correlation that matching was designed to exploit, producing an invalid reference distribution. C is wrong because bootstrapping is an inference tool for estimating sampling variability, not for conducting a randomization test. Bootstrapping doesn't respect the original assignment mechanism and introduces resampling noise unrelated to the null distribution. D applies a single sign to all twelve differences simultaneously — equivalent to just two possible outcomes (+all, −all) instead of 2122^{12}. This collapses the reference distribution to two points, making it nearly useless and completely misrepresenting the independent within-pair randomizations. As a study tip: always ask yourself, "What was the actual random mechanism?" and build your resampling scheme to replicate it precisely — pair by pair, unit by unit, as independence requires.

Question 2

Consider the linear model Y=Zγ+Xβ+εY=Z\gamma+X\beta+\varepsilon, where ZZ contains nuisance covariates and the errors are assumed identically distributed and exchangeable under the null hypothesis H0:β=0H_0:\beta=0. The column XX is correlated with columns of ZZ.

Which procedure best describes the Freedman-Lane approach to testing H0H_0?

  1. Permute XX across observations, refit the full model to each permuted dataset, and record the resulting test statistic for β\beta.
  2. Fit the full model to the observed data, permute its residuals, add them to the full-model fitted values, and retest β\beta in the augmented response.
  3. Fit the reduced model using only ZZ, permute its residuals, add them to the reduced-model fitted values to form a null response, and refit the full model containing both ZZ and XX. (correct answer)
  4. Regress both YY and XX on ZZ, then independently permute the two sets of residuals across observations and compute their association.
Explanation: When testing a coefficient of interest in the presence of nuisance covariates, the key challenge is preserving the null distribution while respecting the correlation structure between XX and ZZ. Simply permuting XX or raw residuals can distort that structure and produce invalid tests. The Freedman-Lane procedure handles this elegantly. First, fit the reduced model Y=Zγ+εY = Z\gamma + \varepsilon and collect its residuals ε^\hat{\varepsilon}. These residuals represent variation in YY unexplained by ZZ — exactly the component that should be exchangeable under H0:β=0H_0: \beta = 0. Permute those residuals, then add them back to the reduced-model fitted values Y^Z=Zγ^\hat{Y}_Z = Z\hat{\gamma} to construct a synthetic null response Y=Zγ^+π(ε^)Y^* = Z\hat{\gamma} + \pi(\hat{\varepsilon}). Finally, refit the full model with both ZZ and XX to YY^* and record the test statistic. Repeating this yields the reference distribution under the null. This is answer C. Answer A permutes XX directly, which destroys its correlation with ZZ and inflates the Type I error rate — the permuted XX no longer reflects the covariate structure observed in the data. Answer B permutes residuals from the full model, which already absorbs the signal from β\beta. This tends to be overly conservative and does not correctly isolate the null component. Answer D describes the Manly or residualization approach — permuting residuals of both YY and XX after partialing out ZZ — a different (and less standard) strategy that conflates two separate residualization steps. Remember: Freedman-Lane always uses reduced-model residuals added back to reduced-model fitted values, then refits the full model — that three-step rhythm is the signature of this method.

Question 3

Two independent populations have the same mean but different variances and otherwise need not have identical distributions. Sample sizes from both populations increase with a limiting ratio bounded away from zero and infinity. An analyst considers permuting group labels to test equality of means.

Which statement most accurately compares a raw difference-in-means statistic with a studentized difference-in-means statistic?

  1. Both permutation tests are finite-sample exact because equality of means makes observations exchangeable across groups.
  2. Only the raw statistic is asymptotically valid because studentization introduces random variance estimates after permutation.
  3. Neither statistic can be asymptotically valid unless the two population distributions are identical in every respect.
  4. The studentized statistic can be asymptotically valid, whereas the raw-statistic permutation test can have incorrect size. (correct answer)
Explanation: When you see a permutation test applied under unequal variances, your first instinct should be to ask: what does permuting labels actually assume? A permutation test is exact only when observations are exchangeable under the null — meaning the joint distribution is truly invariant to label swaps. Equal means alone does not guarantee exchangeability when variances differ. The raw difference-in-means statistic, Xˉ1Xˉ2\bar{X}_1 - \bar{X}_2, is the correct choice to scrutinize here. Under permutation, you're redistributing observations as if they came from a single common population. But if the two groups have different variances, this redistribution produces a null distribution that doesn't match the true sampling distribution of the statistic — leading to inflated or deflated Type I error rates. The test loses its size guarantee asymptotically. The studentized statistic, T=Xˉ1Xˉ2σ^12/n1+σ^22/n2,T = \frac{\bar{X}_1 - \bar{X}_2}{\sqrt{\hat{\sigma}_1^2/n_1 + \hat{\sigma}_2^2/n_2}}, rescales by group-specific variance estimates. A key theoretical result (Chung & Romano, and related work) shows that permutation of the studentized statistic is asymptotically valid even under unequal variances, because the studentization accounts for the heteroskedasticity and the permutation null distribution converges correctly. This makes D the right answer. A is wrong because equal means alone doesn't imply exchangeability — you need identical distributions for that. B has it exactly backwards: studentization fixes the heteroskedasticity problem rather than breaking it. C is overly restrictive; studentization allows asymptotic validity without requiring identical distributions. Study tip: On permutation test questions, always ask whether the null implies full exchangeability. When it doesn't (e.g., unequal variances), look for studentization as the asymptotic fix.

Question 4

A randomized experiment records five outcomes. For every treatment reassignment, analysts calculate five studentized treatment statistics and use the maximum absolute statistic, M=maxjTjM=\max_j |T_j|, to adjust the individual outcome tests. They wish to claim strong familywise error control, including configurations in which some outcome null hypotheses are false.

What additional condition most directly supports that strong-control claim for a single-step max-statistic permutation procedure?

  1. Subset pivotality, so the joint null distribution for any subset of true null hypotheses is unaffected by which other null hypotheses happen to be false. (correct answer)
  2. Pairwise independence of the five outcomes, so the joint distribution of the five statistics factors into a product of marginal distributions under any partial null.
  3. Equality of all five observed treatment effect estimates, so every endpoint contributes symmetrically to the permutation maximum.
  4. Marginal normality of each studentized statistic, so the distribution of the permutation maximum converges to a known parametric form.
Explanation: When a multiple-testing procedure must control the familywise error rate (FWER) strongly — meaning under any configuration of true and false null hypotheses, not just the global null — the key question is whether the permutation distribution of your test statistic remains valid even when some hypotheses are false. This is exactly what subset pivotality guarantees. Subset pivotality states that the joint distribution of the test statistics corresponding to the true null hypotheses is the same regardless of which other hypotheses are false. For a max-statistic permutation procedure, this means the permutation distribution of M=maxjTjM = \max_j |T_j| over the true nulls isn't distorted by treatment effects on the false nulls. Without it, false nulls could contaminate the permutation reference distribution, making your critical value invalid and destroying strong FWER control. With it, you can apply a single-step procedure across all endpoints and trust that the familywise rejection threshold is simultaneously valid for every partial null configuration. Answer A captures this precisely. Answer B is wrong because pairwise independence is far too strong a requirement and also unnecessary — permutation procedures exploit the joint dependence structure naturally. Factoring into marginal distributions isn't what you need. Answer C is wrong because equality of effect estimates is not a theoretical requirement for FWER control; it's a symmetry convenience that has no bearing on validity under partial nulls. Answer D is wrong because permutation procedures are nonparametric by design — you never need marginal normality or convergence to a parametric form; the permutation distribution is empirically derived. Study tip: Whenever you see "strong FWER control" with a max-statistic, immediately think subset pivotality — it's the standard sufficient condition in the Westfall–Young framework and a frequent exam target.

Question 5

A policy was implemented in one city on a date selected uniformly at random from twenty prespecified eligible dates. Daily outcomes are serially correlated. The statistic compares the average outcome during the ten days after implementation with the average during the ten days before implementation. Under the sharp null, the policy changes no daily outcome.

Which procedure most faithfully constructs the design-based randomization distribution while retaining the serial structure of the observed outcomes?

  1. Randomly shuffle all daily outcomes, keep the observed implementation date fixed, and recompute the statistic.
  2. Randomly permute residuals from an intercept-only model, keep the observed implementation date fixed, and recompute.
  3. Randomly reorder the ten pre-policy and ten post-policy outcomes separately around the observed date.
  4. Evaluate the statistic at each eligible implementation date and weight the values uniformly over those dates. (correct answer)
Explanation: Whenever you see a question about randomization-based inference, anchor yourself to one core principle: the randomization distribution must mirror exactly how randomization was actually conducted in the study design — nothing more, nothing less. Here, the implementation date was drawn uniformly from 20 prespecified eligible dates. That single randomization act is the only source of design-based uncertainty. Under the sharp null, every daily outcome is fixed — what varies is only which date was chosen. The correct procedure, then, is to ask: "What statistic value would we have observed had each of the other 19 eligible dates been chosen instead?" Answer D does precisely this — it evaluates the before/after difference statistic at every eligible date using the unchanged observed outcome series, then weights all 20 values equally. This perfectly reconstructs the design's randomization distribution and naturally preserves serial correlation because the outcome sequence is never touched. Answer A is fatally flawed because it shuffles the outcome values themselves, destroying the serial dependence structure and inventing a distribution that was never part of the study design. Answer B similarly manipulates the outcomes (via residual permutation), which alters the temporal autocorrelation pattern and departs from the actual randomization mechanism. Answer C reorders outcomes within the pre- and post-windows separately, which also breaks the serial structure and, more critically, does not correspond to any step in the original randomization — the date was randomized, not the outcomes around a fixed date. The key study tip: in design-based inference, ask "what did the randomization actually vary?" Then build your reference distribution by re-running the statistic over those randomized assignments only, leaving outcomes untouched.

Question 6

A trial randomizes fourteen clinics, seven to an intervention and seven to control. Clinic sizes vary substantially, and outcomes are recorded for individual patients. Investigators use a regression-adjusted individual-level treatment coefficient as the statistic.

Under the sharp null, which randomization procedure is appropriate for inference based on the trial's design?

  1. Permute patient-level treatment indicators freely while preserving the total number of treated patients.
  2. Permute outcomes independently within clinics while holding every clinic's treatment assignment fixed.
  3. Permute intervention labels among clinics, preserving seven treated clinics, and recompute the regression statistic. (correct answer)
  4. Permute clinic labels among patients, preserving the observed intervention label attached to each patient.
Explanation: When a trial randomizes clusters (here, clinics) rather than individuals, the unit of randomization governs what a valid permutation test must mimic. The sharp null says the intervention had no effect on any outcome — but to test it honestly, you must rerandomize the data exactly as the original randomization was conducted. Since the investigators randomized clinics to treatment or control (not individual patients), the permutation distribution must be built by shuffling treatment labels among clinics, not among patients. Choice C is correct because it respects this structure: permute the intervention label across the fourteen clinics (keeping seven treated), recompute the regression-adjusted statistic each time, and compare the observed statistic to that reference distribution. This faithfully replicates how chance operated in the actual design. Choice A is wrong because it permutes treatment at the patient level, ignoring the cluster structure entirely. This inflates the effective sample size and produces a reference distribution that is far too narrow — leading to anticonservative inference. Patients within the same clinic share common exposures and are not exchangeable across clinics under the design. Choice B is wrong in a different way: permuting outcomes within clinics while holding clinic assignments fixed tests nothing about the treatment effect — it destroys only within-clinic variation while preserving the clinic-level treatment contrast that actually matters. Choice D is wrong because shuffling clinic labels among patients scrambles cluster membership, not treatment assignment, and doesn't correspond to any step in the randomization process. The key strategy: always match the permutation scheme to the unit of randomization. In cluster-randomized trials, that unit is the cluster — full stop.

Question 7

To approximate a permutation null distribution, an analyst generates B=99B=99 random permutations without forcing the observed arrangement to be among them. Exactly four permuted statistics are at least as extreme as the observed statistic. The permutations are sampled independently from the null randomization distribution.

Which reported Monte Carlo permutation value uses the standard finite-simulation correction that avoids assigning a zero value and yields a valid randomized-test calibration?

  1. 4/994/99, because only the simulated permutations belong in the denominator.
  2. 4/1004/100, because the observed arrangement is added only to the denominator.
  3. 5/1005/100, because one is added to both the extreme count and total count. (correct answer)
  4. 5/995/99, because the observed arrangement is added only to the extreme count.
Explanation: Whenever you see a question about permutation tests and Monte Carlo p-values, focus on a foundational principle: the observed test statistic itself represents one valid realization under the null hypothesis — namely, the original (unpermuted) data arrangement. Ignoring it creates a p-value that can equal zero, which is statistically invalid and philosophically incoherent (it would claim the data are impossible under the null). The standard finite-simulation correction, formalized by Davison & Hinkley and others, defines the Monte Carlo p-value as p^=b+1B+1\hat{p} = \frac{b + 1}{B + 1}, where bb is the number of simulated permutation statistics at least as extreme as the observed statistic, and BB is the number of simulated permutations. Here, b=4b = 4 and B=99B = 99, giving 4+199+1=5100=0.05\frac{4+1}{99+1} = \frac{5}{100} = 0.05. This is answer C, and it's correct because adding 1 to both numerator and denominator effectively includes the observed arrangement as an additional "permutation" that trivially equals itself — guaranteeing p^1/(B+1)>0\hat{p} \geq 1/(B+1) > 0. Answer A (4/994/99) omits the observed arrangement entirely, risks underestimating the true p-value, and can produce zero in other scenarios. Answer B (4/1004/100) adds 1 only to the denominator, inflating it without crediting the observed statistic as an extreme case, which underestimates the p-value. Answer D (5/995/99) adds 1 only to the numerator while keeping the old denominator, which overcounts relative to the actual reference set and yields an inconsistent estimator. Your study tip: memorize the formula p^=(b+1)/(B+1)\hat{p} = (b+1)/(B+1) and remember the justification — both counts shift by 1 because the observed data is one legitimate member of the full permutation reference set.

Question 8

In an experiment with eight units, exactly three units receive treatment. Because of ethical constraints, the design does not select the treatment set uniformly: each admissible set has a known probability determined by the units' baseline risk scores. The observed statistic is the treated-minus-control difference in mean outcomes.

Which procedure gives the appropriate randomization-test null distribution for the sharp null hypothesis of no treatment effect on any unit?

  1. Enumerate admissible treatment sets and weight each statistic by that set's probability under the actual assignment design. (correct answer)
  2. Enumerate admissible treatment sets and assign them equal weight because each set contains exactly three treated units.
  3. Permute the observed outcomes uniformly while holding the realized treatment indicators and baseline risk scores fixed.
  4. Resample units with replacement within the realized treatment groups and compare the resulting differences in means.
Explanation: Whenever you see a question about randomization (permutation) tests, anchor yourself to this core principle: the null distribution must mirror the actual probability mechanism used to assign treatments. The sharp null hypothesis of no effect means every unit's outcome is fixed regardless of assignment — so the test statistic's distribution comes entirely from re-randomizing according to the true design. Here, the design is non-uniform: different treatment sets carry different probabilities based on baseline risk scores. Under the sharp null, the correct null distribution weights each possible treatment set's statistic by the probability that set would actually be selected — exactly what A prescribes. This preserves the integrity of the design and produces a valid pp-value as p=zZ1[T(z)Tobs]P(Z=z).p = \sum_{\mathbf{z} \in \mathcal{Z}} \mathbf{1}[T(\mathbf{z}) \geq T_{\text{obs}}] \cdot P(\mathbf{Z} = \mathbf{z}). B is the classic trap for non-uniform designs. Equal weighting is only valid when the assignment mechanism is itself uniform — collapsing all sets to equal weight discards the actual design and yields a test that is neither valid nor properly sized. C confuses randomization inference with something closer to a label-shuffling permutation test. Holding treatment indicators fixed while permuting outcomes doesn't enumerate the reference distribution generated by the assignment mechanism, and it conflates the roles of outcomes and assignments. D describes a bootstrap procedure, which relies on sampling assumptions and large-sample approximations. Randomization tests make no distributional assumptions — they derive exact finite-sample distributions from the design itself, so resampling with replacement is conceptually mismatched here. The key study takeaway: always ask "what was the actual assignment mechanism?" and let that — not convenience or symmetry — define your reference distribution.

Question 9

An observational study compares an exposed and an unexposed group. Exposure probabilities differ substantially by a binary baseline variable XX, and the outcome distribution also depends on XX. Within each level of XX, investigators are willing to assume that exposure labels are exchangeable under the null of no exposure-outcome association conditional on XX.

Which permutation strategy most directly uses the stated exchangeability assumption?

  1. Permute exposure labels across all subjects while preserving only the total number of exposed subjects.
  2. Permute outcomes across all subjects while keeping exposure labels and the observed values of XX fixed.
  3. Permute exposure labels separately within each level of XX, preserving each stratum's exposed count. (correct answer)
  4. Permute the values of XX within exposure groups, preserving both groups' observed outcome distributions.
Explanation: When a question mentions exchangeability conditional on a covariate, your first instinct should be to ask: "Where exactly does the randomization argument apply?" Conditional exchangeability means that within strata defined by XX, the exposed and unexposed units are interchangeable under the null — not across the entire sample. This is the core of stratified permutation testing. Choice C is correct because it operationalizes the assumption precisely: by permuting exposure labels separately within each level of XX, you respect the fact that XX is a confounder that must be held fixed. Within each stratum, the exposed count is preserved, reflecting the conditional null hypothesis that exposure assignment within that stratum is arbitrary. The resulting reference distribution accounts for XX's influence on both exposure and outcome. Choice A is the classic trap: permuting labels across all subjects ignores the confounding structure of XX. Because exposure probability varies substantially by XX, unrestricted permutation mixes strata and produces a reference distribution contaminated by XX's effect, inflating or deflating the test statistic under the null. Choice B permutes outcomes rather than exposure labels, which addresses a different kind of exchangeability — it would be relevant if you were testing whether the outcome distribution differs by XX, not whether exposure causes the outcome. Choice D permutes XX itself, which destroys the covariate structure entirely and is unrelated to any standard inference strategy described in the passage. The study tip: whenever you see "exchangeability conditional on XX," that phrase is a direct signal to stratify your permutation by XX. Unconditional permutation tests ignore confounders; conditional (stratified) permutation tests control for them by design.

Question 10

In a completely randomized experiment, investigators perform a Fisher randomization test by imputing every missing potential outcome under the hypothesis that treatment changes no unit's outcome. The test rejects at level 0.050.05. Treatment effects could differ across units.

Which conclusion is justified by this result?

  1. The null of zero average treatment effect is rejected, because it is equivalent to the imputed sharp null.
  2. The sharp null is rejected, but a zero average treatment effect may still be compatible with heterogeneous effects. (correct answer)
  3. Every unit must have a nonzero treatment effect, although the signs of those effects may differ.
  4. The sample average treatment effect must be nonzero, but no conclusion about the sharp null is available.
Explanation: Whenever you see a question involving Fisher's randomization test, your first instinct should be to identify exactly which null hypothesis is being tested — because the conclusion is strictly limited to that hypothesis. Fisher's randomization test evaluates the sharp null hypothesis: that every unit's treatment effect is exactly zero, i.e., Yi(1)=Yi(0)Y_i(1) = Y_i(0) for all ii. Under this hypothesis, all potential outcomes are fully observed (the treated units' control outcomes are imputed as equal to what was observed), making the test statistic's exact distribution computable via rerandomization. When the test rejects at level 0.05, you have evidence against this specific sharp null — nothing more, nothing less. B is correct because it precisely captures this logic. Rejecting the sharp null tells you it is not the case that every unit has a zero effect. However, a zero average treatment effect (ATE) is still compatible with the data: some units could have large positive effects and others large negative effects, canceling out to an average of zero. Heterogeneity in effects is the key nuance. A is wrong because the sharp null and the null of zero average treatment effect are not equivalent. The sharp null is strictly stronger — it implies zero ATE, but zero ATE does not imply the sharp null. C is wrong because rejecting the sharp null doesn't mean every unit has a nonzero effect; it only means not all effects are simultaneously zero. D is wrong because the sharp null rejection does carry information — it is the sharp null, not the ATE null, that is directly tested. Study tip: Always distinguish the sharp null (unit-level) from the weak null (average-level) — these appear repeatedly in causal inference questions and require different testing frameworks.