Home

Tutoring

Subjects

Live Classes

Study Coach

Essay Review

On-Demand Courses

Colleges

Games


Sign up

Log in

Opening subject page...

Loading your content

Practice

  • All Subjects
  • Algebra Flashcards
  • SAT Math Practice Tests
  • Math Question of the Day
  • Live Classes
  • On-Demand Courses

Varsity Tutors

  • Find a Tutor
  • Test Prep
  • Online Classes
  • K-12 Learning
  • College Search
  • VarsityTutors.com

© 2026 Varsity Tutors. All rights reserved.

← Back to quizzes

USMLE Step 1 Quiz

USMLE Step 1 Quiz: Statistical Inference

Practice Statistical Inference in USMLE Step 1 with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

Question 1 / 20

0 of 20 answered

Researchers investigate the effect of a new statin on LDL cholesterol levels. In a clinical trial, the mean difference in LDL reduction between the statin group and the placebo group was 25 mg/dL. The 95% confidence interval for this mean difference was calculated to be (15 mg/dL, 35 mg/dL).

Which of the following is the most accurate interpretation of this 95% confidence interval?

Select an answer to continue

What this quiz covers

This quiz focuses on Statistical Inference, giving you a quick way to practice the rules, question types, and explanations that matter most for USMLE Step 1.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Researchers investigate the effect of a new statin on LDL cholesterol levels. In a clinical trial, the mean difference in LDL reduction between the statin group and the placebo group was 25 mg/dL. The 95% confidence interval for this mean difference was calculated to be (15 mg/dL, 35 mg/dL).

Which of the following is the most accurate interpretation of this 95% confidence interval?

  1. There is a 95% probability that the true mean difference in LDL reduction lies between 15 and 35 mg/dL.
  2. If the study were repeated many times, 95% of the calculated confidence intervals would contain the true mean difference. (correct answer)
  3. 95% of patients in the study had an LDL reduction between 15 and 35 mg/dL.
  4. There is a 5% chance that the new statin is ineffective.

Explanation: A 95% confidence interval provides a range of plausible values for the true population parameter. The correct frequentist interpretation is that if the same study were conducted 100 times, 95 of the resulting confidence intervals would be expected to contain the true mean difference in the population.

Question 2

A study is designed to determine if a new educational intervention improves medical student scores on a standardized exam. The researchers set the significance level (α) to 0.05. After the intervention, they find that the intervention group scores significantly higher than the control group (p=0.04) and conclude the intervention is effective. However, in reality, the intervention has no true effect on exam scores.

In this scenario, the researchers have made which of the following types of error?

  1. Type I error (correct answer)
  2. Type II error
  3. Selection bias
  4. Confounding

Explanation: A Type I error occurs when the null hypothesis is incorrectly rejected. The null hypothesis states there is no difference between the groups. Here, the researchers concluded there was a difference (rejected the null) when in fact no true difference existed. This is the definition of a Type I error (false positive). The probability of making a Type I error is denoted by α.

Question 3

A pharmaceutical company develops a new medication intended to prevent migraines. A clinical trial is conducted, but the results show no statistically significant difference in migraine frequency between the medication group and the placebo group (p = 0.15). The company decides not to pursue further development. Later, several larger studies confirm that the medication is, in fact, modestly effective.

The initial study's failure to detect a real effect is an example of which of the following?

  1. Type I error
  2. Type II error (correct answer)
  3. Recall bias
  4. Lead-time bias

Explanation: A Type II error occurs when one fails to reject a null hypothesis that is actually false. In this case, the initial study failed to find a significant effect (did not reject the null hypothesis of no difference) when a true effect existed. This is the definition of a Type II error (false negative). The probability of making a Type II error is denoted by β.

Question 4

A public health study reports the annual income of residents in a community with a large academic medical center. The data show that a few highly paid surgeons earn substantially more than the majority of residents, who have modest incomes. The distribution of income is highly skewed.

Which of the following measures of central tendency would be the most appropriate to describe the typical income in this community?

  1. Mean
  2. Median (correct answer)
  3. Mode
  4. Range

Explanation: In a skewed distribution, the mean is heavily influenced by extreme values (outliers). The median represents the 50th percentile and is resistant to outliers. Therefore, for skewed data such as income, the median provides a more accurate representation of the central or 'typical' value than the mean.

Question 5

A research team is planning a study to compare a new drug for diabetes with a standard treatment. They want to ensure their study has a high probability of detecting a clinically meaningful difference in HbA1c levels if one truly exists. They aim for a study power of 80%.

Which of the following is the most effective way for the researchers to increase the power of their study?

  1. Increase the p-value threshold for significance (e.g., from 0.05 to 0.10)
  2. Decrease the number of participants in the study
  3. Increase the number of participants in the study (correct answer)
  4. Enroll a more heterogeneous patient population

Explanation: Power is the ability of a study to detect a true effect (1 - β). The power of a study is primarily influenced by the sample size, effect size, and alpha level. Increasing the sample size is the most common and effective method to increase statistical power, as it reduces the standard error and makes it easier to detect a true difference between groups.

Question 6

A study is published on the weights of newborns in a specific hospital. The mean weight is 3400 grams, with a standard deviation of 400 grams. A second publication reports on the mean birth weight from 100 different hospitals, giving a mean of 3400 grams and a standard error of the mean of 40 grams.

Which of the following best describes the standard error of the mean (SEM)?

  1. It measures the variability of individual newborn weights around the population mean.
  2. It is an estimate of the standard deviation of the distribution of sample means. (correct answer)
  3. It is always larger than the standard deviation.
  4. It describes the 95% range for the individual measurements.

Explanation: The standard deviation (SD) measures the variability or spread of individual data points within a single sample. The standard error of the mean (SEM) estimates the variability of the means of multiple samples taken from the same population. It quantifies how precisely the sample mean estimates the true population mean and is calculated as SD / √n.

Question 7

A survey asks physicians to report the number of hours they sleep per night. The results are plotted, and the distribution is found to have a tail extending to the left. The peak of the distribution is at 8 hours, but a number of physicians report sleeping only 4 or 5 hours, pulling the average down.

Which of the following best describes the relationship between the measures of central tendency for this distribution?

  1. Mean = Median = Mode
  2. Mean < Median < Mode (correct answer)
  3. Mode < Median < Mean
  4. Mean = Mode, but different from Median

Explanation: This describes a negatively skewed (left-skewed) distribution. In such a distribution, the outliers are on the lower end. The mode is the most frequent value (the peak), which is 8 hours. The median is less affected by the low-value outliers than the mean. The mean is pulled downward by the low values. Therefore, the relationship is Mean < Median < Mode.

Question 8

Two different studies evaluate the same new surgical procedure. Study A has 50 participants, while Study B has 5000 participants. Both studies find the same point estimate for the reduction in recovery time and are free of bias. Assume all other factors are equal.

How would the 95% confidence interval (CI) for the mean reduction in recovery time in Study B compare to that of Study A?

  1. The CI in Study B would be wider.
  2. The CI in Study B would be narrower. (correct answer)
  3. The CI width would be identical in both studies.
  4. The CI in Study B would be shifted to the right.

Explanation: The width of a confidence interval is inversely related to the square root of the sample size. A larger sample size (like in Study B) leads to a smaller standard error of the mean, resulting in a more precise estimate of the true population parameter. This increased precision is reflected by a narrower confidence interval.

Question 9

A clinical trial compares a new cancer therapy to a standard therapy. The primary endpoint is 5-year survival. The results show a 5-year survival of 45% with the new therapy and 42% with the standard therapy. The p-value for this difference is 0.04. The researchers claim the new therapy is superior.

Which of the following statements represents the most critical consideration when interpreting this result?

  1. The result is not valid because the p-value is close to 0.05.
  2. A Type II error may have occurred.
  3. Statistical significance does not necessarily imply clinical significance. (correct answer)
  4. The study must have been underpowered.

Explanation: While the result is statistically significant (p < 0.05), the absolute difference in survival is only 3%. A critical step in interpreting research is to evaluate whether a statistically significant finding is also clinically meaningful. A small, clinically unimportant difference can become statistically significant if the sample size is very large. Clinicians must decide if a 3% survival benefit justifies the potential costs, side effects, and risks of the new therapy.

Question 10

Before starting a randomized controlled trial for a new drug, the investigators must specify the probability of making a Type I error that they are willing to accept. This threshold is used to determine statistical significance at the end of the study.

This pre-specified probability threshold is known as which of the following?

  1. The alpha (α) level (correct answer)
  2. The beta (β) level
  3. The power of the study
  4. The effect size

Explanation: The alpha (α) level, or significance level, is the pre-specified probability of committing a Type I error. It is the threshold below which a p-value is considered statistically significant. By convention, α is typically set to 0.05, meaning the researchers accept up to a 5% chance of incorrectly rejecting a true null hypothesis.

Question 11

Researchers are planning a cohort study to investigate the effect of a new lifestyle intervention on the incidence of type 2 diabetes. They anticipate that the effect of the intervention will be small. They want to ensure their study has adequate power to detect this small effect.

Besides increasing the sample size, which of the following would also increase the power of the study?

  1. Decreasing the expected effect size
  2. Increasing the significance level (α) from 0.05 to 0.10 (correct answer)
  3. Decreasing the significance level (α) from 0.05 to 0.01
  4. Enrolling participants with less variability in baseline characteristics

Explanation: Power (1-β) is the probability of correctly rejecting a false null hypothesis. There is a trade-off between Type I (α) and Type II (β) errors. By increasing the alpha level (e.g., from 0.05 to 0.10), the threshold for significance becomes less strict, making it easier to reject the null hypothesis. This reduces the probability of a Type II error (β) and therefore increases the power of the study, at the cost of increasing the risk of a Type I error.

Question 12

A clinical laboratory establishes a reference range for serum potassium. They measure the level in thousands of healthy volunteers and find that the values are normally distributed. The reference range is defined as the central 95% of these values.

This reference range corresponds to which of the following statistical intervals?

  1. The mean ± 1 standard deviation
  2. The mean ± 2 standard deviations (correct answer)
  3. The 95% confidence interval for the mean
  4. The interquartile range

Explanation: For a normally distributed variable, approximately 95% of all individual values lie within 2 standard deviations (more precisely, 1.96 SDs) of the population mean. This is the standard definition used for creating reference ranges for many laboratory tests. A 95% confidence interval describes the range for the population mean, not the range for individual values.

Question 13

A large pharmaceutical company conducts a well-designed, randomized trial to test a new cholesterol-lowering drug. The results show a reduction in LDL cholesterol that is not statistically significant (p = 0.25). The company shelves the drug. A junior researcher argues that because the trial had a relatively small sample size, a clinically important effect might have been missed.

The researcher is suggesting that the study may have resulted in which type of error due to insufficient power?

  1. Type I error
  2. Type II error (correct answer)
  3. Observer bias
  4. Ecologic fallacy

Explanation: The study failed to reject the null hypothesis (p > 0.05). The researcher's concern is that a true effect exists but was not detected. This scenario—failing to detect a real effect—is a Type II error. Studies with small sample sizes often lack sufficient statistical power to detect small or moderate effects, increasing the risk of a Type II error.

Question 14

A new rapid screening test for a certain viral infection is developed. In a trial, the test fails to identify the infection in a small number of patients who are later confirmed to have the disease by a gold-standard PCR test. The company is concerned about the consequences of these false negatives.

The failure of the test to detect a disease when it is truly present corresponds to which statistical concept?

  1. A Type I error (α)
  2. A Type II error (β) (correct answer)
  3. Low specificity
  4. Low positive predictive value

Explanation: In the context of diagnostic testing, the null hypothesis is that the patient does not have the disease. A false negative occurs when the test result is negative, but the disease is present. This is analogous to a Type II error: failing to reject the null hypothesis (of no disease) when it is false (disease is present). Low sensitivity, not specificity, is the test characteristic associated with high rates of false negatives.

Question 15

A randomized controlled trial is conducted to evaluate the efficacy of a new antihypertensive drug, Drug X, compared to a placebo. After 12 weeks, the mean reduction in systolic blood pressure (SBP) was 10 mmHg in the Drug X group and 2 mmHg in the placebo group. The difference in mean SBP reduction between the two groups was 8 mmHg. A statistical analysis yields a p-value of 0.03 for this difference.

Based on this p-value, which of the following is the most appropriate conclusion?

  1. There is a 3% probability that the observed difference is due to chance.
  2. There is a 3% probability that Drug X is effective.
  3. If the null hypothesis is true, there is a 3% probability of observing a difference of 8 mmHg or more. (correct answer)
  4. The study proves that Drug X is clinically superior to placebo.

Explanation: The p-value is the probability of observing a result at least as extreme as the one obtained, assuming the null hypothesis is true. In this case, the null hypothesis is that there is no difference in SBP reduction between Drug X and placebo. Therefore, a p-value of 0.03 means there is a 3% chance of seeing a difference of 8 mmHg or greater if the drug truly has no effect.

Question 16

The fasting glucose levels of a healthy adult population are known to be approximately normally distributed with a mean of 90 mg/dL and a standard deviation of 8 mg/dL.

Based on this information, approximately what percentage of this population would be expected to have a fasting glucose level between 74 mg/dL and 106 mg/dL?

  1. 34%
  2. 68%
  3. 95% (correct answer)
  4. 99.70%

Explanation: For a normal distribution, approximately 68% of the data falls within 1 standard deviation (SD) of the mean, 95% falls within 2 SDs, and 99.7% falls within 3 SDs. The range from 74 mg/dL to 106 mg/dL represents the mean (90) plus or minus 16 mg/dL. Since the SD is 8 mg/dL, this range corresponds to the mean ± 2 SDs (90 ± 2*8). Therefore, approximately 95% of the population falls within this range.

Question 17

A study measures the incubation period for a viral illness in a group of 1,000 infected individuals. The mean incubation period is 10 days, the median is 8 days, and the mode is 7 days. The distribution is asymmetrical.

Based on these measures of central tendency, what is the shape of this distribution?

  1. Normal
  2. Bimodal
  3. Positively skewed (right-skewed) (correct answer)
  4. Negatively skewed (left-skewed)

Explanation: In a positively skewed (right-skewed) distribution, there is a long tail of high values. These high values pull the mean to the right of the median. The mode is the most frequent value and is typically at the peak of the distribution, to the left of the median. The relationship Mode < Median < Mean (7 < 8 < 10) is characteristic of a positively skewed distribution.

Question 18

A researcher is studying serum cholesterol levels in a population of healthy adults. The data is collected and analyzed from 1000 participants. The distribution of cholesterol levels is found to be symmetrical and bell-shaped.

For this distribution, which of the following relationships between the mean, median, and mode is expected?

  1. Mean > Median > Mode
  2. Mean < Median < Mode
  3. Mean = Median = Mode (correct answer)
  4. Mean = Median, but Mode is different

Explanation: A symmetrical, bell-shaped distribution is characteristic of a normal distribution. In a perfect normal distribution, the data is symmetrically distributed around the center. As a result, the mean (the arithmetic average), the median (the middle value), and the mode (the most frequent value) are all equal and located at the center of the distribution.

Question 19

A case-control study investigated the association between daily consumption of a specific artificial sweetener and the risk of bladder cancer. The study found an odds ratio of 1.5. The 95% confidence interval for the odds ratio was (0.9, 2.5). The p-value was 0.08.

Which of the following is the best interpretation of these findings?

  1. The artificial sweetener is proven to cause bladder cancer.
  2. The results are statistically significant at the α = 0.05 level.
  3. The study demonstrates a protective effect of the sweetener.
  4. The association is not statistically significant because the confidence interval includes the null value. (correct answer)

Explanation: For an odds ratio or relative risk, the null value is 1.0, which indicates no association between the exposure and the outcome. Since the 95% confidence interval (0.9, 2.5) includes 1.0, the result is not statistically significant at the α = 0.05 level. This is consistent with the p-value of 0.08, which is greater than 0.05.

Question 20

A study reports a mean serum sodium level of 140 mEq/L in a sample of 100 patients. The standard deviation (SD) is 5 mEq/L. The researchers also report the 95% confidence interval for the true mean sodium level in the population from which the sample was drawn.

The calculation of this confidence interval is based on the point estimate (the sample mean) and which of the following measures of variability?

  1. Standard deviation (SD)
  2. Variance
  3. Standard error of the mean (SEM) (correct answer)
  4. Interquartile range (IQR)

Explanation: A confidence interval for a mean is calculated as: Sample Mean ± (Critical Value * SEM). The standard error of the mean (SEM) is used because the goal is to estimate the precision of the sample mean as an estimate of the true population mean. The SEM (calculated as SD/√n) quantifies this uncertainty. The SD measures variability of individual data points, not the precision of the mean.