College Statistics Quiz: Conditions For Inference
20 questions · exam conditions
0:00
Conditions For InferenceQuestion 1 of 20

A researcher plans to calculate a 90% confidence interval. The study involves collecting data from a simple random sample of individuals to estimate a population parameter. Which of the following conditions is necessary for a one-sample t-interval for a mean but is NOT necessary for a one-sample z-interval for a proportion?

The sample must be selected randomly from the population of interest.
The sample size nn must be less than 10% of the population size NN if sampling is done without replacement.
The number of successes and failures in the sample must both be at least 10.
The population distribution must be approximately Normal, or the sample size must be large enough for the Central Limit Theorem to apply.
← Back to quizzes

College Statistics Quiz

College Statistics Quiz: Conditions For Inference

Practice Conditions For Inference in College Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Conditions For Inference, giving you a quick way to practice the rules, question types, and explanations that matter most for College Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A researcher plans to calculate a 90% confidence interval. The study involves collecting data from a simple random sample of individuals to estimate a population parameter. Which of the following conditions is necessary for a one-sample t-interval for a mean but is NOT necessary for a one-sample z-interval for a proportion?

  1. The sample must be selected randomly from the population of interest.
  2. The sample size nn must be less than 10% of the population size NN if sampling is done without replacement.
  3. The number of successes and failures in the sample must both be at least 10.
  4. The population distribution must be approximately Normal, or the sample size must be large enough for the Central Limit Theorem to apply. (correct answer)
Explanation: This question asks to differentiate the conditions for a confidence interval for a mean versus a proportion.
  • D is the correct answer. The Normality condition—that the population is normal or the sample size is large (CLT)—pertains to the sampling distribution of the sample mean (xˉ\bar{x}) and is essential for a t-interval.
  • A (Randomness) and B (Independence/10% condition) are necessary for both types of intervals.
  • C (Large Counts Condition) is necessary for a z-interval for a proportion, not a t-interval for a mean.

Question 2

A researcher plans to calculate a 90% confidence interval. The study involves collecting data from a simple random sample of individuals to estimate a population parameter. Which of the following conditions is necessary for a one-sample t-interval for a mean but is NOT necessary for a one-sample z-interval for a proportion?

  1. The sample must be selected randomly from the population of interest.
  2. The sample size nn must be less than 10% of the population size NN if sampling is done without replacement.
  3. The number of successes and failures in the sample must both be at least 10.
  4. The population distribution must be approximately Normal, or the sample size must be large enough for the Central Limit Theorem to apply. (correct answer)
Explanation: This question asks to differentiate the conditions for a confidence interval for a mean versus a proportion.
  • D is the correct answer. The Normality condition—that the population is normal or the sample size is large (CLT)—pertains to the sampling distribution of the sample mean (xˉ\bar{x}) and is essential for a t-interval.
  • A (Randomness) and B (Independence/10% condition) are necessary for both types of intervals.
  • C (Large Counts Condition) is necessary for a z-interval for a proportion, not a t-interval for a mean.

Question 3

A researcher is preparing to estimate a population mean using a large-sample (n=100n=100) confidence interval. Which of the following statements correctly distinguishes between assumptions about the population and the sampling distribution?

  1. For the confidence interval to be valid, the population distribution must be approximately normal, which in turn ensures the sampling distribution is normal.
  2. The Central Limit Theorem requires that the sample data itself be approximately normal, which allows us to infer that the population is normal.
  3. The randomness condition ensures that the population is normal, while the large sample size ensures that the sampling distribution is normal.
  4. Because the sample size is large, the sampling distribution of the sample mean is approximately normal, regardless of the shape of the population distribution. (correct answer)
Explanation: When you encounter confidence interval questions, focus on distinguishing between what we need from the population versus what the Central Limit Theorem guarantees about the sampling distribution. The Central Limit Theorem is the key here. For large samples (typically n30n \geq 30, and certainly for n=100n = 100), this theorem tells us that the sampling distribution of xˉ\bar{x} will be approximately normal, regardless of the population's shape. This is what makes large-sample confidence intervals work—we don't need to worry about whether the original population is normal, skewed, or even bimodal. Choice A incorrectly suggests we need the population to be normal first. This reverses the logic—the Central Limit Theorem actually frees us from this requirement when samples are large. Choice B confuses the direction of inference and misunderstands what needs to be normal. We don't need the sample data itself to be normal, nor can we conclude the population is normal from sample normality. Choice C incorrectly links randomness to population normality. Randomness ensures our sample is representative and observations are independent, but it doesn't change the population's shape. While large samples do help with sampling distribution normality, the statement wrongly implies randomness creates population normality. Choice D correctly captures the essence of the Central Limit Theorem: large sample size creates an approximately normal sampling distribution regardless of population shape. Remember this distinction: the Central Limit Theorem is about the sampling distribution's shape, not the population's shape. Large samples make confidence intervals robust against non-normal populations.

Question 4

A biologist is studying the lengths of a certain species of fish. A random sample of 25 fish is collected. A histogram of the sample lengths shows a slight left skew, with no outliers. The biologist wishes to construct a 95% confidence interval for the mean length. Which is the most appropriate course of action?

  1. Do not construct a t-interval because the sample data show evidence of non-normality, violating a key condition.
  2. Proceed with constructing the t-interval because t-procedures are known to be robust to mild deviations from normality, especially without outliers. (correct answer)
  3. Use a z-interval instead of a t-interval, as this is more appropriate for non-normal data.
  4. Increase the sample size until the sample histogram appears perfectly symmetric before constructing any interval.
Explanation: This question addresses the robustness of t-procedures. The t-interval for a mean is reasonably accurate even when the population distribution is not perfectly normal, as long as the sample size is not too small and the data do not contain strong skewness or outliers. A sample size of 25 is considered moderately large enough for the procedure to be robust to a slight skew. Therefore, it is appropriate to proceed.
  • A is an overly strict interpretation of the normality condition and ignores the concept of robustness.
  • C is incorrect; z-intervals also require normality and are only used when the population standard deviation is known.
  • D is impractical and unnecessary. Sample data will rarely be perfectly symmetric.

Question 5

An economist is studying the distribution of annual income in a country, which is known to be strongly skewed to the right. The economist plans to draw a random sample of individuals and construct a confidence interval for the mean annual income. To satisfy the conditions for inference, the sampling distribution of the sample mean must be approximately normal. Which statement regarding the required sample size is most accurate?

  1. A sample size of n=30n=30 is sufficient to ensure the sampling distribution is approximately normal, due to the Central Limit Theorem.
  2. The sampling distribution of the sample mean will be approximately normal for any sample size because the sample will be drawn randomly.
  3. A sample size substantially larger than 30 may be required because the population distribution is strongly skewed. (correct answer)
  4. The shape of the sampling distribution will be the same as the shape of the population distribution, so it will be skewed regardless of sample size.
Explanation: The Central Limit Theorem (CLT) states that the sampling distribution of the sample mean will become approximately normal as the sample size increases. The guideline of n30n \ge 30 is sufficient for many populations that are not heavily skewed. However, if the underlying population distribution is strongly skewed, as income distributions typically are, a much larger sample size is needed for the sampling distribution of the mean to become sufficiently normal for inference.
  • A overstates the n30n \ge 30 rule, which is a guideline, not an absolute law.
  • B incorrectly links random sampling to the shape of the sampling distribution.
  • D incorrectly states that the sampling distribution's shape mirrors the population's shape regardless of nn; this contradicts the CLT.

Question 6

A university has 2,500 students enrolled in an introductory psychology course. The department head wants to estimate the proportion of these students who are first-year students. A simple random sample of 200 students is selected from the course roster.

In verifying the assumptions for a one-proportion z-interval, which of the following conditions has been met?

  1. The Normality condition, because the sample size of 200 is greater than 30.
  2. The Independence condition, because the sample was selected randomly.
  3. The Randomness condition, because some students were sampled from the population.
  4. The 10% condition, because the sample size is less than 10% of the population size. (correct answer)
Explanation: When working with confidence intervals for proportions, you must verify several key assumptions before proceeding. These conditions ensure that the normal approximation to the binomial distribution is valid and that your sample accurately represents the population. The correct answer is D because the 10% condition requires that your sample size be less than 10% of the population size to ensure independence between observations. Here, the sample of 200 students represents 2002500=0.08=8%\frac{200}{2500} = 0.08 = 8\% of the total enrollment, which satisfies this condition. Let's examine why the other options are incorrect. Option A misunderstands the normality condition - the "greater than 30" rule applies to means, not proportions. For proportions, you need to check that both np10np \geq 10 and n(1p)10n(1-p) \geq 10, but you can't verify this without knowing the actual proportion. Option B incorrectly describes the independence condition. While random sampling helps, the independence condition specifically refers to the 10% rule mentioned in D. Option C confuses randomness with proper random sampling - simply selecting "some students" doesn't guarantee the systematic random selection needed for valid inference. Remember that proportion problems have three main conditions: randomness (proper random sampling was used), normality (np10np \geq 10 and n(1p)10n(1-p) \geq 10), and independence (sample size < 10% of population). The 10% condition is often the easiest to verify immediately since you typically know both sample and population sizes from the problem setup.

Question 7

A political polling organization surveys a simple random sample of 400 voters from a small town with only 1,500 registered voters. The 10% condition for independence is not met since 400>0.10×1500=150400 > 0.10 \times 1500 = 150. If the standard formula for the margin of error is used without a finite population correction, how will this calculated value compare to the true margin of error?

  1. It will be an overestimate of the true margin of error. (correct answer)
  2. It will be an underestimate of the true margin of error.
  3. It will be an unbiased estimate, but the confidence level will be incorrect.
  4. The comparison depends on the value of the sample proportion p^\hat{p}.
Explanation: When sampling a large fraction (more than 10%) of a finite population without replacement, the variability in the sample is actually less than what is predicted by the standard error formula, which assumes independence. The standard formula for standard error (e.g., p^(1p^)/n\sqrt{\hat{p}(1-\hat{p})/n}) does not account for this reduction in variability. The true standard error should be smaller. Therefore, using the standard formula results in an artificially large standard error, which in turn leads to an overestimate of the true margin of error.
  • B is incorrect; this is the opposite effect.
  • C is incorrect because the estimate of the margin of error is biased (systematically too large).
  • D is incorrect because the direction of the error from violating the 10% condition does not depend on p^\hat{p}.

Question 8

An online news outlet posts an article about climate change and includes a poll at the end asking, "Do you believe human activity is the primary cause of global warming?" After a week, 150,000 people have responded, with 65% selecting "Yes." The outlet reports, "Based on our poll, we are 95% confident that the true proportion of adults who believe in human-caused climate change is between 64.8% and 65.2%."

Why is this confidence interval report misleading?

  1. The Large Counts Condition is not met because the proportion is far from 0.5.
  2. The sample size of 150,000 is too large, which artificially shrinks the margin of error to be nearly zero.
  3. The sample was not a random sample of the population, leading to a high likelihood of bias that is not accounted for by the margin of error. (correct answer)
  4. The 10% condition is violated, as 150,000 is likely less than 10% of the adult population.
Explanation: The most significant flaw in this report is the use of a voluntary response sample. Participants self-selected to respond, meaning they are likely not representative of the entire adult population. They may have stronger opinions on the topic than the average person. The margin of error in a confidence interval only accounts for random sampling variability, not for biases introduced by a flawed sampling method. Therefore, the interval, despite its precision, is not a reliable estimate for the target population.
  • A is incorrect; the Large Counts condition is easily met with such a large sample.
  • B is incorrect; a large sample size is statistically good because it reduces sampling error. The small margin of error is a correct calculation, but it's applied to a biased sample.
  • D misinterprets the 10% condition; the condition is violated if n>0.10Nn > 0.10N. Here, nn is almost certainly far less than 10% of the population, so the condition is met.

Question 9

A pharmaceutical company is testing a new drug for a rare disease. In a random sample of 200 patients with the disease, only 5 experience a significant positive response. The company wishes to construct a 95% confidence interval for the proportion of all patients who would experience a positive response. Which of the following is the primary consequence of the failure to meet the Large Counts Condition in this scenario?

  1. The true proportion pp will not be contained within the calculated confidence interval.
  2. The sampling distribution of the sample proportion p^\hat{p} is not well-approximated by a Normal distribution, making the z-critical value inappropriate. (correct answer)
  3. The standard error of the proportion cannot be calculated because the sample size is too small.
  4. The sample proportion p^=0.025\hat{p} = 0.025 is not an unbiased estimator of the population proportion pp.
Explanation: The Large Counts Condition (np10np \ge 10 and n(1p)10n(1-p) \ge 10, checked with p^\hat{p}) justifies using a Normal model to approximate the sampling distribution of p^\hat{p}. In this case, the number of successes is np^=5n\hat{p} = 5, which is less than 10. Because this condition fails, the Normal approximation is not accurate. Using a z-critical value from the standard Normal distribution to calculate the margin of error is therefore inappropriate and will likely lead to an interval with an actual confidence level different from the stated 95%.
  • A is incorrect because any single interval may or may not contain the true proportion by chance; the issue is that the method does not have a 95% success rate in the long run.
  • C is incorrect because the formula for standard error can still be used to produce a number; the problem is with the distributional assumption used to create the margin of error.
  • D is incorrect because p^\hat{p} is an unbiased estimator of pp regardless of sample size or the number of successes.

Question 10

To estimate the average number of hours per week that students at a large university spend studying, a researcher stands outside the main library on a Wednesday afternoon and invites the first 100 students who walk by to participate in a survey.

The researcher plans to use the collected data to construct a 95% confidence interval for the mean study time of all students at the university. Which of the following inferential conditions is most critically violated by this study design?

  1. The Normality condition, because the sample size of 100 may not be large enough if the population of study times is heavily skewed.
  2. The Independence condition, because the students were sampled without replacement from the university population.
  3. The Randomness condition, because the sampling method used was not a simple random sample. (correct answer)
  4. The sample size condition, because a sample of 100 is too small relative to the total number of students at a large university.
Explanation: The most critical violation is the failure to obtain a random sample. This is a convenience sample, which is not representative of the entire student population. Students who are at the library on a Wednesday afternoon may study more than the average student, leading to a biased sample. This violation of the Randomness condition invalidates any attempt to generalize the findings to the entire university population.
  • A is a potential but less critical issue; a sample size of 100 is often sufficient for the CLT to apply.
  • B is incorrect because a sample of 100 from a large university would easily satisfy the 10% condition.
  • D is incorrect because a sample of 100 can be perfectly adequate for statistical inference if it is selected randomly.

Question 11

A researcher conducts an experiment to compare the effectiveness of two different fertilizers on tomato plant yield. They obtain 50 tomato seedlings from a local nursery. The seedlings are randomly assigned to two groups of 25: one group receives Fertilizer A, and the other receives Fertilizer B. After a growing season, the mean yield is calculated for each group. The researcher wishes to construct a two-sample t-interval for the difference in mean yields. What is the primary limitation of the resulting confidence interval in terms of generalizing conclusions?

  1. The experiment lacks random assignment, so the Normality condition for the difference in sample means cannot be assumed.
  2. The sample sizes (n=25) are too small to use a t-interval, as the Central Limit Theorem does not apply.
  3. Since the seedlings were not randomly sampled from a larger population of seedlings, inference about a broader population is not justified. (correct answer)
  4. The independence condition between the two groups is violated because both groups of seedlings were grown in the same experimental garden.
Explanation: The study design features random assignment, which allows the researcher to make causal conclusions about the effect of the fertilizers on these specific 50 plants. However, it does not feature random sampling from a larger population of interest (e.g., all tomato seedlings of this variety). Without random sampling, the Randomness condition for inference to a larger population is not met. Therefore, generalizing the results beyond the 50 plants in the experiment is not statistically justified.
  • A is incorrect because the experiment explicitly uses random assignment.
  • B is a possible issue, but the lack of random sampling is a more fundamental limitation for generalization.
  • D is incorrect because random assignment ensures independence between the two treatment groups.

Question 12

To estimate the average number of pages in a novel on a library's fiction shelf, a librarian decides to use systematic sampling. The shelf contains 500 books arranged in a particular order. The librarian randomly chooses the 5th book, then selects every 10th book thereafter (15th, 25th, etc.) until a sample of 50 books is obtained.

When checking the conditions for constructing a one-sample t-interval for the mean number of pages using formulas based on a simple random sample, which assumption is violated by this sampling plan?

  1. The Normality condition, because the number of pages might be skewed.
  2. The Random condition, because not all possible samples of size 50 had an equal chance of being selected. (correct answer)
  3. The Independence condition, because the sample size is at the boundary of the 10% condition.
  4. The Independence condition, due to a possible periodic pattern in the book arrangement.
Explanation: Standard inference procedures for a one-sample t-interval assume the data come from a simple random sample (SRS), where every individual has an equal chance of being selected, and every possible sample of size nn has an equal chance of being selected. In systematic sampling, once the first element is chosen, the entire sample is determined. This means that many possible samples (e.g., the 6th and 7th books) have zero chance of being selected together. Therefore, it is not an SRS, and the Random condition is not met.
  • A is a statement about the data's distribution, not a flaw in the sampling plan itself.
  • C is incorrect because the 10% condition is technically met: 500.10(500)=5050 \le 0.10(500) = 50.
  • D is a potential problem with systematic sampling if a pattern exists, but B describes the more fundamental violation of the SRS assumption that underlies the standard formulas.

Question 13

An engineer is testing the tensile strength of a new alloy. The manufacturing process is designed to produce rods with strengths that are normally distributed around a target mean. The engineer takes a random sample of 8 rods and plans to calculate a 99% t-interval for the mean strength. Which statement correctly assesses the conditions for this inference?

  1. The conditions are not met because the sample size (n=8n=8) is less than 30, so the Central Limit Theorem cannot be applied.
  2. The conditions are met because the population is stated to be normally distributed, making the t-procedure valid even for a very small sample. (correct answer)
  3. The conditions are not met because the sample size is too small to reliably check for outliers or skewness from the sample data.
  4. The conditions are met, but a z-interval should be used instead of a t-interval because the population is Normal.
Explanation: The Normality condition for a t-interval can be met in one of two ways: 1) the population distribution is normal, or 2) the sample size is large enough (CLT). In this problem, we are explicitly told that the population of strengths is normally distributed. This satisfies the condition, and the t-procedure is valid regardless of the small sample size.
  • A is incorrect because the CLT is not needed when the population is already normal.
  • C is incorrect because we are given the population distribution shape, so we do not need to check it using the sample data.
  • D is incorrect because a t-interval is used when the population standard deviation σ\sigma is unknown and estimated by the sample standard deviation ss, which is the case here. The normality of the population does not change this.

Question 14

A school district administrator wants to estimate the proportion of high school students who have a part-time job. They survey all students in 4 randomly selected classrooms out of 60 total classrooms in the district's largest high school. In the surveyed classes, 40 out of 110 students have a part-time job.

The administrator wants to construct a 95% confidence interval for the proportion of all students in the high school who have a part-time job using a standard formula for a simple random sample. Which of the following conditions for that procedure has not been met?

  1. The Random condition is met, because the classrooms were randomly selected.
  2. The Large Counts condition is met, because the number of successes (40) and failures (70) are both greater than 10.
  3. The Independence condition is not met, because students within a classroom are not independent observations. (correct answer)
  4. All conditions for a simple random sample have been met, and it is appropriate to proceed.
Explanation: This study uses cluster sampling, where classrooms are the clusters. Standard confidence interval formulas assume a simple random sample (SRS), where all observations are independent. In a cluster sample, observations within a cluster (classroom) are often more similar to each other than they are to observations in other clusters. This violates the independence assumption. For example, students in the same class might have similar schedules or be of a similar age, influencing their job status. Therefore, the standard formulas, which assume independence, are not appropriate.
  • A is misleading. While there was randomness, it was not an SRS of students, which is what the procedure assumes.
  • B is a correct statement (40 > 10, 70 > 10), but it's not the violated condition.
  • D is incorrect because the independence assumption is violated.

Question 15

A political scientist compares the mean age of voters in two different districts. They take a simple random sample of 40 voters from District A and a separate simple random sample of 50 voters from District B. To construct a two-sample t-interval for the difference in mean ages (μAμB\mu_A - \mu_B), which of the following is a required condition?

  1. The samples must be independent of each other. (correct answer)
  2. The two sample sizes must be equal.
  3. The data from both samples, when combined, must appear to come from a normal distribution.
  4. The population standard deviations for both districts must be known.
Explanation: When you encounter questions about two-sample t-intervals, you're dealing with inference procedures that compare means from two independent populations. The key is identifying which conditions must be satisfied for the procedure to be valid. For a two-sample t-interval to work properly, the samples must be independent of each other (A). This means the selection of voters in District A cannot influence which voters are selected in District B, and the age of any voter in one sample doesn't affect the ages in the other sample. Independence ensures that the sampling distributions behave as the mathematical theory predicts, making your confidence interval reliable. Let's examine why the other options are incorrect. Choice (B) suggests equal sample sizes are required, but two-sample t-procedures are specifically designed to handle unequal sample sizes—that's why we see different formulas for the degrees of freedom when n1n2n_1 \neq n_2. Choice (C) incorrectly focuses on the combined data's normality. What actually matters is that each population is approximately normal, or that sample sizes are large enough for the Central Limit Theorem to apply (generally n30n \geq 30). Choice (D) claims you need known population standard deviations, but that would require a two-sample z-procedure instead—the whole point of using t-procedures is when population standard deviations are unknown. Remember this pattern: for any two-sample procedure, independence is non-negotiable, while other conditions like normality and equal variances have workarounds or can be relaxed under certain circumstances. Always check independence first when evaluating whether a two-sample procedure is appropriate.

Question 16

To estimate the proportion of defective items from a large shipment, a quality control inspector takes a random sample of 120 items and finds 9 defective items. A 95% confidence interval for the proportion of all defective items is to be constructed. The inspector notes that one of the conditions for inference is borderline, as np^=9n\hat{p} = 9. What is the most likely consequence of proceeding with a standard one-proportion z-interval?

  1. The true confidence level of the procedure will be slightly lower than the stated 95%. (correct answer)
  2. The calculated interval will be centered at the wrong value.
  3. The standard error will be zero, making the calculation impossible.
  4. The independence condition is violated, which is a more serious problem than the borderline normality check.
Explanation: When constructing confidence intervals for proportions, you need to check that the sampling distribution of p^\hat{p} is approximately normal. This requires both np^10n\hat{p} \geq 10 and n(1p^)10n(1-\hat{p}) \geq 10. Here, np^=9n\hat{p} = 9, which falls just short of the recommended threshold. When the normality condition is borderline or violated, the sampling distribution of p^\hat{p} doesn't follow a normal distribution as closely as assumed. The standard z-interval relies on critical values from the normal distribution (like z=1.96z^* = 1.96 for 95% confidence), but if the actual sampling distribution has different tail behavior than the normal distribution, these critical values won't capture the true proportion 95% of the time. The actual coverage will be somewhat less than the stated 95%. Choice A is correct because proceeding with borderline normality conditions typically results in slightly lower actual confidence levels than stated. Choice B is wrong because the interval will still be properly centered at p^=9/120=0.075\hat{p} = 9/120 = 0.075. The center doesn't change based on normality conditions. Choice C is incorrect because the standard error formula p^(1p^)n\sqrt{\frac{\hat{p}(1-\hat{p})}{n}} will still yield a positive value. With p^=0.075\hat{p} = 0.075, this calculation remains valid. Choice D misidentifies the issue. The problem specifically mentions the borderline normality condition (np^=9n\hat{p} = 9), not independence. Independence depends on sampling method and population size, which aren't indicated as problematic here. Study tip: Always check both np^n\hat{p} and n(1p^)n(1-\hat{p}) before using z-intervals. When conditions are borderline, expect slightly conservative results (lower actual confidence than stated).

Question 17

A student researcher wants to estimate the mean number of soft drinks consumed per week by students at their high school of 800 students. They post a survey link on their personal social media page and collect 75 responses. The distribution of the 75 responses is heavily right-skewed and contains several high outliers. The student then calculates a 95% t-interval.

Which of the following represents the most critical reason why the resulting confidence interval is not reliable?

  1. The sample size (n=75n=75) is too small to satisfy the Central Limit Theorem for a heavily skewed distribution.
  2. The 10% condition is violated because the sample size of 75 is close to 10% of the population size of 800.
  3. The t-procedure is not appropriate because the sample data show strong skewness and outliers.
  4. The use of a voluntary response survey means the sample is not random and is likely biased. (correct answer)
Explanation: When evaluating the reliability of a confidence interval, you need to check whether the fundamental conditions for statistical inference are met. The most critical requirement is having a representative sample from your population of interest. The correct answer is D because voluntary response sampling creates a fundamental flaw that cannot be corrected. When the researcher posted the survey on their personal social media, only students who follow them and choose to respond will participate. This creates selection bias - students who consume more (or fewer) soft drinks might be more likely to respond, making the sample unrepresentative of all 800 students. No statistical procedure can fix biased data. Choice A is incorrect because with n=75n=75, the Central Limit Theorem generally applies even for skewed distributions, though the skewness does raise some concern about normality assumptions. Choice B misapplies the 10% condition. This condition requires the sample to be less than 10% of the population to ensure independence when sampling without replacement. Here, 75/800=9.375%75/800 = 9.375\%, which actually satisfies the condition. Choice C identifies a real concern - the t-procedure assumes approximate normality, and heavy skewness with outliers violates this. However, this is less critical than having biased data because robust methods or transformations might address distributional issues. Remember this hierarchy: biased sampling trumps all other statistical concerns. Even perfect statistical techniques applied to biased data will produce misleading results. Always evaluate the sampling method first when assessing the reliability of statistical inference.

Question 18

A school district administrator wants to estimate the proportion of high school students who have a part-time job. They survey all students in 4 randomly selected classrooms out of 60 total classrooms in the district's largest high school. In the surveyed classes, 40 out of 110 students have a part-time job.

The administrator wants to construct a 95% confidence interval for the proportion of all students in the high school who have a part-time job using a standard formula for a simple random sample. Which of the following conditions for that procedure has not been met?

  1. The Random condition is met, because the classrooms were randomly selected.
  2. The Large Counts condition is met, because the number of successes (40) and failures (70) are both greater than 10.
  3. The Independence condition is not met, because students within a classroom are not independent observations. (correct answer)
  4. All conditions for a simple random sample have been met, and it is appropriate to proceed.
Explanation: This study uses cluster sampling, where classrooms are the clusters. Standard confidence interval formulas assume a simple random sample (SRS), where all observations are independent. In a cluster sample, observations within a cluster (classroom) are often more similar to each other than they are to observations in other clusters. This violates the independence assumption. For example, students in the same class might have similar schedules or be of a similar age, influencing their job status. Therefore, the standard formulas, which assume independence, are not appropriate.
  • A is misleading. While there was randomness, it was not an SRS of students, which is what the procedure assumes.
  • B is a correct statement (40 > 10, 70 > 10), but it's not the violated condition.
  • D is incorrect because the independence assumption is violated.

Question 19

An economist is studying the distribution of annual income in a country, which is known to be strongly skewed to the right. The economist plans to draw a random sample of individuals and construct a confidence interval for the mean annual income. To satisfy the conditions for inference, the sampling distribution of the sample mean must be approximately normal. Which statement regarding the required sample size is most accurate?

  1. A sample size of n=30n=30 is sufficient to ensure the sampling distribution is approximately normal, due to the Central Limit Theorem.
  2. The sampling distribution of the sample mean will be approximately normal for any sample size because the sample will be drawn randomly.
  3. A sample size substantially larger than 30 may be required because the population distribution is strongly skewed. (correct answer)
  4. The shape of the sampling distribution will be the same as the shape of the population distribution, so it will be skewed regardless of sample size.
Explanation: The Central Limit Theorem (CLT) states that the sampling distribution of the sample mean will become approximately normal as the sample size increases. The guideline of n30n \ge 30 is sufficient for many populations that are not heavily skewed. However, if the underlying population distribution is strongly skewed, as income distributions typically are, a much larger sample size is needed for the sampling distribution of the mean to become sufficiently normal for inference.
  • A overstates the n30n \ge 30 rule, which is a guideline, not an absolute law.
  • B incorrectly links random sampling to the shape of the sampling distribution.
  • D incorrectly states that the sampling distribution's shape mirrors the population's shape regardless of nn; this contradicts the CLT.

Question 20

To estimate the average number of pages in a novel on a library's fiction shelf, a librarian decides to use systematic sampling. The shelf contains 500 books arranged in a particular order. The librarian randomly chooses the 5th book, then selects every 10th book thereafter (15th, 25th, etc.) until a sample of 50 books is obtained.

When checking the conditions for constructing a one-sample t-interval for the mean number of pages using formulas based on a simple random sample, which assumption is violated by this sampling plan?

  1. The Normality condition, because the number of pages might be skewed.
  2. The Random condition, because not all possible samples of size 50 had an equal chance of being selected. (correct answer)
  3. The Independence condition, because the sample size is at the boundary of the 10% condition.
  4. The Independence condition, due to a possible periodic pattern in the book arrangement.
Explanation: Standard inference procedures for a one-sample t-interval assume the data come from a simple random sample (SRS), where every individual has an equal chance of being selected, and every possible sample of size nn has an equal chance of being selected. In systematic sampling, once the first element is chosen, the entire sample is determined. This means that many possible samples (e.g., the 6th and 7th books) have zero chance of being selected together. Therefore, it is not an SRS, and the Random condition is not met.
  • A is a statement about the data's distribution, not a flaw in the sampling plan itself.
  • C is incorrect because the 10% condition is technically met: 500.10(500)=5050 \le 0.10(500) = 50.
  • D is a potential problem with systematic sampling if a pattern exists, but B describes the more fundamental violation of the SRS assumption that underlies the standard formulas.