College Biology Quiz: Probability Statistics In Biology
13 questions · exam conditions
0:00
Probability Statistics In BiologyQuestion 1 of 13

A researcher conducts an experiment to test whether a new fertilizer increases plant growth. She randomly assigns 40 plants to two groups: 20 receive the new fertilizer and 20 receive standard fertilizer. After 30 days, she measures plant height and finds that plants with new fertilizer have a mean height of 25.3 cm (standard deviation = 3.1 cm) while plants with standard fertilizer have a mean height of 22.8 cm (standard deviation = 2.9 cm). The p-value for this comparison is 0.02. What is the most appropriate interpretation of these results?

The new fertilizer definitely causes increased plant growth because the p-value is less than 0.05
There is a 2% probability that the new fertilizer actually works to increase plant growth
The observed difference in mean heights would occur by chance alone in about 2% of similar experiments if the fertilizers were equally effective
The new fertilizer increases plant height by exactly 2.5 cm on average in all plant populations
There is a 98% probability that the new fertilizer is better than the standard fertilizer
← Back to quizzes

College Biology Quiz

College Biology Quiz: Probability Statistics In Biology

Practice Probability Statistics In Biology in College Biology with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Probability Statistics In Biology, giving you a quick way to practice the rules, question types, and explanations that matter most for College Biology.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A researcher conducts an experiment to test whether a new fertilizer increases plant growth. She randomly assigns 40 plants to two groups: 20 receive the new fertilizer and 20 receive standard fertilizer. After 30 days, she measures plant height and finds that plants with new fertilizer have a mean height of 25.3 cm (standard deviation = 3.1 cm) while plants with standard fertilizer have a mean height of 22.8 cm (standard deviation = 2.9 cm). The p-value for this comparison is 0.02. What is the most appropriate interpretation of these results?

  1. The new fertilizer definitely causes increased plant growth because the p-value is less than 0.05
  2. There is a 2% probability that the new fertilizer actually works to increase plant growth
  3. The observed difference in mean heights would occur by chance alone in about 2% of similar experiments if the fertilizers were equally effective (correct answer)
  4. The new fertilizer increases plant height by exactly 2.5 cm on average in all plant populations
  5. There is a 98% probability that the new fertilizer is better than the standard fertilizer
Explanation: When you encounter statistical hypothesis testing questions, focus on what the p-value actually measures rather than falling into common interpretation traps. The p-value of 0.02 tells us the probability of observing a difference in plant heights as large as 2.5 cm (25.3 - 22.8) or larger, assuming the null hypothesis is true—that both fertilizers are equally effective. In other words, if the fertilizers truly had no difference in effectiveness and we repeated this exact experiment many times, we'd see a difference this large or larger in about 2% of those trials purely due to random variation. Answer C correctly captures this interpretation: the observed difference would occur by chance alone in about 2% of similar experiments if the fertilizers were equally effective. Answer A is wrong because statistical significance doesn't prove causation—it only suggests the difference is unlikely due to chance. "Definitely causes" is too strong a claim. Answer B misinterprets the p-value as the probability that the treatment works, when it's actually the probability of seeing these results assuming no real difference exists. Answer D incorrectly treats the sample difference (2.5 cm) as a universal effect size that applies to all plant populations, ignoring that this is just one study with specific conditions. Remember: p-values measure the likelihood of your observed data given the null hypothesis, not the likelihood that your hypothesis is correct. This distinction appears frequently on biology exams when interpreting experimental results.

Question 2

In a genetics experiment, a researcher crosses two heterozygous plants (Aa × Aa) and observes 120 offspring. She expects a 3:1 phenotypic ratio based on Mendelian inheritance. The observed results are 85 dominant phenotype and 35 recessive phenotype. To test whether this deviation from expected results is statistically significant, she calculates a chi-square value of 1.48. Given that the critical value for chi-square with 1 degree of freedom at p = 0.05 is 3.84, what should she conclude?

  1. The deviation is statistically significant, so Mendelian inheritance is definitely not occurring in this cross
  2. The deviation is not statistically significant, providing evidence that the results are consistent with Mendelian inheritance (correct answer)
  3. The chi-square value is too low to draw any meaningful conclusions about inheritance patterns
  4. The observed ratio exactly matches Mendelian expectations because 85:35 simplifies to approximately 3:1
  5. The experiment should be repeated with exactly 100 offspring to get a perfect 3:1 ratio for proper analysis
Explanation: When you encounter chi-square problems in genetics, you're testing whether observed data deviates significantly from expected Mendelian ratios. The key is comparing your calculated chi-square value to the critical value to determine statistical significance. Let's work through this step-by-step. In an Aa × Aa cross, you expect a 3:1 phenotypic ratio. With 120 offspring, this means 90 dominant (¾ × 120) and 30 recessive (¼ × 120). The researcher observed 85 dominant and 35 recessive, giving a chi-square value of 1.48. Since 1.48 < 3.84 (the critical value), the deviation is NOT statistically significant. This means the observed results are consistent with random sampling variation from the expected 3:1 ratio. Option A is wrong because when chi-square is below the critical value, we cannot reject the null hypothesis (Mendelian inheritance). The deviation could easily be due to chance. Option C misunderstands chi-square interpretation—a low value is actually good news, indicating results match expectations within normal variation. Option D makes a mathematical error; 85:35 equals 2.43:1, which is noticeably different from 3:1, though not statistically significant. The correct answer is B because the chi-square test confirms that observed deviations fall within expected random variation. Study tip: Remember that in chi-square tests, values below the critical threshold support your null hypothesis (usually Mendelian inheritance). Higher values indicate significant deviation that suggests something other than simple Mendelian genetics is occurring.

Question 3

A study examines the relationship between exercise frequency and resting heart rate in college students. The correlation coefficient (r) between weekly exercise hours and resting heart rate is -0.68 with a p-value of 0.001. Based on this information, which conclusion is most appropriate?

  1. Exercise directly causes a reduction in resting heart rate, and this causal relationship is very strong
  2. There is a moderately strong negative association between exercise frequency and resting heart rate that is unlikely due to chance (correct answer)
  3. Exactly 68% of the variation in resting heart rate can be explained by exercise frequency
  4. Students who exercise more have resting heart rates that are 0.68 beats per minute lower than those who don't exercise
  5. The probability that exercise actually affects heart rate is 99.9% based on the p-value
Explanation: When you encounter correlation questions, remember that correlation coefficients tell you about association strength and direction, not causation or exact measurements. The correlation coefficient r = -0.68 indicates a moderately strong negative relationship (values closer to -1 or +1 are stronger). The negative sign means as exercise hours increase, resting heart rate tends to decrease. The p-value of 0.001 is much less than 0.05, indicating this relationship is statistically significant and unlikely due to random chance alone. Option B correctly interprets both pieces of information: there's a moderately strong negative association that's statistically significant. Option A incorrectly claims causation. Correlation never proves causation—there could be confounding variables (like overall fitness level, diet, or genetics) that influence both exercise habits and heart rate. The word "directly" also overstates the relationship strength. Option C confuses the correlation coefficient with the coefficient of determination. To find the percentage of variation explained, you'd calculate r2=(0.68)2=0.46r^2 = (-0.68)^2 = 0.46, meaning about 46% of variation is explained, not 68%. Option D treats the correlation coefficient as if it represents actual heart rate differences in beats per minute. Correlation coefficients are unitless measures of association strength—they don't tell you specific measurement differences between groups. Remember: correlation coefficients measure association strength and direction (ranging from -1 to +1), while p-values tell you statistical significance. Never interpret correlation as causation, and don't confuse r with r2r^2 or with actual measurement units.

Question 4

A population geneticist studies allele frequencies in a plant population over time. She finds that allele frequency changes follow a normal distribution with a mean change of 0 and standard deviation of 0.02 per generation due to genetic drift. What is the probability that a particular allele will increase in frequency by more than 0.03 in a single generation?

  1. Approximately 13.4%, because this represents the area under both tails of the normal curve beyond 1.5 standard deviations
  2. Approximately 6.7%, because this represents the area under one tail of the normal curve beyond 1.5 standard deviations (correct answer)
  3. Approximately 3%, because the change requested equals 1.5 times the standard deviation
  4. Approximately 93.3%, because most changes fall within 1.5 standard deviations of the mean
  5. Exactly 50%, because genetic drift causes random changes with equal probability in either direction
Explanation: When you encounter questions about genetic drift and probability distributions, you're dealing with statistical analysis of random changes in allele frequencies. Genetic drift causes random fluctuations that follow a normal distribution around zero change. To solve this problem, you need to standardize the value and find the appropriate tail probability. The question asks for the probability of an increase greater than 0.03, given a normal distribution with mean = 0 and standard deviation = 0.02. First, calculate the z-score: z=0.0300.02=1.5z = \frac{0.03 - 0}{0.02} = 1.5 Since we want P(X > 0.03), we need the area in the upper tail beyond z = 1.5. Using the standard normal distribution, approximately 6.7% of the distribution lies beyond 1.5 standard deviations in one tail. Choice A incorrectly includes both tails of the distribution. While 13.4% represents the combined area beyond ±1.5 standard deviations, the question specifically asks for increases greater than 0.03, which requires only the upper tail. Choice C confuses the standardized value (1.5) with a probability. The fact that 0.03 equals 1.5 standard deviations doesn't mean the probability is 3%. Choice D gives the area within 1.5 standard deviations (approximately 93.3%), but we need the area beyond this range in the upper tail specifically. Study tip: For normal distribution problems, always identify whether you need one tail or both tails, then standardize your value and use the appropriate z-table area. "Greater than" or "less than" typically means one tail, while "different from" suggests both tails.

Question 5

An ecologist studying bird populations collects data on nest success rates across different habitat types. She observes 45 successful nests out of 60 total nests in forest habitat, and 28 successful nests out of 50 total nests in grassland habitat. She wants to test whether nest success rates differ significantly between these habitats. What would be the most appropriate statistical approach?

  1. Chi-square test of independence, because she is comparing proportions between two categorical variables (correct answer)
  2. Two-sample t-test, because she is comparing means between two groups
  3. Correlation analysis, because she wants to determine the relationship between habitat type and nest success
  4. ANOVA, because she is comparing nest success across multiple habitat categories
  5. Regression analysis, because she wants to predict nest success based on habitat characteristics
Explanation: When analyzing biological data involving success/failure outcomes across different categories, you need to identify the type of data and the appropriate statistical test. Here, you're comparing nest success rates (a proportion) between two habitat types (categorical groups). Option A is correct because this is a classic chi-square test of independence scenario. You have two categorical variables: habitat type (forest vs. grassland) and nest outcome (successful vs. unsuccessful). The chi-square test determines whether the proportion of successful nests is independent of habitat type. Your data can be arranged in a 2×2 contingency table: Forest (45 success, 15 failure) vs. Grassland (28 success, 22 failure). Option B is wrong because a two-sample t-test compares means of continuous variables between groups, not proportions of categorical outcomes. You're not comparing average nest sizes or temperatures—you're comparing success rates. Option C is incorrect because correlation analysis examines linear relationships between two continuous variables. While you want to know if habitat and success are related, correlation isn't the right tool for categorical data analysis. Option D is wrong because ANOVA compares means across multiple groups with continuous dependent variables. Additionally, you only have two habitat types, not multiple categories, and your outcome variable is categorical (success/failure), not continuous. Study tip: When you see proportion or percentage data across categories, think chi-square. When you see means or averages being compared, think t-tests or ANOVA. The type of data (categorical vs. continuous) drives your statistical test choice.

Question 6

A researcher investigating antibiotic resistance tests 200 bacterial isolates from hospital patients. She finds that 45 isolates are resistant to antibiotic A, 38 are resistant to antibiotic B, and 12 are resistant to both antibiotics. If resistance to the two antibiotics were independent events, how many isolates would be expected to show resistance to both antibiotics?

  1. Approximately 8.6 isolates (correct answer)
  2. Approximately 17.1 isolates
  3. Approximately 41.5 isolates
  4. Approximately 21.5 isolates
  5. Exactly 12 isolates
Explanation: When you encounter questions about antibiotic resistance and independence, you're dealing with probability theory applied to microbiology. Independent events mean that resistance to one antibiotic doesn't influence resistance to another—a key assumption for calculating expected frequencies. To find the expected number of isolates resistant to both antibiotics under independence, you multiply the individual probabilities. First, calculate the probability of resistance to each antibiotic: antibiotic A affects 45/200 = 0.225 of isolates, and antibiotic B affects 38/200 = 0.19 of isolates. Under independence, the probability of resistance to both is 0.225 × 0.19 = 0.04275. Multiplying by the total sample size: 200 × 0.04275 = 8.55 isolates, which rounds to approximately 8.6. Choice A (8.6 isolates) correctly applies the independence calculation. Choice B (17.1 isolates) appears to incorrectly add the individual probabilities rather than multiply them—a common probability error. Choice C (41.5 isolates) seems to average the two resistance rates, which has no basis in probability theory. Choice D (21.5 isolates) might result from incorrectly using the arithmetic mean of the resistant isolates (45 + 38)/2 = 41.5, then halving it. The key insight here is that the observed dual resistance (12 isolates) is actually higher than expected under independence (8.6), suggesting the resistances may be linked—perhaps through plasmids carrying multiple resistance genes. Always remember: for independent events, multiply probabilities, don't add them.

Question 7

A conservation biologist monitors a population of endangered birds over 10 years, recording the number of breeding pairs each year. The data shows considerable year-to-year variation. To determine if there is a significant long-term trend in population size, she calculates the correlation between year and number of breeding pairs, obtaining r = -0.28 with p = 0.43. Based on these results, what is the most appropriate interpretation?

  1. There is a significant declining trend in the bird population that requires immediate conservation action
  2. The population is stable because the correlation is negative, indicating the decline has stopped
  3. No significant linear trend is detected, but this doesn't rule out other patterns or the need for continued monitoring (correct answer)
  4. The correlation is too weak to provide any useful information about population trends
  5. The population has increased by 28% over the 10-year period based on the correlation coefficient
Explanation: When analyzing population trends in conservation biology, you need to understand both correlation strength and statistical significance. A correlation coefficient (r) tells you the strength and direction of a linear relationship, while the p-value tells you whether that relationship is statistically significant. Here, r = -0.28 indicates a weak negative correlation between year and breeding pairs, suggesting a slight downward trend. However, the p-value of 0.43 is much greater than the standard significance threshold of 0.05, meaning this correlation could easily be due to random chance rather than a true population decline. Answer C is correct because it properly interprets both statistics: no significant linear trend was detected (p > 0.05), but the biologist should continue monitoring since population dynamics can follow non-linear patterns, and longer-term data might reveal significant trends not apparent in 10 years. Answer A incorrectly assumes statistical significance where none exists—you can't conclude there's a real declining trend when p = 0.43. Answer B misunderstands what a negative correlation means and incorrectly suggests the population is stable, when the data shows high year-to-year variation. Answer D dismisses weak correlations entirely, but even non-significant results provide valuable information about what patterns are NOT strongly present in your data. Remember: In conservation biology questions involving statistics, always check both the effect size (correlation coefficient) AND the p-value before drawing conclusions about population trends. Non-significant results still inform conservation decisions.

Question 8

An agricultural researcher tests whether organic fertilizer affects crop yield compared to synthetic fertilizer. She randomly assigns 60 plots to receive either organic (n=30) or synthetic (n=30) fertilizer and measures grain yield in kg per plot. The results show: organic mean = 245 kg (SD = 28 kg), synthetic mean = 267 kg (SD = 31 kg), with t = -2.85 and p = 0.006. Before concluding that synthetic fertilizer is superior, what important consideration should guide her interpretation?

  1. The difference could be due to genetic variation in the crop plants rather than fertilizer type
  2. Statistical significance doesn't necessarily indicate practical significance for farmers (correct answer)
  3. The sample size is too small to detect meaningful differences between fertilizer types
  4. The standard deviations are too similar, suggesting the measurements were imprecise
  5. A t-test is inappropriate for this type of agricultural data, so the results are invalid
Explanation: When interpreting research results, you need to distinguish between statistical significance (whether an effect exists) and practical significance (whether that effect matters in real-world applications). This study found a statistically significant difference with p = 0.006, but the researcher must consider whether this difference is meaningful for actual farming decisions. The correct answer is B because statistical significance alone doesn't guarantee practical importance. While synthetic fertilizer produced 22 kg more grain per plot on average, farmers must weigh this against factors like cost differences, environmental impact, and profit margins. If organic fertilizer costs significantly less or commands premium prices in the market, the 22 kg difference might not justify switching to synthetic fertilizer despite the statistical significance. Answer A is incorrect because the researcher used randomization, which controls for genetic variation by distributing it equally across both groups. Answer C is wrong because n=30 per group provides adequate power to detect meaningful differences—the significant p-value confirms this. Answer D misunderstands measurement precision; similar standard deviations (28 vs 31 kg) actually suggest consistent measurement methods, and the variability is reasonable for agricultural data. Remember that p-values only tell you whether a difference likely exists, not whether it's large enough to matter. On biology exams, look for questions that test this distinction—especially in applied contexts like agriculture, medicine, or conservation where practical significance determines real-world decisions. Always consider effect size and context alongside statistical significance.

Question 9

A medical researcher wants to test whether a new drug reduces blood pressure more effectively than the current standard treatment. She plans to recruit 100 patients and randomly assign them to two groups. To ensure her experimental design can detect a meaningful difference, she performs a power analysis and finds that with her planned sample size, the study has 80% power to detect a 10 mmHg difference between treatments. What does this power analysis tell her about her experimental design?

  1. There is an 80% probability that the new drug will reduce blood pressure by exactly 10 mmHg compared to standard treatment
  2. If the true difference between treatments is 10 mmHg, there is an 80% chance her study will detect a statistically significant effect (correct answer)
  3. Her study design has a 20% chance of making a Type I error (false positive) when testing for significance
  4. The study will definitely detect any difference of 10 mmHg or larger, but smaller differences might be missed
  5. 80% of patients in the study will experience a blood pressure reduction of at least 10 mmHg with the new drug
Explanation: Power analysis is a crucial tool in experimental design that helps researchers determine whether their study can reliably detect an effect if it truly exists. When you encounter power analysis questions, focus on what statistical power actually measures: the probability of detecting a real effect. Statistical power represents the probability that a study will detect a statistically significant effect when that effect actually exists in the population. In this case, 80% power means there's an 80% chance the study will find a statistically significant difference if the true difference between treatments is actually 10 mmHg. This makes option B correct. Let's examine why the other options are wrong. Option A misinterprets power as predicting the exact magnitude of effect the drug will have, but power analysis assumes a hypothetical effect size to calculate detection probability. Option C confuses power with Type I error rate (alpha level). The 20% represents the chance of a Type II error (false negative), not a Type I error (false positive), which is typically set at 5%. Option D incorrectly suggests the study will "definitely" detect effects of 10 mmHg or larger. Power gives probabilities, not guarantees—even with 80% power, there's still a 20% chance of missing a real 10 mmHg difference. Remember that statistical power increases with larger effect sizes, bigger sample sizes, and less variability in data. When studying experimental design, always distinguish between the different error types and what power actually measures versus what it assumes.

Question 10

A researcher studies the relationship between study time and exam scores in a biology class. She collects data from 50 students and finds a correlation coefficient of r = 0.45 with p = 0.001. However, she realizes that students who study more also tend to attend class more frequently, and class attendance might be the true factor affecting exam scores. This scenario best illustrates which important statistical concept?

  1. Type I error, where the researcher incorrectly rejected a true null hypothesis
  2. Confounding variables, where a third variable may explain the apparent relationship between study time and scores (correct answer)
  3. Sampling bias, where the students in the study are not representative of the broader population
  4. Measurement error, where study time was not accurately recorded by the students
  5. Statistical power, where the sample size was too small to detect the true relationship
Explanation: When analyzing relationships between variables in biological research, you must always consider whether other factors might be influencing the observed correlation. This scenario presents a classic case where two variables appear related, but a third variable may actually be driving both. The correct answer is B because this situation perfectly demonstrates confounding variables. A confounding variable is an external factor that influences both the independent variable (study time) and dependent variable (exam scores), potentially creating a false impression of causation. Here, class attendance could be the true driver - students who attend more classes might naturally study more AND perform better on exams, making it appear that study time directly improves scores when attendance is the real factor. Looking at the incorrect options: A is wrong because a Type I error involves rejecting a true null hypothesis, but the researcher hasn't made an error about statistical significance - she's questioning what the correlation actually means. C (sampling bias) is incorrect because there's no indication that the 50 students aren't representative of the class population. D (measurement error) doesn't fit because the issue isn't about inaccurate recording of study time, but rather about interpreting what the correlation represents. The strong correlation (r = 0.45) and high significance (p = 0.001) don't prove causation - they only show association. Remember this key principle for biology research: correlation never implies causation, especially when confounding variables haven't been controlled. Always ask yourself what other factors could explain an observed relationship before drawing conclusions.

Question 11

In an experiment testing the effect of light intensity on photosynthesis rate, a researcher measures oxygen production in aquatic plants. She tests 5 different light intensities with 6 replicates each. The data shows increasing oxygen production with increasing light intensity, but the relationship appears to level off at high intensities. To determine if a linear model adequately describes this relationship, she calculates R² = 0.73. What does this value indicate about the model's appropriateness?

  1. The linear model is excellent because 73% represents a strong correlation between variables
  2. The linear model explains 73% of the variation in oxygen production, but the leveling off suggests a non-linear model might be more appropriate (correct answer)
  3. The linear model is inappropriate because R² values must be above 0.80 to be considered statistically significant
  4. The relationship is perfectly linear for 73% of the data points, while 27% are outliers that should be removed
  5. The R² value indicates that light intensity causes 73% of the changes in photosynthesis rate
Explanation: When analyzing experimental data, understanding what R² means and how to interpret it in context is crucial for drawing valid conclusions about relationships between variables. R² represents the coefficient of determination - it tells you what proportion of the variation in your dependent variable (oxygen production) is explained by your independent variable (light intensity). An R² of 0.73 means the linear model explains 73% of the variation in oxygen production, which is actually quite good. However, the key insight here is that while this explains a substantial portion of the variation, the researcher observed that the relationship "levels off at high intensities." This plateau pattern is characteristic of biological processes that have limiting factors - exactly what you'd expect in photosynthesis where other factors like CO₂ concentration or temperature become limiting at high light intensities. Choice A incorrectly conflates R² with correlation strength without considering the biological context and observed plateau. Choice C is wrong because there's no universal threshold of 0.80 for R² to be "statistically significant" - significance depends on sample size, experimental design, and field of study. Choice D fundamentally misunderstands what R² represents; it doesn't mean 73% of data points follow a linear pattern while 27% are outliers. The correct answer is B because it properly interprets R² as explaining 73% of variation while recognizing that the leveling-off pattern suggests the true relationship is non-linear, making a curvilinear model (like exponential or logarithmic) potentially more appropriate. Remember: Always consider both statistical measures AND biological patterns when evaluating model appropriateness in experimental data.

Question 12

A genetics counselor analyzes family history data to assess disease risk. In a study of 500 families, she finds that 85 families have a history of both diabetes and heart disease, 120 families have diabetes history only, 95 families have heart disease history only, and 200 families have neither condition. She wants to test whether family history of diabetes and heart disease occur independently. What would be the expected frequency of families with both conditions under the null hypothesis of independence?

  1. Approximately 61.6 families
  2. Approximately 73.8 families (correct answer)
  3. Exactly 85 families
  4. Approximately 42.5 families
  5. Approximately 107.5 families
Explanation: When you encounter questions about testing independence between two categorical variables, you're working with chi-square analysis concepts. The key is calculating expected frequencies under the assumption that the variables are independent. To find the expected frequency of families with both conditions, you first need the marginal totals. Families with diabetes history: 85 + 120 = 205. Families with heart disease history: 85 + 95 = 180. Total families: 500. Under independence, the probability of having both conditions equals the product of individual probabilities: P(diabetes and heart disease)=P(diabetes)×P(heart disease)=205500×180500=0.41×0.36=0.1476P(\text{diabetes and heart disease}) = P(\text{diabetes}) \times P(\text{heart disease}) = \frac{205}{500} \times \frac{180}{500} = 0.41 \times 0.36 = 0.1476 Expected frequency = 0.1476 × 500 = 73.8 families, confirming answer B. Answer A (61.6) likely results from calculation errors or using incorrect marginal totals. Answer C (85) represents the observed frequency, not the expected frequency under independence—this is a common trap since students might confuse what the null hypothesis predicts versus what actually occurred. Answer D (42.5) suggests a fundamental misunderstanding of the independence calculation, possibly dividing instead of multiplying probabilities. Remember that in independence testing, you're comparing observed frequencies to what you'd expect if there were no association. The expected frequency formula is: (row total × column total) ÷ grand total. Always calculate marginal totals first, then apply this formula rather than using the observed joint frequency.

Question 13

A researcher conducts a controlled experiment to test whether caffeine affects reaction time. She randomly assigns 30 participants to receive either caffeine or placebo, then measures their reaction times. Based on the graph shown, which conclusion about the experimental design and results is most supported?

  1. The experiment demonstrates that caffeine significantly improves reaction time because the mean difference exceeds the error bars
  2. The large error bars indicate poor experimental design and the results should not be trusted
  3. The overlapping error bars suggest no significant difference between treatments, so caffeine has no effect on reaction time
  4. The error bars represent the range of all data points, indicating high variability in both groups
  5. The results show a meaningful difference, but the overlapping confidence intervals suggest the difference may not be statistically significant (correct answer)
Explanation: The graph shows a clear difference in means between caffeine and placebo groups, but the 95% confidence intervals overlap substantially. While overlapping confidence intervals don't definitively prove no significant difference (formal statistical testing is needed), the substantial overlap suggests the difference may not reach statistical significance. Choice A incorrectly assumes that any difference beyond error bars indicates significance. Choice B wrongly assumes large error bars indicate poor design (they may reflect natural biological variation). Choice C overstates what overlapping error bars prove. Choice D incorrectly identifies what the error bars represent (they appear to be confidence intervals, not ranges).