All questions
Question 1
Before conducting a chi-square test of independence on a 2×3 contingency table, a researcher calculates all expected frequencies. What is the minimum acceptable expected frequency to ensure the validity of the chi-square test?
- At least 10 in every cell, with no exceptions allowed
- At least 5 in every cell, with no exceptions allowed
- At least 5 in every cell, or at least 1 in every cell with average ≥ 5
- At least 1 in every cell, with no other restrictions needed
- At least 5 in 80% of cells, with no cell less than 1 (correct answer)
Explanation: When you encounter chi-square test assumptions, you're dealing with a fundamental requirement that ensures the test statistic follows the expected chi-square distribution. The validity of this test depends critically on having adequate expected frequencies in your contingency table.
The standard rule for chi-square tests requires that at least 80% of cells have expected frequencies of 5 or greater, AND no cell should have an expected frequency less than 1. This is a more nuanced requirement than the absolute rules suggested in most answer choices.
However, there appears to be an issue with this question as written, since the correct answer is listed as "E" but only options A through D are provided. Based on standard biostatistics principles, the correct guideline would be approximately: "At least 80% of cells should have expected frequencies ≥ 5, and no cell should have expected frequency < 1."
Looking at the given options: Choice A is too restrictive, requiring 10 in every cell when 5 is the standard threshold. Choice B incorrectly demands 5 in every single cell with no exceptions, which is overly strict. Choice C mentions an average requirement that isn't part of the standard rule - the 80% guideline is about individual cells, not averages. Choice D is too lenient, as having just 1 in every cell without the 5+ requirement would violate assumptions.
Study tip: Remember the "80% rule" for chi-square tests - this is one of the most commonly tested assumptions in biostatistics. When expected frequencies are too low, consider Fisher's exact test as an alternative.
Question 2
A researcher obtains a chi-square test statistic of 0.42 when testing independence between gender (male/female) and preference (yes/no) in a sample of 100 subjects. What does this result most likely indicate?
- Strong evidence of association between gender and preference
- The sample size is too small to detect any meaningful association
- Little evidence of association between gender and preference (correct answer)
- A calculation error occurred because chi-square cannot be less than 1
- The test assumptions have been violated due to small expected frequencies
Explanation: When interpreting chi-square test statistics, you need to understand that larger values indicate stronger evidence against independence (meaning stronger association), while smaller values suggest the observed data is close to what you'd expect if variables were truly independent.
A chi-square statistic of 0.42 is quite small, indicating that the observed frequencies in your 2×2 table are very close to what you'd expect if gender and preference were completely unrelated. This provides little evidence of any meaningful association between the variables.
Looking at the wrong answers: (A) is backwards—strong evidence of association would require a much larger chi-square value, typically several times bigger than 0.42. (B) misunderstands the result; while sample size affects statistical power, a sample of 100 is reasonably sized for a 2×2 chi-square test, and the small test statistic isn't due to insufficient sample size but rather to the lack of association in the data. (D) reflects a fundamental misconception—chi-square statistics can absolutely be less than 1 and often are when there's little to no association between variables.
The key insight is that chi-square values exist on a continuum from 0 upward, with values closer to 0 indicating independence and larger values (especially those exceeding critical values from chi-square tables) indicating dependence.
Study tip: Remember that chi-square statistics measure departure from independence—small values mean "close to independent," not "invalid test." Always compare your calculated chi-square to critical values from statistical tables to determine significance, rather than assuming any particular threshold like 1.0.
Question 3
A chi-square test of independence yields a p-value of 0.03. If the researcher had chosen α = 0.01 instead of α = 0.05, how would this change the conclusion about the null hypothesis of independence?
- The conclusion would change from rejecting H₀ to failing to reject H₀ (correct answer)
- The conclusion would change from failing to reject H₀ to rejecting H₀
- The conclusion would remain the same because p-values don't depend on α
- The test would become invalid because α cannot be changed after data collection
- The conclusion would remain the same in both cases: reject H₀
Explanation: When you encounter chi-square test questions involving different significance levels, focus on the relationship between p-values and alpha thresholds for decision-making.
The key principle is simple: you reject the null hypothesis when p < α, and fail to reject when p ≥ α. With a p-value of 0.03, let's see what happens under each significance level.
At α = 0.05: Since 0.03 < 0.05, you would reject H₀ (independence) and conclude the variables are associated. At α = 0.01: Since 0.03 > 0.01, you would fail to reject H₀ and cannot conclude the variables are associated. This represents a change from rejecting to failing to reject H₀, making option A correct.
Option B reverses the direction of change, incorrectly suggesting the conclusion moves from failing to reject to rejecting. Option C contains a true statement that p-values don't depend on α, but misses the point entirely—while the p-value stays 0.03, your interpretation of what it means changes based on your chosen threshold. Option D reflects a misconception about hypothesis testing; researchers can absolutely consider different significance levels during analysis, and doing so doesn't invalidate the test.
The p-value represents the strength of evidence against H₀, but α represents your personal threshold for "convincing enough." A more stringent α (like 0.01) requires stronger evidence to reject H₀.
Study tip: Remember that lowering α makes it harder to reject the null hypothesis. You need stronger evidence (smaller p-values) to reach statistical significance with more stringent alpha levels.
Question 4
A researcher tests independence between education level (high school, college, graduate) and income bracket (low, medium, high) using a chi-square test. The analysis yields χ² = 11.8 with p = 0.019. What is the most appropriate interpretation?
- Education level causes differences in income bracket
- There is a 1.9% probability that education and income are independent
- There is evidence of association between education level and income bracket (correct answer)
- Education level and income bracket are proven to be related
- The sample provides insufficient evidence to conclude association exists
Explanation: When you encounter chi-square test results, you're dealing with a test of independence that examines whether two categorical variables are associated. The key is understanding what the p-value tells you about the relationship between variables.
In this study, the chi-square statistic of 11.8 with p = 0.019 means there's only a 1.9% chance of observing this strong an association (or stronger) if education and income were truly independent. Since p < 0.05, you reject the null hypothesis of independence and conclude there's evidence of association between education level and income bracket. This makes C correct.
Let's examine why the other options are wrong. Option A claims causation, but chi-square tests only detect association—they cannot establish that education causes income differences. Correlation doesn't equal causation. Option B misinterprets the p-value. The p = 0.019 doesn't mean there's a 1.9% probability that the variables are independent; rather, it's the probability of seeing these results assuming they are independent. Option D uses absolute language ("proven to be related") which is too strong. Statistical tests provide evidence for associations, but they don't "prove" relationships definitively.
Remember this pattern: chi-square tests detect associations, never causation. The p-value represents the probability of your observed results under the null hypothesis, not the probability that the null hypothesis is true. Always interpret results as "evidence of association" rather than "proof of relationship" to avoid overstating your statistical conclusions.
Question 5
A 4×2 contingency table has a total sample size of n = 80. During analysis, one cell is found to have an expected frequency of 3.2. What is the most appropriate course of action?
- Proceed with the chi-square test since n > 50
- Combine adjacent categories to increase the expected frequency in that cell (correct answer)
- Increase the sample size to at least 100 before conducting the test
- Use Fisher's exact test instead of the chi-square test
- Apply a continuity correction to adjust for the small expected frequency
Explanation: When you encounter contingency table analysis, the key consideration is whether the assumptions for the chi-square test are met. The chi-square test requires that expected frequencies in each cell be sufficiently large—typically at least 5, though some sources accept as low as 1 with most cells having expected frequencies ≥5.
With an expected frequency of 3.2 in one cell, you're below the recommended threshold. The most appropriate solution is to combine adjacent categories to increase expected frequencies. This involves collapsing rows or columns that are conceptually similar, which increases the cell counts while maintaining the validity of your analysis. For example, if you're analyzing age groups, you might combine "18-25" and "26-35" into "18-35."
Option A is incorrect because while a large total sample size (n=80) is helpful, it doesn't guarantee adequate expected frequencies in individual cells. The chi-square assumption focuses on cell-level expected frequencies, not total sample size.
Option C misses the point—simply increasing sample size might help, but it's not the most efficient or practical solution when you can address the problem by restructuring your existing data.
Option D suggests Fisher's exact test, which is typically reserved for 2×2 tables and becomes computationally intensive for larger tables. A 4×2 table would make Fisher's exact test unnecessarily complex when combining categories is a simpler solution.
Study tip: Always check expected frequencies before running a chi-square test. When any expected frequency falls below 5, your first instinct should be to combine logically related categories rather than abandoning the chi-square test entirely.
Question 6
A researcher calculates χ² = 7.25 for a test of independence in a 2×3 contingency table. Using α = 0.05, what conclusion should be reached?
- Reject H₀ because 7.25 > 5.99 (correct answer)
- Fail to reject H₀ because 7.25 < 9.49
- Reject H₀ because 7.25 > 3.84
- The test is inconclusive because the critical value cannot be determined
- Fail to reject H₀ because 7.25 < 12.59
Explanation: When you encounter a chi-square test of independence, you need to determine the correct degrees of freedom to find the appropriate critical value. For a contingency table, degrees of freedom equals (rows - 1) × (columns - 1). In this 2×3 table, that's (2-1) × (3-1) = 1 × 2 = 2 degrees of freedom.
With df = 2 and α = 0.05, the critical value from the chi-square distribution table is 5.99. Since your calculated χ² = 7.25 exceeds this critical value (7.25 > 5.99), you reject the null hypothesis and conclude that the variables are not independent.
Answer A is correct because it uses the proper critical value of 5.99 for 2 degrees of freedom and correctly applies the decision rule. Answer B incorrectly uses 9.49, which would be the critical value for 4 degrees of freedom—this suggests miscalculating df as (2×3) = 6, then subtracting 2, which is wrong. Answer C uses 3.84, the critical value for 1 degree of freedom, indicating the error of calculating df as either (2-1) or (3-1) instead of their product. Answer D is simply incorrect since critical values for chi-square tests are always determinable from the degrees of freedom and significance level.
Study tip: Always double-check your degrees of freedom calculation in contingency tables—it's (rows-1) × (columns-1), not rows × columns. This is one of the most common errors students make on chi-square tests of independence.
Question 7
A chi-square test of independence produces a standardized residual of -2.3 for one cell. What does this indicate about that particular cell?
- The observed frequency is 2.3 times larger than expected
- The observed frequency is 2.3 units smaller than expected
- The observed frequency is significantly smaller than expected (correct answer)
- The cell contributes 2.3 to the overall chi-square statistic
- The cell has a 2.3% probability of occurring by chance
Explanation: When you encounter standardized residuals in chi-square tests, you're looking at how much each cell deviates from independence, measured in standard deviation units. A standardized residual tells you whether the observed frequency in a cell is significantly different from what you'd expect if the variables were independent.
The standardized residual formula is: standard errorobserved−expected
A value of -2.3 means the observed frequency is 2.3 standard deviations below the expected frequency. Since standardized residuals follow approximately a standard normal distribution, values beyond ±2 are typically considered statistically significant. The negative sign indicates the observed count is smaller than expected, and the magnitude (2.3) indicates this difference is statistically significant.
Option A is incorrect because the negative sign means observed is smaller, not larger, than expected. The value also doesn't represent a multiplier. Option B misinterprets the standardized residual as raw units rather than standard deviation units—it's not simply 2.3 fewer observations. Option D confuses standardized residuals with raw residuals; standardized residuals don't directly contribute their value to the chi-square statistic.
The correct answer is C: the observed frequency is significantly smaller than expected, as indicated by the negative sign and the magnitude exceeding the ±2 threshold for significance.
Remember: standardized residuals help you identify which specific cells are driving significant chi-square results. Values beyond ±2 flag cells with meaningful deviations from independence. Question 8
A researcher wants to test if voting preference (Candidate A, B, C) is independent of age group (18-30, 31-50, 51+). After collecting data from 300 voters, which assumption must be verified before conducting the chi-square test?
- Each voter must be randomly selected from the population
- The sample size must be at least 30 per cell
- At least 80% of expected frequencies must be 5 or greater (correct answer)
- The variables must be normally distributed
- The relationship between variables must be linear
Explanation: When you encounter questions about chi-square tests of independence, you're dealing with categorical data analysis that requires specific assumptions to be met for valid results. The chi-square test examines whether two categorical variables are independent by comparing observed frequencies to expected frequencies.
The correct answer is C because the chi-square test requires that at least 80% of expected frequencies be 5 or greater (with no expected frequency below 1). This assumption ensures the chi-square distribution provides a good approximation for the test statistic. When expected frequencies are too small, the test becomes unreliable and may give misleading p-values.
Here's why the other options are incorrect: Option A describes good study design practice but isn't a statistical assumption you verify after data collection—it's a design consideration. Option B confuses the chi-square test with other statistical procedures; there's no requirement for 30 observations per cell. The actual requirement relates to expected frequencies, not sample sizes. Option D applies to tests like t-tests or ANOVA that assume normal distributions, but chi-square tests work with categorical data where normality isn't relevant—you're counting frequencies, not measuring continuous variables.
To check this assumption, you'd calculate expected frequencies using the formula: Expected = (row total × column total) ÷ grand total. With your 3×3 table (3 candidates × 3 age groups), you'd have 9 expected frequencies to examine.
Study tip: For chi-square questions, always think "expected frequencies" when you see assumption-checking. This is the most commonly tested assumption and distinguishes chi-square from other statistical tests.
Question 9
A medical researcher tests whether treatment outcome (cured, improved, unchanged, worse) is independent of patient age group (young, old) in a clinical trial with 200 participants. The chi-square test yields p = 0.08. How should this result be interpreted in the context of medical decision-making?
- Age group definitely does not affect treatment outcome
- There is an 8% chance that age affects treatment outcome
- The evidence for association between age and outcome is not statistically significant at α = 0.05 (correct answer)
- The sample size is too small to detect clinically meaningful associations
- Age group and treatment outcome are moderately correlated
Explanation: When you encounter chi-square test results in medical research, you're dealing with hypothesis testing for independence between categorical variables. Here, the researcher is testing whether treatment outcome and age group are independent (null hypothesis) versus associated (alternative hypothesis).
The p-value of 0.08 tells you the probability of observing data this extreme or more extreme if the null hypothesis (independence) were true. Since p = 0.08 > 0.05 (the conventional significance level), you fail to reject the null hypothesis. This means the evidence for an association between age and treatment outcome is not statistically significant at α = 0.05, making answer C correct.
Answer A is wrong because failing to reject the null hypothesis doesn't prove it's true—it simply means insufficient evidence exists to conclude there's an association. Answer B misinterprets the p-value: it's not the probability that age affects outcome, but rather the probability of seeing these results assuming no association exists. Answer D makes an unfounded assumption about sample size—200 participants may be adequate depending on effect sizes and study design, and the p-value doesn't directly indicate sample size issues.
Remember that p-values measure evidence strength against the null hypothesis, not the probability that your research hypothesis is true. When p > α, you conclude "no statistically significant association" rather than "no association exists." This distinction is crucial in medical research where you're making evidence-based recommendations about treatment protocols.
Question 10
In a chi-square test of independence between drug type (A, B, C) and side effect occurrence (present, absent), the degrees of freedom are calculated as 2. If the same data were analyzed using only drugs A and B, what would be the degrees of freedom?
- 1 (correct answer)
- 2
- 0
- 3
- Cannot be determined without knowing the sample sizes
Explanation: When you encounter chi-square test questions, focus on the contingency table structure to determine degrees of freedom. The formula is always (r−1)×(c−1), where r is the number of rows and c is the number of columns.
In the original scenario, you have a 3×2 contingency table: three drug types (A, B, C) versus two outcomes (present, absent). This gives you (3−1)×(2−1)=2×1=2 degrees of freedom, confirming the given information.
When you reduce the analysis to only drugs A and B, you're creating a smaller 2×2 contingency table. The side effect outcomes remain the same (present, absent), but now you only have two drug types instead of three. Using the same formula: (2−1)×(2−1)=1×1=1 degree of freedom.
Therefore, A) 1 is correct.
B) 2 would be wrong because this assumes you still have three drug categories, but you've eliminated drug C from the analysis. C) 0 is impossible unless you had only one category in both dimensions, which would make the test meaningless. D) 3 would require either four drug types or three outcome categories, neither of which applies here.
Remember this pattern: degrees of freedom in chi-square tests depend entirely on the dimensions of your contingency table. When you reduce categories in either dimension, you proportionally reduce the degrees of freedom. Always count your actual categories in the current analysis, not what you might have had in a previous or larger study. Question 11
In a chi-square test of independence, the expected frequency for cell (i,j) in a contingency table is calculated as 12.5, but only 8 observations fall in this cell. What is the contribution of this cell to the overall chi-square test statistic?
- 1.62 (correct answer)
- 0.36
- 4.50
- 2.25
- 0.64
Explanation: When you encounter chi-square test calculations, you're working with the formula that measures how much observed data deviates from what we'd expect if variables were independent. The contribution of each cell to the overall test statistic follows the formula: E(O−E)2, where O is the observed frequency and E is the expected frequency.
In this problem, you have an expected frequency (E) of 12.5 and an observed frequency (O) of 8. Plugging into the formula: 12.5(8−12.5)2=12.5(−4.5)2=12.520.25=1.62
This confirms that A) 1.62 is correct.
Let's examine why the other options are wrong. B) 0.36 represents a common error where students might incorrectly calculate (12.5)2(8−12.5)2 - they're squaring the denominator when they shouldn't. C) 4.50 is simply the absolute difference ∣8−12.5∣ without any squaring or division, showing a fundamental misunderstanding of the formula. D) 2.25 appears to be 9(8−12.5)2, suggesting confusion about which value serves as the denominator - perhaps mixing up observed and expected frequencies.
Remember that chi-square contributions are always non-negative since we're squaring the numerator, and larger deviations from expected values create larger contributions to the test statistic. Always double-check that you're using the expected frequency as the denominator, not the observed frequency. Question 12
Two researchers independently conduct chi-square tests of independence on the same dataset. Researcher A uses the raw counts, while Researcher B multiplies all frequencies by 2 before analysis. How will their test statistics compare?
- Researcher B's chi-square statistic will be twice as large as Researcher A's (correct answer)
- Researcher B's chi-square statistic will be four times as large as Researcher A's
- Both researchers will obtain identical chi-square statistics
- Researcher B's chi-square statistic will be half as large as Researcher A's
- The relationship depends on the specific cell frequencies and cannot be determined
Explanation: When you encounter questions about how data transformations affect chi-square statistics, focus on understanding the mathematical relationship in the test formula: χ2=∑E(O−E)2, where O represents observed frequencies and E represents expected frequencies.
Let's trace what happens when all frequencies are multiplied by 2. If the original observed frequencies are doubled, the expected frequencies (calculated from marginal totals) are also doubled. In the chi-square formula, both the numerator and denominator are affected: the numerator (O−E)2 becomes (2O−2E)2=4(O−E)2, while the denominator E becomes 2E. Therefore, each term becomes 2E4(O−E)2=2⋅E(O−E)2, making the overall chi-square statistic exactly twice as large.
This confirms that A is correct - Researcher B's statistic will be twice Researcher A's.
B incorrectly assumes only the numerator is affected by the doubling, ignoring that expected frequencies also double. C reflects a common misconception that proportional scaling doesn't affect the test statistic - while the relationship between variables remains the same conceptually, the mathematical calculation does change. D gets the direction wrong, perhaps confusing this with situations where larger sample sizes might reduce certain statistics.
Study tip: Remember that chi-square statistics are not scale-invariant. When all frequencies are multiplied by a constant k, the chi-square statistic is multiplied by that same constant k. This is why using raw counts versus percentages matters in chi-square analysis. Question 13
Two contingency tables are constructed from the same dataset: Table 1 uses the original categories, while Table 2 combines two adjacent categories. If Table 1 yields χ² = 12.4 with df = 6, what can be said about the chi-square statistic for Table 2?
- It will definitely be larger than 12.4
- It will definitely be smaller than 12.4
- It will be exactly equal to 12.4
- It cannot be determined without additional information (correct answer)
- It will have the same p-value as Table 1
Explanation: When you encounter questions about modifying contingency tables, you need to understand that combining categories fundamentally changes the data structure and can affect the chi-square statistic in unpredictable ways.
The relationship between the original and modified chi-square statistics depends on several factors you cannot determine from the given information. When you combine two adjacent categories, you reduce the degrees of freedom (Table 2 will have df = 4 instead of 6), but the actual chi-square value depends on how the observed and expected frequencies redistribute across the new table structure.
The chi-square statistic could increase, decrease, or theoretically remain similar depending on the specific data patterns. If the combined categories had similar deviation patterns from expected values, the statistic might decrease. Conversely, if combining categories creates larger discrepancies between observed and expected frequencies in the remaining cells, the statistic could increase.
Answer A is wrong because there's no mathematical principle guaranteeing the statistic will be larger when categories are combined. Answer B is incorrect because the statistic doesn't necessarily decrease - the redistribution of frequencies could amplify differences. Answer C is wrong because exact equality would be an extraordinary coincidence with no theoretical basis.
Answer D is correct because without knowing the actual frequency distributions in both tables, you cannot predict the direction or magnitude of change in the chi-square statistic.
Study tip: Remember that combining categories in contingency tables affects both the degrees of freedom AND the chi-square statistic value, but the direction of change in the statistic cannot be predicted without examining the actual data structure.
Question 14
A study examines the relationship between exercise frequency (daily, weekly, monthly, never) and health status (excellent, good, fair, poor). The resulting 4×4 contingency table shows χ² = 18.7. Using α = 0.01, what is the appropriate conclusion?
- Reject H₀; exercise frequency and health status are associated
- Fail to reject H₀; there is insufficient evidence of association (correct answer)
- Reject H₀; exercise frequency causes changes in health status
- Fail to reject H₀; exercise frequency and health status are independent
- The test is invalid due to too many categories
Explanation: When you encounter a chi-square test of independence with a contingency table, you're testing whether two categorical variables are associated. The key is comparing your calculated test statistic to the critical value at your chosen significance level.
For a 4×4 contingency table, the degrees of freedom equal (rows - 1) × (columns - 1) = (4-1) × (4-1) = 9. With α = 0.01 and df = 9, the critical value is χ0.01,92=21.67. Since your calculated χ2=18.7 is less than 21.67, you fail to reject the null hypothesis.
Choice A is incorrect because 18.7 < 21.67, so we don't have sufficient evidence to reject H₀ at the 0.01 level. Choice C makes a critical error by confusing association with causation – chi-square tests can only detect association, never establish causality. Even if we had rejected H₀, we couldn't conclude that exercise causes health changes. Choice D uses incorrect terminology by stating the variables "are independent" rather than saying we lack evidence against independence.
Choice B correctly states we fail to reject H₀ and acknowledges insufficient evidence of association, which is the proper interpretation when our test statistic doesn't exceed the critical value.
Remember: Chi-square tests require you to compare your calculated statistic against the critical value for your chosen α level. Always check your degrees of freedom calculation, and never interpret a significant result as proof of causation – association and causation are fundamentally different concepts in statistics. Question 15
A researcher performs a chi-square test on a 3×2 contingency table and obtains χ² = 5.8. A colleague suggests that the same data could be analyzed using a different statistical test. Which alternative test would be most appropriate for the same research question?
- Two-sample t-test
- Fisher's exact test (correct answer)
- Mann-Whitney U test
- Analysis of variance (ANOVA)
- Pearson correlation
Explanation: When you encounter a chi-square test on a contingency table, you're dealing with categorical data analysis to test for independence between two variables. The key insight here is recognizing that multiple statistical tests can sometimes address the same research question, each with different assumptions and conditions.
Fisher's exact test (B) is the most appropriate alternative because it tests the exact same null hypothesis as the chi-square test: whether two categorical variables are independent. Both tests work on contingency tables and answer identical research questions. The main difference is that Fisher's exact test calculates exact p-values rather than using the chi-square approximation, making it especially valuable when sample sizes are small or expected cell counts are low (typically when any expected count is less than 5).
Option A (two-sample t-test) is incorrect because it compares means of continuous variables between two groups, not categorical associations. Option C (Mann-Whitney U test) is wrong because it compares distributions of ordinal or continuous data between two groups, not categorical independence. Option D (ANOVA) is incorrect because it compares means across multiple groups for continuous outcomes, not categorical relationships.
The chi-square value of 5.8 with 2 degrees of freedom (calculated as (3-1)×(2-1) = 2) suggests the researcher might want to verify results with Fisher's exact test, especially if any expected cell counts are small.
Study tip: Remember that chi-square tests and Fisher's exact tests are interchangeable for testing independence in contingency tables—Fisher's exact is your go-to when sample sizes are small or assumptions are violated.
Question 16
A medical researcher conducts a chi-square test of independence to examine the relationship between age group (Under 30, 30-50, Over 50) and vaccine response (Strong, Moderate, Weak). The analysis reveals χ2=13.28 with df = 4. However, the researcher realizes that the 'Weak' response category has very low frequencies across all age groups. What is the most appropriate way to address this issue while maintaining the integrity of the analysis?
- Transform the data using log-linear modeling to handle the sparse contingency table structure
- Remove participants with 'Weak' responses from the analysis and proceed with a 3×2 table
- Apply a continuity correction to the existing chi-square statistic to account for low expected frequencies
- Combine 'Moderate' and 'Weak' response categories to create a 3×2 table and recalculate the test statistic (correct answer)
Explanation: When dealing with chi-square tests of independence, you need adequate expected frequencies in each cell (typically ≥5) for the test to be valid. Low frequencies in contingency table cells violate this assumption and can lead to unreliable results.
The most appropriate solution when facing sparse cells is to combine logically related categories to increase cell frequencies while preserving the meaningful structure of your data. Option D does exactly this by merging 'Moderate' and 'Weak' responses into a single category, creating a 3×2 table that likely meets the expected frequency requirements. This combination makes conceptual sense since both represent non-strong responses, and you'd recalculate the chi-square statistic with the proper degrees of freedom (df = 2).
Option A is unnecessarily complex for this situation. Log-linear modeling is used for more sophisticated analyses of multi-way tables, not for fixing basic frequency problems in a simple independence test. Option B removes valuable data from your analysis, reducing power and potentially introducing bias by excluding an entire response category. This wastes information and changes the research question. Option C misunderstands the problem entirely—continuity corrections are applied to 2×2 tables to account for the discrete nature of the chi-square distribution, not to fix low expected frequencies.
Remember this principle: when chi-square assumptions are violated due to sparse cells, your first strategy should be to combine categories that make logical sense rather than discarding data or switching to more complex methods. Always ensure your combined categories are conceptually meaningful to maintain interpretability.
Question 17
A chi-square test of independence is conducted on a 3×4 contingency table with n = 240 participants. The calculated test statistic is χ2=18.52. Using α=0.01, what is the appropriate critical value, and what conclusion should be drawn about the null hypothesis?
- Critical value = 21.67; reject H₀ and conclude that the variables are significantly associated
- Critical value = 21.67; fail to reject H₀ and conclude that there is insufficient evidence of association
- Critical value = 16.92; fail to reject H₀ and conclude that there is insufficient evidence of association
- Critical value = 16.92; reject H₀ and conclude that the variables are significantly associated (correct answer)
Explanation: When you encounter a chi-square test of independence, you need to determine the degrees of freedom, find the critical value, and compare it to your test statistic to make a decision about the null hypothesis.
For a contingency table, degrees of freedom equals (r−1)×(c−1) where r is the number of rows and c is the number of columns. With a 3×4 table, you get (3−1)×(4−1)=2×3=6 degrees of freedom.
Using α=0.01 and 6 degrees of freedom, the critical value from the chi-square distribution table is 16.92. Since your calculated test statistic (18.52) exceeds this critical value, you reject the null hypothesis and conclude the variables are significantly associated.
Option A uses the wrong critical value (21.67), which corresponds to a different degrees of freedom—likely someone miscalculated as (3×4)−3=9 degrees of freedom. Option B makes the same critical value error but also reaches the wrong conclusion. Option C uses the correct critical value (16.92) but incorrectly fails to reject the null hypothesis, missing that 18.52 > 16.92.
The key strategy here is remembering the correct degrees of freedom formula for contingency tables: (rows−1)×(columns−1), not the total number of cells minus something else. Always double-check your degrees of freedom calculation first, as this drives everything else in the analysis. Many students trip up on this fundamental step, leading to wrong critical values and incorrect conclusions. Question 18
In a study of 240 patients, researchers want to test if treatment response (improved/no change/worse) is independent of age group (young/middle/old). The observed frequencies are shown in the table. What is the expected frequency for middle-aged patients who improved?
- 32
- 28.5 (correct answer)
- 36
- 30.4
- 25.6
Explanation: Expected frequency = (row total × column total) / grand total. From the table: Middle age total = 90, Improved total = 76, Grand total = 240. Expected = (90 × 76) / 240 = 6840 / 240 = 28.5. This matches choice B.
Question 19
A researcher conducts a chi-square test of independence to examine the relationship between smoking status (smoker, non-smoker) and lung cancer diagnosis (yes, no) in a sample of 500 adults. The calculated test statistic is χ2=8.64. However, upon reviewing the data, the researcher discovers that 15% of the cells in the contingency table have expected frequencies less than 5. What is the most appropriate next step?
- Proceed with the chi-square test since the test statistic exceeds the critical value at α=0.01
- Apply Yates' continuity correction to adjust the test statistic before interpreting results
- Use Fisher's exact test instead of the chi-square test to ensure valid inference (correct answer)
- Increase the sample size proportionally until all expected frequencies exceed 5 participants
Explanation: When more than 20% of cells (or any cell in a 2×2 table) have expected frequencies less than 5, the chi-square test assumptions are violated. Since this is a 2×2 table (smoking vs. lung cancer), having 25% of cells (1 out of 4) with expected frequencies <5 violates the assumption. Fisher's exact test is the appropriate alternative as it doesn't rely on large-sample approximations. Option A ignores the assumption violation. Option B (Yates' correction) is for different purposes and doesn't address the expected frequency issue. Option D is impractical and doesn't solve the fundamental distributional problem.
Question 20
In a study examining the association between treatment type (Drug A, Drug B, Placebo) and treatment outcome (Improved, No Change, Worsened), a chi-square test of independence yields χ2=12.84 with p<0.05. The researcher concludes there is a significant association and wants to determine which specific treatment-outcome combinations contribute most to this association. Which approach would provide the most appropriate post-hoc analysis?
- Calculate standardized residuals for each cell and identify those with absolute values greater than 2.0 (correct answer)
- Perform pairwise chi-square tests between all possible treatment pairs using Bonferroni correction
- Conduct separate one-way ANOVA tests for each outcome category across treatment groups
- Apply Tukey's HSD test to compare mean outcome scores between treatment groups
Explanation: Standardized residuals (also called adjusted residuals) indicate which cells contribute most to the overall chi-square statistic. Values with absolute magnitude >2.0 suggest significant departure from independence in that cell. This is the standard post-hoc approach for significant chi-square tests. Option B is incorrect because pairwise chi-square tests would collapse the 3×3 table inappropriately and lose information. Option C is wrong because ANOVA requires continuous dependent variables, not categorical outcomes. Option D is incorrect for the same reason - Tukey's HSD is for continuous variables and requires ANOVA assumptions.