All questions
Question 1
A researcher compares the variability in daily step counts between office workers and manual laborers. Office workers show a mean of 6,200 steps with MAD = 1,100 steps. Manual laborers show a mean of 12,800 steps with MAD = 1,400 steps. To fairly compare variability between these groups, which approach is most appropriate?
- Compare MAD values directly since 1,400 > 1,100, manual laborers have more variable step counts
- Convert MAD to standard deviation using the relationship MAD ≈ 0.8σ for better comparison
- Calculate relative MAD (MAD/mean) since the groups have substantially different baseline activity levels (correct answer)
- Use the difference in MAD values (300 steps) to quantify how much more variable one group is than the other
Explanation: When comparing variability between groups with very different baseline values, you need to account for the scale differences to make a fair comparison. Raw measures of spread can be misleading when the underlying means differ substantially.
The correct approach is C) Calculate relative MAD (MAD/mean) because it standardizes the variability measure relative to each group's baseline activity level. For office workers: 6,2001,100=0.178 (17.8%). For manual laborers: 12,8001,400=0.109 (10.9%). This reveals that office workers actually have more variable step counts relative to their typical activity level, despite having a smaller absolute MAD.
A is wrong because directly comparing MAD values ignores the dramatic difference in baseline activity levels (6,200 vs 12,800 steps). A 1,400-step variation around 12,800 steps represents much less relative variability than 1,100 steps around 6,200 steps.
B is incorrect because converting MAD to standard deviation doesn't solve the fundamental problem of comparing groups with different scales. You'd still need to standardize the comparison somehow.
D is flawed because the 300-step difference in MAD values is meaningless without considering the context of each group's activity level. This approach still falls into the same trap as option A.
Study tip: Whenever you're comparing variability between groups with substantially different means, always consider using relative measures (like coefficient of variation or relative MAD) rather than absolute measures. This pattern appears frequently in statistics problems involving different scales or units. Question 2
A quality control manager measures the weights (in grams) of 20 manufactured bolts and calculates a mean absolute deviation (MAD) of 0.8 grams. If the target weight is 50 grams and the manager wants to reject any bolt that deviates from the target by more than 2.5 times the MAD, what percentage of bolts would typically be rejected assuming the deviations follow a roughly normal pattern?
- Approximately 5% because 2.5 MAD corresponds to about 2 standard deviations from the mean (correct answer)
- Approximately 12% because 2.5 MAD corresponds to about 1.5 standard deviations from the mean
- Approximately 20% because 2.5 MAD corresponds to about 1 standard deviation from the mean
- Approximately 32% because 2.5 MAD corresponds to about 0.67 standard deviations from the mean
Explanation: For a normal distribution, MAD ≈ 0.8σ, where σ is the standard deviation. So 2.5 × MAD ≈ 2.5 × 0.8σ = 2σ. In a normal distribution, approximately 95% of values fall within 2 standard deviations of the mean, meaning about 5% fall outside this range and would be rejected. Choice B corresponds to 1.5σ (about 13% outside), choice C to 1σ (about 32% outside), and choice D reverses the relationship incorrectly.
Question 3
A company's monthly sales data for 24 months shows a median of $45,000 and an IQR of $8,000. Due to a calculation error, each month's sales figure was recorded as exactly $2,000 higher than the actual amount. When this error is corrected by subtracting $2,000 from each value, what will be the new median and IQR respectively?
- Median: $43,000; IQR: $6,000 because both measures shift proportionally downward
- Median: $43,000; IQR: $8,000 because only the median shifts while spread measures remain constant (correct answer)
- Median: $45,000; IQR: $6,000 because the median stays fixed while variability decreases from the correction
- Median: $43,000; IQR: $10,000 because the correction increases the relative spread between quartiles
Explanation: When a constant value is added to or subtracted from every data point, measures of center (like median) shift by that same constant, but measures of spread remain unchanged. The median decreases by $2,000 to $43,000. However, the IQR stays $8,000 because the difference between Q3 and Q1 is unaffected when both quartiles shift down by the same amount. Choice A incorrectly changes the IQR. Choice C incorrectly keeps the median unchanged. Choice D incorrectly increases the IQR.
Question 4
Two manufacturing processes produce bolts with the same target diameter. Process A has a standard deviation of 0.003 inches, while Process B has a standard deviation of 0.007 inches. If both processes are adjusted so their ranges become identical, what can be concluded about their new standard deviations?
- Process A will have a larger standard deviation than Process B after the adjustment to equalize ranges
- Both processes will have identical standard deviations since equal ranges imply equal standard deviations
- Process B will still have a larger standard deviation, but the difference will be smaller than originally
- The relationship between the standard deviations cannot be determined without knowing how the ranges were equalized (correct answer)
Explanation: Range and standard deviation measure different aspects of spread. Range depends only on extreme values, while standard deviation reflects the spread of all data points around the mean. Making ranges equal could be achieved in many ways: stretching/compressing the distributions, removing outliers, adding extreme values, etc. Each method would affect the standard deviations differently. For example, if Process A's range is increased by adding outliers while keeping most data clustered, its standard deviation might increase only slightly. If Process B's range is decreased by removing outliers, its standard deviation could decrease substantially. Without knowing the specific method used to equalize ranges, we cannot determine the relationship between the new standard deviations.
Question 5
An economist studying income inequality calculated the IQR for household incomes in two cities. City A has an IQR of $23,000 and City B has an IQR of $31,000. Both cities have approximately the same median income and similar-sized populations. The economist concludes that City B has greater income inequality. What assumption is the economist making, and is it necessarily valid?
- The economist assumes larger IQR always indicates greater inequality, which is valid since IQR directly measures spread in the middle income ranges
- The economist assumes equal population sizes matter for IQR comparison, which is invalid since IQR is unaffected by sample size differences
- The economist assumes income distributions are similar in shape, which may not be valid if the cities have different income distribution patterns (correct answer)
- The economist assumes median income similarity is sufficient for comparison, which is valid since IQR measures spread around the median effectively
Explanation: When comparing measures of variability like IQR across different groups, you need to consider whether the underlying distributions are comparable. The IQR measures the spread of the middle 50% of data, but this spread can be influenced by the overall shape and pattern of the distribution.
The economist's conclusion is correct only if both cities have similarly shaped income distributions. If City B's larger IQR results from a different distribution pattern—perhaps with income clustering at certain levels or different skewness—then the IQR comparison becomes misleading. For example, City B might have many middle-class families with incomes clustered around specific salary bands, creating gaps that inflate the IQR without necessarily indicating greater inequality. Answer C correctly identifies this critical assumption.
Answer A is incorrect because while IQR does measure middle-range spread, a larger IQR doesn't automatically mean greater overall inequality if distribution shapes differ significantly. Answer B misses the point—population size isn't the key issue here since both cities have similar populations, and the real concern is about distribution shape, not sample size effects. Answer D is wrong because having similar medians doesn't guarantee that IQR comparisons are meaningful; the shapes of the distributions around those medians could be vastly different.
When comparing statistical measures across groups, always ask yourself: "What assumptions make this comparison valid?" The groups should have similar distribution characteristics for the comparison to be meaningful. Don't just focus on the numbers—consider the underlying data patterns that generate those numbers.
Question 6
A teacher recorded the number of books read by students in two different classes over a semester. Class A (15 students) had an IQR of 4 books, while Class B (25 students) had an IQR of 6 books. When the classes are combined into one group of 40 students, which statement about the combined IQR is most likely correct?
- The combined IQR will be exactly 5.0 books as the weighted average of the two class IQRs
- The combined IQR will be between 4 and 6 books but closer to 6 due to Class B's larger size
- The combined IQR could be less than 4 or greater than 6 depending on the overlap between the classes (correct answer)
- The combined IQR will be approximately 4.9 books using the formula for combining standard deviations
Explanation: IQR is not a linear measure that can be averaged or weighted by sample size. When combining data sets, the new quartiles depend on the actual distribution of all 40 values, not just the individual IQRs. If Class A read 0-8 books (IQR=4) and Class B read 20-26 books (IQR=6), the combined IQR could be much larger than 6. Conversely, if the distributions overlap significantly, the combined IQR could be smaller than either individual IQR. Choices A, B, and D all incorrectly assume IQR can be calculated from the component IQRs without knowing the actual data distributions.
Question 7
A research study tracked the daily calorie intake of 100 participants over one month. The data showed a mean of 2,200 calories with a standard deviation of 300 calories. The distribution was approximately normal.
After reviewing the data, researchers discovered that 5 participants had been recording their intake incorrectly, consistently reporting 500 calories less than their actual consumption. When this error is corrected, how will the measures of spread be affected?
- The standard deviation will increase by approximately 25 calories due to the upward shift in those values
- The IQR will remain exactly the same since the correction maintains the relative ordering of all participants (correct answer)
- Both the standard deviation and IQR will increase proportionally since 5% of the data moved significantly upward
- The standard deviation will decrease slightly while the IQR may increase depending on quartile positions
Explanation: Since the 5 participants' values are all shifted by the same amount (+500), their relative positions in the ordered data set remain unchanged. The IQR depends only on the relative positions of the quartiles, not the actual values, so it stays the same. The standard deviation also remains unchanged because it measures spread around the mean - when both individual values and the mean shift by the same amount, the deviations from the mean stay constant. Choice A is wrong because SD doesn't change. Choice C incorrectly assumes both measures increase. Choice D incorrectly predicts changes in both measures.
Question 8
A data set of 50 values has Q1 = 12, Q3 = 28, and contains no outliers when using the 1.5×IQR rule. If the minimum value is 2 and a new data point with value x is added to make it an outlier, what is the smallest possible integer value of x?
- x=52 because this exceeds the upper fence by the minimum amount for integer values
- x=53 because outliers must be strictly greater than the upper fence boundary value
- x=−12 because this exceeds the lower fence and is the most extreme possible outlier
- x=−13 because the lower fence calculation requires strict inequality for outlier classification (correct answer)
Explanation: IQR = 28 - 12 = 16. Lower fence = 12 - 1.5(16) = 12 - 24 = -12. Upper fence = 28 + 1.5(16) = 28 + 24 = 52. For a value to be an outlier, it must be either less than -12 or greater than 52. Since the question asks for the smallest possible integer value of x that makes it an outlier, we need the smallest integer less than -12, which is -13. The upper outliers (choices A and B) would be much larger than the smallest possible outlier. Choice C gives -12, but this is exactly the fence value and typically not considered an outlier (outliers are usually defined as strictly outside the fences).
Question 9
A quality inspector measures the lengths of 200 metal rods. The data shows Q1 = 14.2 cm, median = 15.8 cm, Q3 = 17.1 cm, and range = 8.4 cm. Using the 1.5×IQR rule, she identifies 8 outliers. If these outliers are removed, which measure of spread will change the most dramatically?
- The range will change most because it depends entirely on extreme values that include the outliers (correct answer)
- The IQR will change most because removing outliers shifts the positions of the quartiles significantly
- The standard deviation will change most because it is most sensitive to extreme values in the calculation
- All measures will change equally since removing 4% of the data affects each measure proportionally
Explanation: Range is calculated using only the minimum and maximum values. Since the outliers include the most extreme values in the data set, removing them will dramatically change the range. The new range will be much smaller as it will be based on the most extreme non-outlier values. IQR is relatively robust to outliers since it's based on quartiles representing the middle 50% of data. Standard deviation is affected by outliers but not as dramatically as range since it incorporates all data points. Choice B is incorrect because quartiles are resistant to outliers. Choice C is wrong because while SD is sensitive to outliers, range changes more dramatically. Choice D is incorrect because different measures have different sensitivities to outliers.
Question 10
Two data sets have the same mean and the same range. Data Set A has a standard deviation of 3.2, while Data Set B has a standard deviation of 5.1. If both sets are approximately normal, which statement about the interquartile ranges (IQR) is most accurate?
- Data Set A has a larger IQR because smaller standard deviation indicates more values near the quartiles
- Data Set B has a larger IQR because higher standard deviation indicates greater spread around the median (correct answer)
- The IQRs are approximately equal since both sets have the same range and mean values
- The relationship cannot be determined without knowing the actual quartile positions in each distribution
Explanation: For normal distributions, IQR ≈ 1.35σ, where σ is the standard deviation. Since Data Set B has a larger standard deviation (5.1 vs 3.2), it will have a larger IQR. The fact that both sets have the same range is misleading - range only considers extreme values, while IQR measures the spread of the middle 50% of the data. Choice A is incorrect because smaller standard deviation means less spread, not more values near quartiles. Choice C is wrong because equal range doesn't imply equal IQR. Choice D is incorrect because for normal distributions, the relationship between standard deviation and IQR is well-established.
Question 11
A researcher studying reaction times found that the middle 50% of participants had reaction times between 0.18 and 0.34 seconds. If the data follows a normal distribution with mean 0.26 seconds, what is the approximate standard deviation, and how does this relate to the range of the full data set?
- Standard deviation ≈ 0.12 seconds; the full range is approximately 6 times the standard deviation (correct answer)
- Standard deviation ≈ 0.08 seconds; the full range is approximately 8 times the standard deviation
- Standard deviation ≈ 0.06 seconds; the full range is approximately 10 times the standard deviation
- Standard deviation ≈ 0.04 seconds; the full range is approximately 12 times the standard deviation
Explanation: The middle 50% corresponds to the IQR. Given Q1 = 0.18 and Q3 = 0.34, IQR = 0.16 seconds. For a normal distribution, IQR ≈ 1.35σ, so σ ≈ 0.16/1.35 ≈ 0.12 seconds. The full range of a normal distribution is approximately 6 standard deviations (from about -3σ to +3σ from the mean), so range ≈ 6(0.12) = 0.72 seconds. Choice B uses the wrong conversion factor. Choices C and D underestimate the standard deviation significantly.
Question 12
Two datasets have the same mean of 50. Dataset X has a standard deviation of 8, while Dataset Y has a standard deviation of 12. A new data point of 74 is added to each dataset. How does this addition affect the relative variability of the two datasets?
- Dataset X will have greater variability than Dataset Y because 74 is further from its original distribution
- Dataset Y will still have greater variability than Dataset X, and the difference between their standard deviations will increase
- Both datasets will have identical variability since the same value was added to each dataset
- Dataset Y will still have greater variability than Dataset X, but the difference between their standard deviations will decrease (correct answer)
Explanation: When you encounter questions about adding data points to datasets with different variabilities, focus on how extreme values affect standard deviation differently depending on the existing spread.
Dataset Y starts with greater variability (standard deviation of 12 vs. 8), since both datasets have the same mean of 50. The new data point 74 is 24 units above the mean for both datasets, making it an outlier that will increase both standard deviations.
However, the key insight is that adding the same outlier to datasets with different spreads creates a "convergence effect." The dataset that was originally less variable (Dataset X) experiences a proportionally larger increase in standard deviation because the outlier represents a bigger disruption to its tighter distribution. Dataset Y, already more spread out, is less dramatically affected by the same outlier.
Think of it this way: adding one tall person to a group of people with similar heights changes the group's height variability more than adding that same person to a group already containing people of very different heights.
Choice A incorrectly suggests Dataset X will become more variable than Dataset Y - while X's variability increases more proportionally, Y still remains more variable overall. Choice B wrongly claims the difference between standard deviations will increase, when it actually decreases due to the convergence effect. Choice C makes the fundamental error of assuming identical additions create identical variability changes, ignoring how the same value affects different distributions differently.
Remember: outliers have greater relative impact on datasets with lower initial variability, causing standard deviations to converge when the same extreme value is added to different datasets.
Question 13
The waiting times (in minutes) at a doctor's office for 12 patients are: 5, 8, 12, 15, 18, 22, 25, 28, 32, 35, 38, 45. The office manager wants to report a measure of spread that best represents the typical deviation from the center for most patients, excluding extreme cases. Which measure and value should be reported?
- Report the range of 40 minutes, as it shows the full extent of waiting time variation
- Report the standard deviation of approximately 13.2 minutes, as it measures typical deviation from the mean
- Report the interquartile range of 20 minutes, as it represents the spread of the middle 50% of patients (correct answer)
- Report the mean absolute deviation of approximately 11.8 minutes, as it shows average distance from the center
Explanation: When you encounter questions about measures of spread, you need to consider both what each measure tells you and what the context requires. Here, the key phrase is "excluding extreme cases" - this immediately signals that you want a resistant measure of spread.
The interquartile range (IQR) is perfect for this situation. To find it, you first locate the quartiles. With 12 data points, Q1 is the median of the first 6 values (between 12 and 15, so Q1 = 13.5) and Q3 is the median of the last 6 values (between 32 and 35, so Q3 = 33.5). The IQR = Q3 - Q1 = 33.5 - 13.5 = 20 minutes. This represents the spread of the middle 50% of patients, automatically excluding potential outliers.
Option A's range (45 - 5 = 40 minutes) includes the most extreme values, which is exactly what the problem asks you to avoid. Option B's standard deviation, while measuring typical deviation from the mean, is heavily influenced by extreme values since it squares all deviations. Option D's mean absolute deviation is less affected by extremes than standard deviation, but it still includes all data points rather than focusing on the central portion.
The IQR in option C gives you exactly what's requested: a measure of spread for the typical middle group of patients, automatically filtering out any unusually long or short wait times.
Study tip: When you see "excluding extremes" or "typical" in spread questions, think quartiles and IQR. When you see "all data points" or "overall variation," consider range or standard deviation.
Question 14
A teacher calculates that her class's quiz scores have a mean of 82 and a range of 24 points. After discovering a grading error, she adds 3 points to every student's score. What are the new mean and range?
- The new mean is 85 and the new range is 24 points, since adding a constant affects central tendency but not spread (correct answer)
- The new mean is 85 and the new range is 27 points, since adding a constant increases both measures proportionally
- The new mean is 82 and the new range is 27 points, since the relative positions remain unchanged
- The new mean is 85 and the new range is 21 points, since adding points reduces the relative spread of scores
Explanation: Adding a constant to every data value shifts all values by that amount, changing measures of central tendency (mean increases by 3) but not measures of spread. The range remains 24 because the difference between the highest and lowest scores doesn't change. Choice B incorrectly adds 3 to the range. Choice C incorrectly keeps the mean unchanged. Choice D incorrectly decreases the range.
Question 15
A researcher collected data on the number of hours students spend on homework per week. The dataset has a mean of 12 hours and a standard deviation of 3.2 hours. If the distribution is approximately normal, what percentage of students spend between 8.8 and 15.2 hours on homework per week?
- Approximately 68% of students spend between 8.8 and 15.2 hours on homework per week (correct answer)
- Approximately 95% of students spend between 8.8 and 15.2 hours on homework per week
- Approximately 99.7% of students spend between 8.8 and 15.2 hours on homework per week
- Approximately 34% of students spend between 8.8 and 15.2 hours on homework per week
Explanation: The values 8.8 and 15.2 are exactly one standard deviation below and above the mean (12 - 3.2 = 8.8 and 12 + 3.2 = 15.2). By the empirical rule, approximately 68% of data falls within one standard deviation of the mean. Choice B represents two standard deviations (95%), choice C represents three standard deviations (99.7%), and choice D represents only one side of the distribution (34%).
Question 16
The test scores for two classes are shown with their five-number summaries. Class A: Min = 65, Q1 = 72, Median = 78, Q3 = 85, Max = 95. Class B: Min = 60, Q1 = 70, Median = 78, Q3 = 86, Max = 98. Which statement about the spread of the test scores is most accurate?
- Class A has greater variability because its interquartile range is larger than Class B's interquartile range
- Class B has greater variability because its range is larger than Class A's range, despite having a similar IQR (correct answer)
- Both classes have identical variability because they have the same median score of 78 points
- Class A has greater variability because its minimum score is higher, indicating more spread in the data
Explanation: Class A: IQR = 85 - 72 = 13, Range = 95 - 65 = 30. Class B: IQR = 86 - 70 = 16, Range = 98 - 60 = 38. Class B has both a larger IQR (16 vs 13) and larger range (38 vs 30), indicating greater variability. Choice A incorrectly states Class A has larger IQR. Choice C confuses central tendency with spread. Choice D incorrectly interprets a higher minimum as indicating more spread.