Biostatistics Quiz: Distribution Shape And Outliers
20 questions · exam conditions
0:00
Distribution Shape And OutliersQuestion 1 of 20

A researcher creates a stem-and-leaf plot for blood pressure readings and notices that most values cluster between 120-140 mmHg, but there are isolated readings at 98, 102, and 180, 185 mmHg. The median is 132 mmHg and the mean is 136 mmHg. What is the most accurate assessment of this distribution?

Symmetric distribution with outliers balanced on both extremes, explaining the mean-median similarity
Left-skewed distribution with low-end outliers having greater influence than high-end outliers
Right-skewed distribution where high-end outliers outweigh the influence of low-end outliers
Normal distribution with random measurement errors creating apparent but non-significant outliers
Bimodal distribution with outliers representing a separate population of hypertensive patients
← Back to quizzes

Biostatistics Quiz

Biostatistics Quiz: Distribution Shape And Outliers

Practice Distribution Shape And Outliers in Biostatistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Distribution Shape And Outliers, giving you a quick way to practice the rules, question types, and explanations that matter most for Biostatistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A researcher creates a stem-and-leaf plot for blood pressure readings and notices that most values cluster between 120-140 mmHg, but there are isolated readings at 98, 102, and 180, 185 mmHg. The median is 132 mmHg and the mean is 136 mmHg. What is the most accurate assessment of this distribution?

  1. Symmetric distribution with outliers balanced on both extremes, explaining the mean-median similarity
  2. Left-skewed distribution with low-end outliers having greater influence than high-end outliers
  3. Right-skewed distribution where high-end outliers outweigh the influence of low-end outliers (correct answer)
  4. Normal distribution with random measurement errors creating apparent but non-significant outliers
  5. Bimodal distribution with outliers representing a separate population of hypertensive patients
Explanation: When analyzing distributions with outliers, focus on the relationship between mean and median, along with the direction and magnitude of extreme values. The mean is sensitive to outliers while the median remains resistant to them. Here, the mean (136 mmHg) is greater than the median (132 mmHg), which immediately suggests right skewness. When extreme values pull the distribution's tail toward higher numbers, they drag the mean in that direction while leaving the median relatively unaffected. The outliers at 180 and 185 mmHg are much farther from the central cluster (120-140 mmHg) than the low outliers at 98 and 102 mmHg. These high-end outliers have a disproportionate effect on the mean because they're more extreme relative to the main distribution. Option A is incorrect because the mean and median aren't similar enough to indicate symmetry—a 4 mmHg difference with outliers present suggests asymmetry. Option B misidentifies the skew direction; left-skewed distributions have mean less than median, not greater. Option D incorrectly assumes normality when the mean-median relationship and outlier pattern clearly indicate skewness, not random measurement error around a normal distribution. The key insight is that right skewness occurs when the tail extends toward higher values, pulling the mean upward past the median. The high-end outliers (180, 185) are proportionally more extreme than the low-end ones (98, 102), creating this rightward pull. Study tip: Remember "mean chases the tail"—in right-skewed distributions, extreme high values pull the mean above the median, while in left-skewed distributions, extreme low values pull the mean below the median.

Question 2

A researcher examining hospital length of stay data finds that 80% of patients stay 1-5 days, 15% stay 6-10 days, and 5% stay 15-45 days. The mean stay is 6.2 days while the median is 3.0 days. What does this pattern indicate about outliers and their effect on distribution measures?

  1. The distribution is symmetric with outliers having minimal impact on central tendency measures
  2. Right-skewed distribution where outliers substantially inflate the mean compared to the median (correct answer)
  3. Left-skewed distribution with outliers creating an artificially low median relative to the mean
  4. Bimodal distribution where outliers represent a distinct population requiring separate analysis
  5. Normal distribution with the mean-median difference explained by measurement error rather than outliers
Explanation: When analyzing distribution shapes and the impact of outliers, focus on the relationship between mean and median. These measures respond differently to extreme values, giving you clues about the data's distribution pattern. Here, you have a clear right-skewed distribution. The data shows 80% of patients with short stays (1-5 days), but 5% have very long stays (15-45 days). These lengthy stays are outliers that pull the mean (6.2 days) substantially higher than the median (3.0 days). In right-skewed distributions, the mean is always greater than the median because outliers on the right side inflate the average while having no effect on the middle value. Option A is incorrect because the distribution is clearly not symmetric - a symmetric distribution would have roughly equal means and medians. Option C describes left-skewedness, which would show the opposite pattern (mean less than median) with outliers on the left side pulling the mean down. Option D suggests bimodality, but the data shows a single peak with a long right tail, not two distinct peaks. The key indicator here is mean > median, which signals right skewness. The extreme difference (6.2 vs 3.0) confirms that outliers are having a substantial impact. Study tip: Remember the skewness rule: when mean > median, think right-skewed with high outliers inflating the mean. When median > mean, think left-skewed with low outliers deflating the mean. This pattern appears frequently on biostatistics exams, especially with healthcare data like length of stay, costs, or lab values.

Question 3

A medical researcher analyzes patient recovery times and finds that the 5-number summary is: Min=2 days, Q1=8 days, Median=12 days, Q3=18 days, Max=65 days. Three patients had recovery times of 45, 52, and 65 days. What is the most accurate interpretation of this distribution?

  1. Symmetric distribution with the three longest recovery times representing normal variation
  2. Right-skewed distribution with all three extended recovery times qualifying as outliers
  3. Left-skewed distribution where the maximum values indicate measurement or recording errors
  4. Right-skewed distribution with confirmed outliers, but the maximum value may not be an outlier (correct answer)
  5. Bimodal distribution where the extended times represent a secondary population of complicated cases
Explanation: When analyzing distributions with outliers, you need to examine both the shape characteristics and systematically identify outliers using the 1.5×IQR rule. First, determine the distribution shape. With Min=2, Q1=8, Median=12, Q3=18, Max=65, notice that the distance from median to Q3 (6 days) is much smaller than from Q3 to maximum (47 days). This indicates right skewness, where the tail extends toward higher values. Next, identify outliers using the standard criterion. Calculate IQR = Q3 - Q1 = 18 - 8 = 10 days. The upper fence is Q3 + 1.5×IQR = 18 + 15 = 33 days. Any value above 33 days qualifies as an outlier. Among the three patients with times of 45, 52, and 65 days, both 45 and 52 exceed the 33-day threshold, confirming they're outliers. However, 65 days, while extreme, might represent natural variation in a right-skewed medical distribution rather than a true outlier requiring investigation. Answer A is wrong because the distribution is clearly right-skewed, not symmetric. Answer B incorrectly assumes all three values are outliers when systematic calculation shows otherwise. Answer C misidentifies the skew direction—left skewness would show a long tail toward lower values, opposite of what we see here. Study tip: Always calculate the actual outlier boundaries using 1.5×IQR rather than making visual judgments. In medical data, extreme values often represent real biological variation, so distinguish between statistical outliers and clinically meaningful extremes.

Question 4

A pharmaceutical company tests drug absorption rates and obtains the following percentile information: 10th percentile = 2.1 hours, 25th percentile = 2.8 hours, 50th percentile = 3.5 hours, 75th percentile = 4.6 hours, 90th percentile = 6.2 hours. Individual observations at 8.5, 9.1, and 12.3 hours were also recorded. What distribution characteristics are most apparent?

  1. Symmetric distribution with the extreme values representing acceptable pharmacological variation
  2. Right-skewed distribution with multiple confirmed outliers in the delayed absorption range (correct answer)
  3. Left-skewed distribution where rapid absorption creates the primary outlier concern
  4. Normal distribution with the extreme values likely due to patient compliance issues
  5. Bimodal distribution indicating two distinct absorption pathways requiring separate analysis
Explanation: When analyzing drug absorption data, you need to examine the distribution shape by looking at percentile spacing and identifying potential outliers. This involves understanding how data spreads around the median and recognizing when extreme values fall outside expected ranges. The percentile data reveals key distribution characteristics. Notice how the distances between percentiles increase as you move upward: from the 25th to 50th percentile is 0.7 hours (2.8 to 3.5), while from the 50th to 75th percentile is 1.1 hours (3.5 to 4.6), and from the 75th to 90th percentile jumps to 1.6 hours (4.6 to 6.2). This increasing spread indicates a right-skewed distribution with a longer tail extending toward higher values. The individual observations at 8.5, 9.1, and 12.3 hours are substantially beyond the 90th percentile of 6.2 hours, confirming these as outliers in the delayed absorption range. These extreme values further support the right-skewed pattern. Answer A is incorrect because the unequal percentile spacing and extreme outliers contradict symmetry. Answer C is wrong since left skew would show the longest tail toward lower values, not higher ones. Answer D fails because normal distributions have symmetric, equal percentile spacing, which this data clearly lacks. Study tip: When evaluating distribution shape, always compare percentile intervals. Equal spacing suggests normality, while progressively increasing intervals toward one direction indicate skewness in that direction. Values far beyond the 90th percentile are strong outlier candidates.

Question 5

A researcher studying exam scores finds that removing the 5 lowest scores (ranging from 15-25 points) changes the mean from 78.2 to 79.8 points, while removing the 5 highest scores (ranging from 95-98 points) changes the mean from 78.2 to 76.9 points. What does this sensitivity analysis reveal about outliers and distribution shape?

  1. The distribution is symmetric since both sets of extreme scores have similar impacts on the mean (correct answer)
  2. Right skewness dominates because high scores have greater influence on the central tendency measures
  3. Left skewness is evident since low scores create larger changes in the mean when removed
  4. The distribution is normal with both sets representing random measurement errors rather than true outliers
  5. Bimodal characteristics exist with outliers representing distinct performance groups requiring separate analysis
Explanation: When analyzing outliers and distribution shape, focus on the magnitude and direction of changes when extreme values are removed. This sensitivity analysis helps reveal whether a distribution is skewed and in which direction. Let's examine what happens when we remove extreme scores. Removing the 5 lowest scores changes the mean from 78.2 to 79.8, an increase of 1.6 points. Removing the 5 highest scores changes the mean from 78.2 to 76.9, a decrease of 1.3 points. The key insight is that both changes are remarkably similar in magnitude (1.6 vs 1.3 points), indicating that the extreme values on both ends have roughly equal influence on the central tendency. Why answer A is correct: The similar impact magnitudes (1.6 and 1.3 points) suggest the distribution is approximately symmetric. In a symmetric distribution, extreme values on both tails pull the mean away from the center with roughly equal force. Why the other options are wrong: B suggests right skewness dominates, but the low scores actually created a slightly larger change (1.6 vs 1.3). C claims left skewness based on the larger impact of removing low scores, but a 0.3-point difference is negligible and within normal variation. D incorrectly assumes the distribution is normal and dismisses these as measurement errors rather than recognizing this as a legitimate symmetry assessment. Study tip: When evaluating distribution shape through outlier removal, compare the magnitude of mean changes, not just the direction. Similar magnitudes suggest symmetry, while dramatically different impacts indicate skewness toward the more influential tail.

Question 6

A laboratory analyzes protein concentrations and reports that the data follows a distribution where Q1 = 12.5 mg/mL, median = 15.2 mg/mL, Q3 = 18.8 mg/mL, with a mean of 16.7 mg/mL. The laboratory also notes concentrations of 8.1, 9.2, 28.5, and 31.2 mg/mL in their dataset. What pattern emerges when considering both outliers and distributional asymmetry?

  1. Symmetric distribution with outliers equally balanced, explaining the close mean-median agreement
  2. Moderate right skewness with outliers on both ends, but high-end outliers having greater impact (correct answer)
  3. Strong left skewness dominated by the low concentration outliers pulling measures downward
  4. Normal distribution where the extreme values represent analytical measurement errors
  5. Bimodal pattern with outliers defining secondary peaks in the concentration distribution
Explanation: When analyzing distributions in biostatistics, you need to examine both the relationship between central tendency measures and the presence of outliers to understand the data's shape and behavior. Here, the mean (16.7 mg/mL) exceeds the median (15.2 mg/mL), indicating right skewness. The quartiles (Q1=12.5, Q3=18.8) show the data extends further toward higher values. Looking at the outliers, you have low values (8.1, 9.2 mg/mL) and high values (28.5, 31.2 mg/mL), but the high-end outliers are more extreme relative to the distribution's center. The mean being pulled above the median confirms that these upper outliers have greater influence on the distribution's shape. Option A is incorrect because symmetric distributions have mean ≈ median, but here the mean exceeds the median by 1.5 mg/mL. Option C misinterprets the skewness direction - if low outliers dominated, the mean would be below the median, not above it. Option D incorrectly assumes normality when the mean-median relationship clearly indicates asymmetry, and dismisses the outliers as errors rather than analyzing their impact. Option B correctly identifies moderate right skewness (mean > median) and recognizes that while outliers exist on both ends, the high-end outliers (28.5, 31.2) have greater impact on pulling the distribution's tail and elevating the mean above the median. Study tip: Always compare mean to median first - when mean > median, suspect right skewness. Then examine outlier locations and magnitudes to understand which ones drive the distributional asymmetry.

Question 7

A researcher examines a dataset of 200 patient cholesterol levels and finds that the mean is 185 mg/dL, the median is 180 mg/dL, and the mode is 175 mg/dL. Additionally, there are 8 values above 300 mg/dL in the dataset. What can be concluded about the distribution shape and presence of outliers?

  1. The distribution is left-skewed with potential outliers on the high end
  2. The distribution is right-skewed with confirmed outliers on the high end (correct answer)
  3. The distribution is approximately normal with no significant outliers present
  4. The distribution is right-skewed with potential outliers on the low end
  5. The distribution is bimodal with outliers on both ends of the spectrum
Explanation: When analyzing distribution shape, you need to examine the relationship between measures of central tendency and identify potential outliers based on extreme values. The key indicators here point to right skewness: the mean (185) is greater than the median (180), which is greater than the mode (175). In a right-skewed distribution, the tail extends toward higher values, pulling the mean upward beyond the median and mode. Additionally, the presence of 8 values above 300 mg/dL in a cholesterol dataset represents extreme values that are highly unusual - normal cholesterol levels rarely exceed 240 mg/dL, making values above 300 mg/dL confirmed outliers rather than just potential ones. Option A incorrectly identifies left skewness, which would show mean < median < mode - the opposite pattern from what we observe here. Option C is wrong because a normal distribution would have mean ≈ median ≈ mode, and the presence of 8 extreme high values clearly indicates significant outliers. Option D correctly identifies right skewness but incorrectly places the outliers on the low end, when the extreme values (>300 mg/dL) are clearly on the high end of the distribution. The correct answer is B because the mean-median-mode relationship confirms right skewness, and the 8 values above 300 mg/dL are definitively outliers given the clinical context of cholesterol measurements. Study tip: Remember the skewness pattern: right-skewed means mean > median > mode, while left-skewed means mode > median > mean. Always consider the clinical or practical context when determining whether extreme values constitute outliers.

Question 8

A dataset of 150 patient weights has a first quartile of 65 kg, median of 72 kg, and third quartile of 81 kg. Using the 1.5×IQR criterion, the calculated fences are at 41 kg and 105 kg. However, the actual data range extends from 45 kg to 125 kg with 3 observations above 110 kg. What is the most appropriate interpretation?

  1. No outliers exist since all values fall within the extended whisker range of typical box plots
  2. Three outliers are confirmed above 105 kg, and the distribution shows moderate right skewness (correct answer)
  3. The distribution is symmetric with outliers equally balanced, requiring no special treatment
  4. Outliers exist on both ends, but the high-end outliers are more statistically significant
  5. The fencing method is inappropriate here due to the apparent bimodal nature of the data
Explanation: When analyzing distributions for outliers, you need to systematically apply the 1.5×IQR criterion and assess distributional shape. The interquartile range here is 8165=1681 - 65 = 16 kg, making the outlier fences 651.5(16)=4165 - 1.5(16) = 41 kg and 81+1.5(16)=10581 + 1.5(16) = 105 kg. Since three observations fall above 110 kg, they clearly exceed the upper fence of 105 kg, confirming them as statistical outliers. The distribution shows right skewness because the outliers extend only in the upper direction (125 kg maximum vs. 45 kg minimum), and the tail stretches farther above the median than below it. Option A incorrectly assumes that falling within some "extended whisker range" negates outlier status, but the 1.5×IQR criterion definitively identifies values beyond the fences as outliers. Option C mischaracterizes the distribution as symmetric when the data clearly shows asymmetry with outliers concentrated on the high end only. Option D incorrectly claims outliers exist on both ends—while 45 kg is close to the lower fence of 41 kg, it doesn't actually cross it, so no lower outliers exist. The phrase "more statistically significant" in option D is also problematic since outlier detection isn't about statistical significance testing but about exceeding calculated thresholds. Study tip: Always calculate the exact fence values using 1.5×IQR, then check which observations actually cross these boundaries. Remember that outliers on just one side of the distribution indicate skewness in that direction.

Question 9

Examine the scatter plot showing the relationship between age and reaction time. What can be determined about outliers and their potential impact on distribution analysis?

  1. Multiple outliers exist primarily in younger age groups, creating artificial left skewness in reaction time
  2. Outliers appear randomly distributed across age groups with no systematic pattern affecting distribution shape
  3. Several outliers in older age groups may contribute to right skewness in reaction time distribution (correct answer)
  4. Outliers form a secondary cluster indicating a bimodal age distribution with normal reaction times
  5. No clear outliers are present, suggesting a normal distribution for both variables independently
Explanation: The scatter plot shows several points in the older age ranges (65-80 years) with unusually high reaction times that fall well above the general trend. These outliers in the upper age groups would contribute to right skewness in the reaction time distribution by extending the upper tail. Choice A incorrectly locates outliers in younger groups and misidentifies skew direction. Choice B ignores the clear clustering pattern. Choice D confuses reaction time outliers with age distribution patterns. Choice E fails to recognize the obvious outlying points.

Question 10

The frequency polygon shown displays test scores for two different classes. When analyzing outliers and distribution shape, what pattern emerges?

  1. Class A shows normal distribution with no outliers, while Class B exhibits right skewness with high-end outliers
  2. Both classes display symmetric distributions, but Class B has more extreme outliers on both ends
  3. Class A demonstrates left skewness with potential outliers, while Class B appears normally distributed
  4. Class A exhibits bimodal characteristics with outliers between modes, Class B shows uniform distribution
  5. Both distributions show right skewness, but Class A has fewer outliers than Class B (correct answer)
Explanation: Both frequency polygons show longer tails extending to the right, indicating right skewness in both classes. However, Class A has fewer isolated points in the extreme right tail compared to Class B, suggesting fewer outliers. The peaks of both distributions are shifted toward the lower scores with extended right tails. Choice A incorrectly identifies Class A as normal. Choice B incorrectly assumes symmetry. Choice C misidentifies the skew directions. Choice D incorrectly identifies bimodal and uniform patterns.

Question 11

In the dot plot displayed below, each dot represents one observation. What conclusion about distribution shape and outliers requires the most careful consideration?

  1. The apparent right skewness may be artificially created by a small number of extreme outliers (correct answer)
  2. The distribution appears symmetric, but outliers on the left side are creating statistical bias
  3. Multiple modes are present, with outliers helping to define distinct population subgroups
  4. The distribution is clearly uniform with a few random outliers that should be excluded
  5. Left skewness dominates the pattern, with right-side outliers being statistically insignificant
Explanation: The dot plot shows most data clustered in the lower values with a few isolated points extending far to the right. This pattern suggests that what appears to be right skewness might be primarily driven by these extreme outliers rather than a natural distributional shape. Removing these outliers might reveal a different underlying distribution. Choice B incorrectly identifies left-side outliers and symmetry. Choice C assumes multiple modes without clear evidence. Choice D incorrectly identifies uniformity. Choice E misidentifies the skew direction.

Question 12

The following comparative box plots show response times for three different experimental conditions. Based on the patterns shown, what is the most important observation about outliers across conditions?

  1. All conditions show identical outlier patterns, indicating consistent measurement precision across treatments
  2. Condition A has the most outliers, suggesting this treatment creates the most variable responses
  3. Outlier frequency decreases from Condition A to C, but the magnitude of extreme outliers increases
  4. Conditions B and C show no outliers, while A demonstrates normal experimental variation
Explanation: C

Question 13

A clinical study measures serum cholesterol levels in 300 patients. The histogram shows a distribution with the following characteristics: most values cluster between 180-220 mg/dL, there's a gradual decline in frequency from 220-280 mg/dL, and three isolated bars appear at 320, 340, and 380 mg/dL with frequencies of 2, 1, and 1 respectively. Which analytical approach would be most appropriate for handling these extreme values?

  1. Remove all values above 280 mg/dL as they represent data entry errors since cholesterol cannot reach such levels
  2. Apply logarithmic transformation to normalize the distribution before conducting any statistical analyses on the complete dataset
  3. Investigate the three highest values for clinical context while reporting results both with and without these potential outliers (correct answer)
  4. Use robust statistical methods that minimize outlier influence without removing any data points from the analysis
Explanation: The extreme values at 320-380 mg/dL are potential outliers but represent clinically possible cholesterol levels. Best practice involves investigating their clinical context and reporting sensitivity analyses. Choice A incorrectly assumes these are impossible values. Choice B applies transformation without first understanding the nature of extreme values. Choice D uses robust methods but doesn't address the need to understand whether these represent meaningful clinical findings versus measurement issues.

Question 14

A researcher calculates the following statistics for reaction times (milliseconds) in a cognitive test: Q1 = 245, Q3 = 315, IQR = 70. Using the 1.5×IQR rule, any value below 140 or above 420 would be considered an outlier. The dataset contains values at 425, 430, and 135 ms. Considering the biological context of human reaction times, how should these values be interpreted?

  1. All three values are statistical outliers and should be excluded since they exceed the calculated fences by substantial margins
  2. The high values (425, 430) represent plausible slow responses and should be retained, while 135 ms should be investigated as potentially impossible (correct answer)
  3. The low value (135) represents exceptional fast responses and should be retained, while the high values should be investigated as attention lapses
  4. All three values represent natural human variation in reaction times and should be retained despite being statistical outliers
Explanation: While all three values are statistical outliers by the 1.5×IQR rule, biological plausibility matters. Reaction times of 425-430 ms, though slow, are within human capability. However, 135 ms approaches the physiological limit for simple reactions and warrants investigation. Choice A mechanically applies statistical rules without biological context. Choice C incorrectly prioritizes the implausibly fast value. Choice D ignores that 135 ms may represent measurement error or invalid responses.

Question 15

Two datasets of patient ages both have identical means (45 years) and standard deviations (12 years), but their histograms show markedly different shapes. Dataset A appears roughly bell-shaped, while Dataset B shows a prominent peak at 35 years with a long tail extending toward higher ages. When applying the same outlier detection method (values beyond 2 standard deviations from the mean) to both datasets, what outcome should be expected?

  1. Both datasets will identify the same proportion of outliers since they have identical means and standard deviations regardless of distribution shape
  2. Dataset A will identify fewer outliers because normal distributions naturally contain fewer extreme values within the 2-SD range
  3. Dataset B will identify fewer outliers because the 2-SD rule is more appropriate for skewed distributions than normal distributions
  4. Dataset B may identify fewer true outliers because the 2-SD rule assumes normality and may misclassify tail values as outliers (correct answer)
Explanation: The 2-SD rule assumes normal distribution. In the right-skewed Dataset B, values in the natural tail of the distribution may be flagged as outliers when they're actually part of the expected distribution shape. Dataset A (normal) is appropriate for this method. Choice A ignores that distribution shape affects outlier detection validity. Choice B incorrectly suggests normal distributions have fewer extreme values within 2-SD. Choice C incorrectly claims the 2-SD rule is more appropriate for skewed distributions.

Question 16

Two research teams analyze the same dataset of tumor sizes but report different numbers of outliers. Team A uses the z-score method (|z| > 3) and identifies 2 outliers. Team B uses the modified z-score method with median absolute deviation and identifies 7 outliers. The dataset shows moderate right skewness. Which explanation most accurately accounts for this discrepancy?

  1. Team B made calculation errors since the modified z-score method should identify fewer outliers than the standard z-score method
  2. Team A's method is inappropriate for skewed data and failed to detect outliers that the more robust method identified correctly (correct answer)
  3. The difference reflects different statistical software packages using varying default thresholds for outlier classification methods
  4. Both methods are equally valid, and the difference represents the inherent subjectivity in outlier detection requiring clinical judgment
Explanation: The standard z-score method assumes normality and uses mean/SD, which are sensitive to outliers and skewness. In right-skewed data, outliers inflate the mean and SD, making the z-score method less sensitive. The modified z-score using median and MAD is robust and more appropriate for skewed data. Choice A incorrectly suggests Team B made errors. Choice C attributes the difference to software rather than methodological appropriateness. Choice D ignores that one method is clearly more suitable for skewed data.

Question 17

A researcher collects blood pressure measurements from 200 patients and finds that the distribution has a mean of 125 mmHg, median of 122 mmHg, and mode of 118 mmHg. Additionally, the 95th percentile is 165 mmHg while the 5th percentile is 95 mmHg. Based on this information, what can be concluded about the distribution shape and the likelihood of extreme values?

  1. The distribution is left-skewed, and values above 165 mmHg should be considered potential outliers requiring clinical investigation
  2. The distribution is right-skewed, and values above 165 mmHg should be considered potential outliers requiring clinical investigation
  3. The distribution is right-skewed, but values above 165 mmHg are within normal variation and not outliers (correct answer)
  4. The distribution is approximately normal, and outlier detection requires additional statistical measures beyond percentiles
Explanation: The mean (125) > median (122) > mode (118) indicates right skewness. However, the 95th percentile represents the upper 5% of the distribution, so values at or slightly above 165 mmHg are part of the expected range, not outliers. True outliers would be substantially beyond typical percentile ranges. Choice A has wrong skew direction. Choice B correctly identifies skewness but incorrectly classifies the 95th percentile as outlier threshold. Choice D incorrectly assumes normality when the mean-median-mode relationship clearly shows skewness.

Question 18

A laboratory reports the following summary statistics for white blood cell counts (×10³/μL) from 500 patients: mean = 7.2, median = 6.8, standard deviation = 2.1, skewness coefficient = +0.8. Three patients have values of 15.2, 16.8, and 17.1. Before determining whether these represent outliers, which additional information is most critical for proper interpretation?

  1. The clinical context and medical history of patients with the three highest values to assess biological plausibility (correct answer)
  2. The kurtosis coefficient to determine whether the distribution has heavier tails than expected for normal distributions
  3. The coefficient of variation to assess whether the standard deviation is appropriate for this type of measurement scale
  4. The 95th and 99th percentiles of the distribution to establish more appropriate thresholds than standard deviation-based methods
Explanation: When evaluating potential outliers in biomedical data, statistical methods alone are insufficient—you must consider the clinical context to determine whether extreme values represent meaningful biological variation or true anomalies requiring investigation. Option A is correct because outlier assessment in medical settings requires understanding whether extreme values are biologically plausible. White blood cell counts of 15.2-17.1 (×10³/μL) could represent normal physiological responses to infection, inflammation, stress, or certain medications, or they might indicate pathological conditions like leukemia. Without knowing the patients' clinical histories, you cannot determine if these values warrant concern or represent expected biological variation. Option B is incorrect because while kurtosis describes tail behavior, it doesn't provide the clinical context needed to interpret whether specific values are medically significant. Heavy tails might be expected in laboratory data due to underlying medical conditions. Option C is wrong because the coefficient of variation (CV = standard deviation/mean) assesses relative variability but doesn't help determine if specific values are clinically meaningful outliers. A CV calculation won't tell you whether 15.2 is a concerning finding or normal variation. Option D is incorrect because percentile-based thresholds, while statistically more appropriate than standard deviation methods for skewed data, still don't address the fundamental question of biological plausibility. Knowing that a value exceeds the 99th percentile doesn't indicate whether it's clinically significant. Study tip: In biostatistics questions involving outliers in medical data, always prioritize clinical interpretation over purely statistical approaches. Context matters more than mathematical thresholds when patient care is involved.

Question 19

A researcher analyzing medication adherence rates (percentage of prescribed doses taken) obtains the following five-number summary: Min = 12%, Q1 = 78%, Median = 89%, Q3 = 96%, Max = 100%. The distribution contains several values below 30%. Given the bounded nature of this variable and the summary statistics, how should these low values be interpreted?

  1. Values below 30% are extreme but clinically meaningful, representing patients with severe adherence problems requiring targeted intervention (correct answer)
  2. Values below 30% represent legitimate poor adherence cases, but the left-skewed distribution requires transformation before outlier assessment
  3. Values below 30% are statistical outliers by the IQR method and likely represent data collection errors requiring removal from analysis
  4. Values below 30% cannot be properly evaluated without knowing the exact sample size and distribution of values in this range
Explanation: When analyzing bounded variables like medication adherence percentages, you need to distinguish between statistical outliers and clinically meaningful extreme values. The key insight is that real-world constraints and clinical significance should guide your interpretation, not just statistical rules. Looking at this five-number summary, the data shows a left-skewed distribution with most patients having good adherence (Q1=78%, median=89%, Q3=96%). However, the minimum of 12% indicates some patients have very poor adherence. Since the question states there are "several values below 30%," these represent a meaningful subgroup of patients with severe adherence problems. Option A correctly recognizes that these low values, while extreme, are clinically meaningful and represent patients needing targeted intervention. Poor medication adherence is a well-documented clinical problem, and values below 30% indicate serious non-compliance requiring attention. Option B incorrectly assumes transformation is needed before outlier assessment. While the distribution is left-skewed, transformation isn't automatically required for bounded variables, especially when extreme values have clinical meaning. Option C falls into the trap of applying the IQR outlier rule mechanically. Using the 1.5×IQR rule: Q11.5(Q3Q1)=781.5(18)=51%Q1 - 1.5(Q3-Q1) = 78 - 1.5(18) = 51\%. Values below 51% would be statistical outliers, but this doesn't mean they're errors requiring removal. Option D suggests insufficient information, but the clinical context and summary statistics provide adequate information for interpretation. Study tip: With bounded variables in healthcare, always consider clinical significance before applying statistical outlier rules. Extreme values often represent important patient subgroups rather than data errors.

Question 20

Based on the histogram shown, which statement best describes the distribution's characteristics?

  1. Symmetric distribution with two potential outliers in the right tail
  2. Left-skewed distribution with outliers clustered in the middle range
  3. Right-skewed distribution with gaps suggesting potential data entry errors
  4. Bimodal distribution with the second mode created by outlying observations (correct answer)
  5. Uniform distribution with isolated extreme values on both ends
Explanation: The histogram shows two distinct peaks, indicating bimodality. The second, smaller peak appears to be formed by outlying observations that are separated from the main distribution. Choice A incorrectly identifies symmetry. Choice B misidentifies the skew direction and outlier location. Choice C suggests data errors without evidence. Choice E incorrectly identifies a uniform pattern.