Business Statistics Quiz: Data Quality Issues
20 questions · exam conditions
0:00
Data Quality IssuesQuestion 1 of 20

A management consulting firm assesses employee productivity by having consultants observe employees and rate their 'focus level' on a 1-10 scale. Analysis reveals that employees observed by a specific senior consultant received systematically lower focus scores than those observed by other consultants. This discrepancy most strongly suggests which data quality issue?

Low inter-rater reliability, which is a form of random measurement error.
The presence of outliers in the dataset corresponding to the senior consultant's ratings.
Systematic measurement error, specifically observer bias, related to the senior consultant.
A non-representative sample, as the employees observed by the senior consultant may have been different.
← Back to quizzes

Business Statistics Quiz

Business Statistics Quiz: Data Quality Issues

Practice Data Quality Issues in Business Statistics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Data Quality Issues, giving you a quick way to practice the rules, question types, and explanations that matter most for Business Statistics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A management consulting firm assesses employee productivity by having consultants observe employees and rate their 'focus level' on a 1-10 scale. Analysis reveals that employees observed by a specific senior consultant received systematically lower focus scores than those observed by other consultants. This discrepancy most strongly suggests which data quality issue?

  1. Low inter-rater reliability, which is a form of random measurement error.
  2. The presence of outliers in the dataset corresponding to the senior consultant's ratings.
  3. Systematic measurement error, specifically observer bias, related to the senior consultant. (correct answer)
  4. A non-representative sample, as the employees observed by the senior consultant may have been different.
Explanation: Observer bias occurs when the characteristics or expectations of the person collecting the data influence the measurements. Here, the scores are systematically lower for one specific observer. This indicates a bias, not just random variation between raters (which would be random error / low reliability). The pattern is consistent and tied to a single source, making it a systematic error.

Question 2

An analyst examining customer transaction data discovers that purchase amounts show a bimodal distribution with peaks at $49.99 and $99.99, while amounts between $50-$98 are notably sparse. Further investigation reveals that the company's pricing strategy uses these specific price points, and the data collection system automatically rounds any manually entered prices to the nearest $0.99 format. How should this measurement characteristic be handled in analysis?

  1. Treat this as a measurement error requiring correction, since the rounding introduces artificial clustering that distorts the true distribution of customer spending
  2. Treat this as an outlier detection issue where the peak values represent unusual purchasing behavior that should be investigated separately from normal transaction patterns
  3. Treat this as missing data problem where the 'true' prices between the peaks are unobserved and require imputation to restore the underlying continuous distribution
  4. Treat this as valid business data requiring analysis techniques appropriate for discrete rather than continuous variables due to the systematic price constraints (correct answer)
Explanation: When analyzing business data, you need to distinguish between data artifacts that represent measurement problems versus legitimate business constraints that shape the actual data-generating process. This question tests your ability to recognize when apparent "irregularities" in data distributions are actually meaningful business phenomena. The bimodal distribution at $49.99 and $99.99 reflects genuine business reality - the company's deliberate pricing strategy creates these natural clustering points, and the automated rounding system is part of their operational process. This isn't a flaw to fix but rather the actual mechanism by which prices are determined. Since customer purchases can only occur at these systematically constrained price points, the data is inherently discrete rather than continuous, making option D correct. Option A incorrectly treats systematic business rules as measurement error. The rounding isn't distorting some "true" underlying distribution - it's creating the actual distribution through legitimate business processes. Option B misidentifies the peaks as outliers when they're actually the most common, expected values given the pricing strategy. Option C assumes there are "true" missing prices between peaks that need restoration, but those intermediate prices literally cannot exist in this business system - there's nothing missing to impute. The key insight is that business constraints often create discrete distributions even when the underlying economic variable (customer willingness to pay) might be continuous. Your analytical approach must match the data's true nature, not impose assumptions about what you think it should look like. Study tip: Always investigate whether apparent data irregularities reflect real business processes before assuming they're problems requiring correction.

Question 3

A human resources analyst studying employee turnover finds that exit interview data is missing for 40% of departing employees. Among employees who completed exit interviews, 85% cite 'career growth' as their primary reason for leaving, while industry benchmarks suggest this reason typically accounts for only 35% of departures. What data quality issue is most likely affecting the analysis?

  1. Systematic measurement error due to leading questions in the exit interview process that bias responses toward socially acceptable answers
  2. Sampling bias due to inadequate representation of different employee demographics in the exit interview volunteer pool
  3. Random measurement error due to inconsistent administration of exit interviews across different departments and time periods
  4. Selection bias combined with response bias, where employees who complete interviews may differ systematically from those who don't, and may provide socially desirable responses (correct answer)
Explanation: When analyzing data quality issues, you need to identify whether problems stem from how data is collected, who provides it, or how it's measured. This question presents a classic scenario where multiple biases compound to create misleading results. The key insight is recognizing that 40% missing data combined with results dramatically different from industry benchmarks (85% vs 35%) suggests multiple systematic problems. Answer D correctly identifies this as selection bias (certain types of employees are more likely to complete exit interviews) combined with response bias (those who do participate may give socially acceptable answers rather than truthful ones). Employees citing "career growth" sounds more professional than admitting they're fleeing bad management or toxic culture. Answer A focuses only on measurement error from leading questions, but doesn't address the substantial missing data problem that's likely creating the bigger distortion. Answer B mentions sampling bias but incorrectly attributes it to demographic representation rather than the fundamental difference between employees who do versus don't complete interviews. Answer C suggests random measurement error from inconsistent administration, but the consistently high percentage citing career growth indicates systematic rather than random error. The 40% missing data rate is the crucial clue here. When combined with results that deviate so significantly from industry norms, this strongly suggests that the employees completing interviews are fundamentally different from those who don't participate, and they're providing responses that paint their departure in the best possible light. Study tip: When you see large amounts of missing data paired with unexpected results, always consider whether the missing data is random or if certain groups systematically opt out.

Question 4

A market researcher notices that income data in a consumer survey shows suspicious patterns: exactly 47% of respondents reported incomes ending in '000' (like $50,000, 75,000),andtheincomedistributionshowsunusualspikesatroundnumbers.Additionally,1275,000), and the income distribution shows unusual spikes at round numbers. Additionally, 12% of high-income respondents (>150,000) refused to answer the income question, while only 3% of lower-income respondents had missing income data. What analytical strategy best addresses these data quality issues?

  1. Use income ranges instead of exact values and apply multiple imputation for missing data, treating missingness as Missing Not at Random (MNAR) (correct answer)
  2. Use income ranges instead of exact values and apply multiple imputation for missing data, treating missingness as Missing at Random (MAR)
  3. Remove all observations with round-number incomes and use listwise deletion for missing income data to ensure analysis uses only precise measurements
  4. Apply smoothing techniques to eliminate round-number spikes and use hot-deck imputation to replace missing values with similar respondents' income data
Explanation: When you encounter data quality issues in survey research, you need to recognize different types of measurement problems and missing data patterns to choose appropriate analytical strategies. The scenario presents two key issues: heaping (respondents rounding to convenient numbers) and systematic missingness (high earners are more likely to skip income questions). The 47% reporting round numbers indicates substantial heaping, making exact income values unreliable. More critically, the differential missing rates (12% vs 3%) suggest the missingness depends on the unobserved income values themselves—high earners systematically avoid disclosure regardless of other observed characteristics. Answer A correctly addresses both problems. Income ranges reduce heaping bias by acknowledging that many responses aren't precise anyway. Multiple imputation with Missing Not at Random (MNAR) assumptions is essential because the probability of missing data depends on the actual (unobserved) income level, not just observed respondent characteristics. Answer B fails because treating this as Missing at Random (MAR) assumes missingness only depends on observed variables, ignoring the clear systematic pattern where income level itself drives non-response. Answer C discards too much valuable data—removing nearly half your sample (round numbers) and all high-income non-responders creates severe bias and reduces power. Answer D's smoothing techniques artificially manipulate the underlying distribution, while hot-deck imputation doesn't account for the systematic missingness pattern. Remember: when missing data rates vary dramatically across the variable of interest itself, always suspect MNAR and choose methods that explicitly model this dependency.

Question 5

An operations manager finds that employee productivity measurements are missing for 25% of the workforce, with missing data concentrated among employees who work flexible schedules or remote arrangements. The manager plans to use listwise deletion (complete case analysis) for upcoming productivity analysis. What is the most significant concern with this approach?

  1. The analysis will have reduced statistical power due to smaller sample size, but parameter estimates will remain unbiased since the missing data appears random
  2. The analysis will produce biased estimates of average productivity because flexible and remote workers likely have different productivity patterns than office workers (correct answer)
  3. The analysis will violate normality assumptions required for most statistical tests due to the systematic pattern of missing observations
  4. The analysis will have increased Type I error rates because listwise deletion artificially inflates the correlation between productivity and work arrangement variables
Explanation: Listwise deletion with non-random missing data (MNAR/MAR) introduces bias because the remaining sample no longer represents the full population. Since flexible/remote workers may have different productivity patterns, excluding them systematically biases estimates. Choice A is wrong because the missing data is not random, so estimates will be biased. Choice C is wrong because missing data patterns don't directly affect normality assumptions of the analyzed variables. Choice D is wrong because listwise deletion typically doesn't inflate Type I error rates in this manner.

Question 6

A quality control specialist analyzing manufacturing defect rates notices that measurement precision varies systematically across production shifts: day shift measurements are recorded to 0.01% precision, evening shift to 0.1% precision, and night shift to 1% precision due to different calibration protocols. This measurement inconsistency is most likely to create which analytical problem?

  1. Heteroscedasticity in regression models, where error variance increases with less precise measurements from night shift data (correct answer)
  2. Multicollinearity between shift variables and defect rates, making it difficult to isolate the true effect of production timing
  3. Autocorrelation in time series analysis, where measurement errors become correlated across consecutive time periods within shifts
  4. Selection bias in comparative analysis, where apparent differences between shifts reflect measurement quality rather than actual performance
Explanation: Different measurement precision across shifts creates heteroscedasticity because the variance of measurement errors differs systematically (night shift has larger measurement variance than day shift). This violates the constant variance assumption in regression. Choice B is wrong because multicollinearity refers to correlation between predictor variables, not measurement precision issues. Choice C is wrong because autocorrelation involves temporal correlation, not precision differences. Choice D describes a valid concern but is less direct than the heteroscedasticity issue.

Question 7

A marketing firm is building a predictive model for customer churn. The dataset has missing values for the 'customer_tenure' variable, which is a key predictor. An analyst uses mean imputation to fill these missing values before training the model. What is the most likely consequence of this action on the final model?

  1. The model will become overfitted to the mean tenure value, leading to poor generalization on new data.
  2. The predictive power of 'customer_tenure' will be artificially weakened, potentially reducing the model's overall accuracy. (correct answer)
  3. The standard error of the 'customer_tenure' coefficient will decrease, suggesting a more precise estimate than is actually the case.
  4. Multicollinearity will be introduced if the mean tenure value is correlated with other predictor variables in the model.
Explanation: Mean imputation replaces all missing values with a single constant (the mean). This artificially reduces the natural variation in the 'customer_tenure' variable. By concentrating many data points at the mean, it weakens the observed relationship between tenure and churn, as the model sees many instances with the same tenure value but different churn outcomes. This dilution of the relationship reduces the variable's predictive power.

Question 8

A company uses a new employee engagement survey. When administered to the same group of employees one month apart, the survey yields highly consistent scores for each employee. However, the survey scores show no significant correlation with employee turnover rates, a key indicator the survey was intended to predict. Which statement best describes the quality of the survey data?

  1. The survey has high validity but low reliability.
  2. The survey has low validity and low reliability.
  3. The survey has high reliability but low validity. (correct answer)
  4. The survey suffers from systematic error but not random error.
Explanation: Reliability refers to the consistency of a measurement. Since the survey gives consistent scores over time, it has high reliability. Validity refers to whether the instrument measures what it is intended to measure. Since the scores do not correlate with a relevant outcome (turnover), its predictive validity is low. Therefore, the survey is consistent but does not appear to be measuring the construct of 'engagement' in a way that is useful for prediction.

Question 9

A hospital is analyzing patient recovery times. Data on 'post-operative complications' is frequently missing. An administrator discovers that patients who experience severe complications are often too ill to complete the follow-up paperwork where this data is recorded. How should this missing data mechanism be classified, and what is a major risk if records with missing complication data are simply deleted?

  1. MCAR; the analysis will have less statistical power but the results will remain unbiased.
  2. MAR; the bias can be corrected using imputation based on other patient variables like age or gender.
  3. MNAR; the analysis will be biased, likely underestimating the average recovery time and rate of complications. (correct answer)
  4. MNAR; the sample size reduction is the primary issue, but estimates will be more precise for the remaining cohort.
Explanation: The missingness is directly related to the unobserved value itself (patients with severe complications are more likely to have missing data on complications). This is the definition of Missing Not At Random (MNAR). If these records are deleted (listwise deletion), the most severe cases are systematically removed from the dataset. This will lead to a biased analysis that underrepresents the negative impact of complications, making recovery times appear shorter and complication rates seem lower than they truly are.

Question 10

An analyst is screening a dataset of 1,000 transaction values for outliers. The data distribution is symmetric but has heavier tails than a normal distribution (leptokurtic). The analyst considers using either the Z-score method (flagging values where Z>3|Z| > 3) or the IQR method (flagging values outside Q11.5×IQRQ_1 - 1.5 \times IQR and Q3+1.5×IQRQ_3 + 1.5 \times IQR). Which statement accurately compares the expected outcomes?

  1. Both methods will identify the exact same set of outliers because the data is symmetric.
  2. The Z-score method may flag fewer outliers because the heavy tails will inflate the standard deviation used in its calculation. (correct answer)
  3. The IQR method is inappropriate for symmetric data; the Z-score method should always be preferred in this case.
  4. The Z-score method will flag more outliers because the heavy tails mean more data points are far from the mean.
Explanation: The Z-score is calculated as (xμ)/σ(x - \mu) / \sigma. In a distribution with heavy tails, the extreme values will inflate the calculated standard deviation (σ\sigma). This larger denominator will result in smaller Z-scores for all points, including the extreme ones. Consequently, fewer points may cross the Z>3|Z| > 3 threshold. The IQR method, based on quartiles, is robust to extreme values and will not be similarly affected, making it more likely to flag outliers in a heavy-tailed distribution.

Question 11

A financial analyst is comparing the quarterly returns of two portfolios. Portfolio A's returns are approximately normally distributed. Portfolio B's returns have a similar mean but are subject to occasional extreme positive and negative events, creating a distribution with fat tails. To accurately compare the central tendency and risk of the two portfolios, which pair of statistics would be most appropriate and robust?

  1. Mean for central tendency and Standard Deviation for risk.
  2. Median for central tendency and Standard Deviation for risk.
  3. Mean for central tendency and Interquartile Range (IQR) for risk.
  4. Median for central tendency and Interquartile Range (IQR) for risk. (correct answer)
Explanation: For Portfolio B, the presence of extreme events (outliers) makes the mean and standard deviation unreliable. The mean is sensitive to extreme values, and the standard deviation is even more so. Robust statistics are needed. The median is a robust measure of central tendency, and the interquartile range (IQR) is a robust measure of dispersion. For a fair comparison, the same set of statistics should be used for both portfolios. Therefore, median and IQR are the most appropriate choices.

Question 12

A scatter plot of employee performance score versus age shows a weak, positive, linear association for employees aged 20-60. The dataset also includes one 65-year-old senior executive with a very high performance score. This data point is located far from the other points but in the same general direction of the trend. How will the inclusion of this single outlier most likely affect the Pearson correlation coefficient (r)?

  1. It will substantially increase the correlation coefficient, making the linear relationship appear stronger than it is for the bulk of the data. (correct answer)
  2. It will decrease the correlation coefficient, as all outliers weaken the measure of linear association.
  3. It will have a negligible effect on the correlation coefficient because it is only a single data point.
  4. It will shift the correlation coefficient closer to zero, indicating a weaker relationship.
Explanation: This type of outlier is an 'influential point'. Because it lies far from the mean of the predictor variable (age) and reinforces the existing trend (high age, high performance), it will 'pull' the regression line towards it and increase the magnitude of the Pearson correlation coefficient. It creates a stronger statistical relationship than what is representative of the rest of the employees.

Question 13

A city's economic development office wants to measure the 'financial health' of its small business community. Lacking direct access to revenue data, they decide to use the 'number of social media posts per week by businesses' as a proxy variable. What is the most significant potential data quality issue with using this proxy?

  1. Low reliability, because the number of posts can be counted inconsistently across different social media platforms.
  2. A potential lack of construct validity, as social media activity may not accurately reflect a business's actual financial performance. (correct answer)
  3. The presence of outliers, such as a single business 'going viral' and posting hundreds of times in one week.
  4. Systematic bias, as businesses in certain industries (e.g., retail) are naturally more active on social media than others (e.g., manufacturing).
Explanation: Construct validity is the degree to which a variable measures the theoretical construct it is intended to measure. Here, the construct is 'financial health'. While there might be some correlation, social media activity is not a direct or necessarily accurate reflection of revenue, profit, or stability. A struggling business might increase posting out of desperation, while a highly successful one might not need to. This mismatch between the proxy and the construct is the most fundamental issue.

Question 14

A company surveys employees, asking them to rate their job satisfaction and also whether they are actively seeking a new job. The company finds that employees who report lower job satisfaction are significantly more likely to leave the 'actively seeking' question blank. How would an analyst classify this pattern of missing data?

  1. Missing Completely At Random (MCAR), as the responses are missing due to individual employee choice.
  2. Missing Not At Random (MNAR), as the missingness is related to the unobserved value of actively seeking a job.
  3. Missing At Random (MAR), as the probability of missingness depends on another observed variable in the dataset. (correct answer)
  4. Systematic measurement error, as the survey instrument is failing to capture data from a specific subgroup.
Explanation: This scenario is a classic example of Missing At Random (MAR). The data is not missing completely at random, because the likelihood of a missing value depends on another variable. However, it is not necessarily MNAR because the missingness depends on an observed variable (job satisfaction), not the unobserved value of the 'actively seeking' variable itself. With MAR, the reason for missingness is captured by other information in the dataset, which can potentially be used in more advanced imputation methods.

Question 15

An analyst receives a small dataset of 20 executive salaries. One value is recorded as $25,000, while all others are between $250,000 and $750,000. One other salary value is missing entirely. The analyst removes the $25,000 value, judging it a data entry error, and then imputes the missing value using the mean of the remaining 18 data points. Which of the following statements best evaluates this two-step procedure?

  1. The procedure is flawed; the missing value should have been imputed first before the outlier was removed.
  2. The procedure is optimal; removing an obvious error and then using mean imputation is the standard best practice.
  3. The outlier should have been retained to avoid data manipulation, and the missing value should have been left blank.
  4. Removing the outlier is a justifiable step, but mean imputation may be suboptimal if the remaining salary distribution is skewed. (correct answer)
Explanation: The first step, removing the $25,000 value, is a reasonable action. Given the context of executive salaries, it is highly probable this is a typographical error. Leaving it in would severely bias any calculations. However, the second step is questionable. Executive salary data is typically right-skewed. Using the mean for imputation is sensitive to this skew. A more robust approach would be to use median imputation, which would be less affected by the high salaries at the upper end of the remaining distribution.

Question 16

A call center wants to assess agent performance using an AI algorithm that scores transcribed calls for 'empathy'. To test the algorithm, the manager feeds the same 100 call recordings through the system on five different occasions. The resulting 'empathy' scores for each individual call vary significantly across the five runs. This variability points to a significant problem with which aspect of data quality?

  1. Low reliability, as the measurement tool does not produce consistent results under identical conditions. (correct answer)
  2. Low validity, as the AI's definition of 'empathy' may not align with a human's.
  3. Sampling bias, as the 100 calls selected may not be representative of all calls.
  4. Systematic error, as the algorithm may be consistently over- or under-scoring empathy across all calls.
Explanation: When evaluating data quality in business statistics, you need to distinguish between reliability (consistency) and validity (accuracy). This question tests your ability to diagnose measurement problems based on the symptoms described. The key clue here is that the same 100 recordings produced significantly different empathy scores across five identical runs. This inconsistency under identical conditions is the textbook definition of poor reliability. A reliable measurement tool should produce nearly identical results when measuring the same thing repeatedly. Think of a scale that shows different weights each time you step on it within minutes - that's a reliability problem. Answer A correctly identifies this as a reliability issue because the algorithm fails to produce consistent results under identical conditions. Answer B describes a validity problem, which would involve the algorithm consistently measuring something other than what we intend (empathy). However, the question doesn't suggest the AI is measuring the wrong thing - just that it's measuring inconsistently. Answer C focuses on sampling bias, but the problem isn't about whether the 100 calls represent all calls well. The issue is that the same calls produce different scores each run. Answer D suggests systematic error, which would show up as consistent over- or under-scoring across all measurements. But systematic errors are typically consistent, not variable like described here. Remember: Reliability = consistency, Validity = accuracy. When you see measurement results varying dramatically under identical conditions, think reliability first.

Question 17

In analyzing its sales data, a company's business intelligence team encounters three distinct issues:

  1. One sales entry is 100 times larger than any other, which they attribute to a misplaced decimal point.

  2. For 10% of entries, the 'customer region' field is blank because a legacy data entry system often crashed when used by new sales associates.

  3. The 'sales amount' is self-reported by salespersons who informally round to the nearest hundred dollars. Which sequence correctly classifies these three data quality issues?

  1. Outlier, followed by Missing At Random (MAR) data, followed by measurement error. (correct answer)
  2. Measurement error, followed by Missing Completely At Random (MCAR) data, followed by an outlier.
  3. Outlier, followed by Missing Not At Random (MNAR) data, followed by sampling bias.
  4. Systematic error, followed by random error, followed by measurement error.
Explanation:
  1. The extremely large sales entry is an outlier, likely due to a data entry error. 2. The missing 'customer region' data is dependent on an observable characteristic (the newness of the sales associate), making it Missing At Random (MAR). It's not MCAR because it's not completely random. 3. The practice of rounding introduces a discrepancy between the true sales amount and the recorded amount, which is a form of measurement error (specifically, reduced precision).

Question 18

A financial analyst examining quarterly earnings data notices that 12 companies in a dataset of 200 have earnings-per-share (EPS) values that are more than 4 standard deviations above the mean, while only 1 company has EPS more than 4 standard deviations below the mean. Normal distribution theory predicts approximately 0.01% of observations beyond 4 standard deviations in each tail. What data quality assessment should the analyst prioritize?

  1. Investigate potential data entry errors for the 12 high-EPS companies, as this asymmetric pattern suggests systematic input mistakes
  2. Investigate measurement validity issues, as the asymmetric extreme value pattern likely reflects legitimate business phenomena rather than data errors
  3. Investigate both potential data errors and legitimate outliers, as the asymmetric pattern could reflect either measurement issues or true skewness in earnings distributions (correct answer)
  4. Investigate sampling bias issues, as the asymmetric pattern suggests the sample may not be representative of the broader population
Explanation: The analyst should investigate both possibilities because earnings data can legitimately exhibit positive skewness (some companies can have exceptionally high earnings while earnings cannot go infinitely negative), but the extreme asymmetry also warrants checking for data entry errors. Choice A assumes the pattern must be errors without considering that earnings distributions are often naturally skewed. Choice B assumes the pattern must be legitimate without verifying data quality. Choice D focuses on sampling bias when the issue is about extreme values in the collected sample.

Question 19

An analyst examining sales performance data across regions discovers that the Western region consistently reports sales figures with much higher precision (to the cent) compared to other regions, which report rounded values to the nearest $100. Additionally, the Western region has a 5% missing data rate while other regions have 15-20% missing data rates. When comparing regional performance, what methodological concern should receive highest priority?

  1. The precision differences will create heteroscedasticity in regression models, making statistical inference unreliable across regions
  2. The missing data rate differences suggest systematic reporting quality issues that may confound regional performance comparisons
  3. Both precision and missing data differences indicate that regional comparisons may reflect data quality rather than actual performance differences (correct answer)
  4. The precision differences will cause attenuation bias in correlation analyses, understating relationships between variables in lower-precision regions
Explanation: Both issues together create a situation where apparent regional differences might reflect data collection quality rather than true performance differences, making any regional comparison potentially misleading. Choice A identifies a real statistical issue but doesn't address the broader interpretation problem. Choice B focuses only on missing data while ignoring precision issues. Choice D correctly identifies attenuation bias but doesn't address the fundamental problem that regional comparisons may be confounded by data quality differences.

Question 20

An analyst is studying the effect of a new training program by comparing the performance scores of employees who completed the training with those who did not. The training was optional. The analyst finds that 20% of the scores are missing from the 'trained' group, while only 5% are missing from the 'untrained' group. It is likely that lower-performing employees in the 'trained' group were less inclined to report their scores. What is the most significant risk in this analysis?

  1. The different rates of missingness will reduce the statistical power of the comparison between the two groups.
  2. The data are MAR, and listwise deletion will cause the effect of the training program to be underestimated.
  3. The data are MNAR, and listwise deletion will likely cause the average score of the 'trained' group to be artificially inflated. (correct answer)
  4. The measurement of performance scores is unreliable, as evidenced by the high number of missing values.
Explanation: The missingness of scores is related to the scores themselves (lower-performers not reporting), which is the definition of Missing Not At Random (MNAR). If the analyst uses listwise deletion, they will disproportionately remove the lower scores from the 'trained' group. This will artificially inflate the average score of the remaining members of the 'trained' group, potentially making the training program look far more effective than it actually was.