EPPP: Part 2, Skills Quiz: Research Appraisal
20 questions ยท exam conditions
0:00
Research AppraisalQuestion 1 of 20

A behavior analyst uses a multiple-baseline-across-subjects design to evaluate a new intervention for reducing self-injurious behavior (SIB) in three adolescents. The analyst collects baseline data on SIB for all three participants. The intervention is then introduced for Participant 1, while baseline conditions continue for Participants 2 and 3. After a stable change is observed in Participant 1, the intervention is introduced for Participant 2, and so on.

Which of the following patterns of results would provide the strongest evidence for a causal link between the intervention and the reduction in SIB?

A gradual reduction in SIB is observed in all three participants, starting from the first day of the study and continuing steadily.
A sharp reduction in SIB occurs for all three participants as soon as the intervention is implemented for Participant 1.
For each participant, a sharp reduction in SIB occurs only after the intervention is specifically introduced for that individual.
SIB is reduced for Participant 1 after the intervention, but the behavior of Participants 2 and 3 remains unchanged throughout the study.
โ† Back to quizzes

EPPP: Part 2, Skills Quiz

EPPP: Part 2, Skills Quiz: Research Appraisal

Practice Research Appraisal in EPPP: Part 2, Skills with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Research Appraisal, giving you a quick way to practice the rules, question types, and explanations that matter most for EPPP: Part 2, Skills.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A behavior analyst uses a multiple-baseline-across-subjects design to evaluate a new intervention for reducing self-injurious behavior (SIB) in three adolescents. The analyst collects baseline data on SIB for all three participants. The intervention is then introduced for Participant 1, while baseline conditions continue for Participants 2 and 3. After a stable change is observed in Participant 1, the intervention is introduced for Participant 2, and so on.

Which of the following patterns of results would provide the strongest evidence for a causal link between the intervention and the reduction in SIB?

  1. A gradual reduction in SIB is observed in all three participants, starting from the first day of the study and continuing steadily.
  2. A sharp reduction in SIB occurs for all three participants as soon as the intervention is implemented for Participant 1.
  3. For each participant, a sharp reduction in SIB occurs only after the intervention is specifically introduced for that individual. (correct answer)
  4. SIB is reduced for Participant 1 after the intervention, but the behavior of Participants 2 and 3 remains unchanged throughout the study.
Explanation: The correct answer is C. The logic of the multiple-baseline design is to demonstrate experimental control by showing that the behavior changes only when the intervention is applied. A staggered introduction of the treatment across different subjects, settings, or behaviors serves as its own control. The strongest evidence of causality is when the change in the dependent variable is tightly linked in time to the introduction of the independent variable for each separate baseline. A would suggest a maturation or history effect. B would strongly suggest a history effect or diffusion of treatment. D would show the intervention worked for one participant but failed to generalize to others, weakening the overall claim of effectiveness.

Question 2

A psychologist is investigating a potential link between adverse childhood experiences (ACEs) and the development of personality disorders in adulthood. The study protocol involves recruiting a sample of adults already diagnosed with borderline personality disorder and a control group without the diagnosis. All participants are then administered a detailed questionnaire asking them to recall and report on the frequency and severity of various traumatic events from their childhood.

What is the most significant methodological limitation inherent in this research design for establishing a link between ACEs and the disorder?

  1. The correlation found between ACEs and the disorder does not prove that the ACEs caused the disorder.
  2. The reliance on retrospective recall is subject to memory biases that may be systematically different between the two groups. (correct answer)
  3. The sample size may not be large enough to detect a statistically significant relationship between the variables.
  4. The definition of adverse childhood experiences may not be consistent with the way participants understand their own history.
Explanation: The correct answer is B. This is a retrospective case-control design. Its greatest weakness is its reliance on retrospective data, which is prone to recall bias. This bias is especially problematic here because an individual's current mental state (e.g., having BPD) can significantly influence how they remember, interpret, and report past events. The BPD group might recall more negative events, or interpret ambiguous events more negatively, than the control group, creating a spurious association. A is a general limitation of correlational designs, but B points to a more specific and potent flaw in this particular data collection method. C is a question of power, not a flaw in the design itself. D is a measurement concern, but recall bias is a more fundamental design issue.

Question 3

A school psychologist identifies the 10% of students with the highest scores on a validated measure of test anxiety. These students are enrolled in a newly developed, eight-week anxiety reduction program. At the end of the program, a re-administration of the test anxiety measure shows a statistically significant decrease in the group's average score. The psychologist concludes the program was effective.

When critically evaluating the psychologist's conclusion, which of the following is the most significant threat to the study's internal validity?

  1. History, as an external event like a change in school exam policy could have occurred during the eight-week program.
  2. Testing effect, as students' familiarity with the anxiety measure from the pre-test may have influenced their responses on the post-test.
  3. Maturation, as students may have naturally become less anxious over the eight-week period due to normal developmental processes.
  4. Statistical regression, as an extreme group selected on a pre-test will, on average, score closer to the mean on a subsequent post-test. (correct answer)
Explanation: The correct answer is D. Statistical regression to the mean is a major threat to internal validity when participants are selected based on extreme scores. Because of measurement error, individuals who score very high on a pre-test are likely to score lower (i.e., closer to the population mean) on a post-test, regardless of any intervention. This artifact could fully account for the observed decrease in anxiety scores. A is incorrect because while history is a possible threat in any pre-post design without a control group, the selection of an extreme group makes statistical regression the most salient and probable threat. B is incorrect because while a testing effect is possible, it is less likely to account for a large, systematic decrease in scores than regression to the mean in this specific design. C is incorrect because maturation is a less specific explanation than statistical regression, which is directly tied to the method of selecting participants based on extreme scores.

Question 4

A research team investigates the efficacy of a new antidepressant medication. In a double-blind, placebo-controlled trial, the sole outcome measure used to assess changes in depressive symptoms is the Beck Depression Inventory-II (BDI-II), a self-report questionnaire.

When critiquing the methodology of this study, which of the following represents the most significant concern regarding the assessment of the primary outcome?

  1. Instrumentation threat, because repeated use of the BDI-II could lead to changes in how the instrument is interpreted by participants over time.
  2. Mono-operation bias, because relying on a single self-report measure fails to capture the multifaceted construct of depression. (correct answer)
  3. Low reliability, because self-report measures like the BDI-II are known to be less consistent than clinician-administered interviews.
  4. Demand characteristics, because participants completing a self-report measure may guess the study's purpose and alter their responses accordingly.
Explanation: The correct answer is B. Mono-operation bias is a threat to construct validity that occurs when only a single operationalization of a construct is used. Depression is a complex construct with behavioral, cognitive, affective, and physiological components. By using only the BDI-II (a single self-report measure), the study may be inadequately capturing the full scope of the construct, thus limiting the validity of the conclusions about the drug's effect on 'depression' as a whole. A is incorrect because an instrumentation threat refers to changes in the measuring instrument itself or its administration, not just repeated use. C is incorrect as the BDI-II is a well-established instrument with good reliability; the issue is its narrowness, not its consistency. D is a potential issue with any non-blinded measure, but the core methodological flaw in the construct measurement strategy is the reliance on a single operationalization.

Question 5

A large corporation implements a mandatory mindfulness meditation program to reduce employee stress. A consultant psychologist is hired to evaluate its effectiveness. The psychologist obtains weekly stress rating data from all employees for one year prior to the program's implementation and for one year following it. The data show a sharp decline in average stress levels immediately after the program's start, which is sustained for the following year. There was no control group.

This interrupted time-series design is strongest at controlling for which of the following threats to internal validity?

  1. History, because any significant external event would be visible as a separate discontinuity in the data series.
  2. Instrumentation, because the measurement of stress was consistent across the entire two-year period of data collection.
  3. Maturation, because the pre-existing trend in stress levels is established before the intervention begins. (correct answer)
  4. Selection, because the entire population of employees was included, eliminating differences between groups.
Explanation: The correct answer is C. A major strength of the interrupted time-series design is its ability to control for maturation. By collecting multiple data points before the intervention, the researcher can establish the natural trend of the dependent variable over time. If the post-intervention trend is different from the pre-intervention trend, it is less likely that the change is due to simple maturation. A is incorrect because history is actually the primary weakness of this design; a single, concurrent event (e.g., a company-wide bonus announcement) could perfectly co-occur with the intervention and explain the change. B is an assumption of the study, not a threat controlled by the design itself. D is incorrect because selection threats are relevant when comparing groups, and this design does not have a comparison group.

Question 6

A psychologist working at a university counseling center is considering implementing a new, brief intervention for academic procrastination that has shown excellent results in a recent study. The study was conducted with highly motivated, high-achieving undergraduate volunteers at an elite private university. The psychologist's clients, however, are often mandated to attend counseling due to academic probation and come from a large, public state university with a diverse student body.

In evaluating the generalizability of the study's findings to the clinic's population, which is the most critical question the psychologist must consider?

  1. Could differences in student motivation and academic standing between the samples act as moderators of the intervention's effect? (correct answer)
  2. Was the effect size reported in the original study large enough to suggest clinical significance?
  3. Did the original study use random assignment and an appropriate control group to establish internal validity?
  4. Was the original study published in a high-impact, peer-reviewed journal, indicating a rigorous review process?
Explanation: The correct answer is C. This question is about external validity, specifically the applicability of research from one population to another. The most critical issue is whether key differences between the research sample (highly motivated, high-achieving volunteers) and the target population (mandated, struggling students) could moderate the treatment's effectiveness. It is plausible that an intervention requiring high motivation will not work for a population with low motivation. A addresses internal validity, which is important, but not the central issue of generalizability. B is a relevant question, but it is secondary to whether the effect will be present at all in the new population. D relates to the overall quality of the study but does not directly address the specific population mismatch.

Question 7

A clinical psychology doctoral student designed a novel intervention to improve emotion regulation in adolescents. In her dissertation study, she randomly assigned participants to her new intervention or a waitlist control group. She personally conducted all intervention sessions and also administered and scored the primary outcome measure, a behavioral observation task of emotion regulation, for all participants in both groups at post-test.

Which of the following is the most prominent and direct threat to the study's validity introduced by this assessment procedure?

  1. Demand characteristics, as participants in the intervention group may feel pressured to perform better on the post-test.
  2. Attrition, as participants in the waitlist control group may be more likely to drop out of the study before the post-test.
  3. Experimenter expectancy effects, as the student's investment in the intervention could unconsciously bias her scoring of the behavioral task. (correct answer)
  4. Selection bias, as the student may have unintentionally placed more promising participants into her intervention group.
Explanation: The correct answer is C. Experimenter expectancy (or Rosenthal) effects occur when a researcher's beliefs or expectations about the outcome of a study unconsciously influence how they treat participants or measure the outcome. Because the student is not blind to the participants' group assignment and is invested in her intervention's success, there is a high risk that she will subtly and unintentionally score the behavioral observations of the intervention group more favorably. A is a participant-based effect, whereas the question asks about the assessment procedure conducted by the researcher. B is a threat, but it's not related to the assessment procedure itself. D is a threat to the initial group composition, which should have been controlled by random assignment, and is not a feature of the post-test procedure.

Question 8

A researcher is studying the efficacy of a new group therapy protocol for social anxiety. Ten therapy groups are run, each with eight participants. To analyze the data, the researcher performs an independent samples t-test comparing the post-treatment anxiety scores of the 80 participants in the therapy condition to 80 participants in a no-treatment control group.

What is the most critical methodological flaw in the researcher's statistical analysis plan?

  1. The assumption of homogeneity of variance is likely to be violated between a treatment and no-treatment group.
  2. The data from participants within the same therapy group are not independent, violating a core assumption of the t-test. (correct answer)
  3. The sample size is insufficient for a t-test and a more powerful test like an ANOVA should have been used.
  4. A pre-test/post-test design would be necessary to draw any valid conclusions about the therapy's effectiveness.
Explanation: The correct answer is B. The t-test assumes that all observations are independent of one another. In this study, participants within the same therapy group share a common experience (the same therapist, group dynamics, etc.), and their outcomes are likely to be more similar to each other than to participants in other groups. This 'nesting' of participants within groups violates the assumption of independence. The consequence is a deflated standard error, which greatly increases the risk of a Type I error (a false positive). A is a possible but less certain violation, and t-tests can be robust to it. C is incorrect, as N=160 is a large sample for a t-test. D is incorrect as a post-test only design with random assignment is a valid experimental design, though the statistical analysis is flawed.

Question 9

A study investigated the effectiveness of a new therapy for insomnia. Results showed that the therapy significantly reduced sleep latency (time to fall asleep). A subsequent analysis revealed that this effect was much stronger for participants who reported high levels of motivation to change at baseline, while the therapy had almost no effect for participants with low motivation.

In this study, what is the role of 'motivation to change'?

  1. It is a confounding variable that should have been controlled for during random assignment.
  2. It is a mediating variable that explains the process through which the therapy works.
  3. It is a moderating variable that influences the strength of the relationship between the therapy and sleep latency. (correct answer)
  4. It is a dependent variable that was unintentionally influenced by the therapeutic intervention.
Explanation: The correct answer is C. A moderator is a variable that affects the direction or strength of the relationship between an independent variable (the therapy) and a dependent variable (sleep latency). In this case, motivation to change determines when or for whom the therapy is effective. The therapy's effect is conditional upon the level of motivation. B is incorrect because a mediator would explain how the therapy works (e.g., therapy -> reduces cognitive arousal -> reduces sleep latency). A is incorrect because while motivation might be a confounding variable in a non-randomized study, here it is identified as an effect modifier. D is incorrect because sleep latency is the dependent variable.

Question 10

A psychologist is reviewing literature to select an evidence-based treatment for obsessive-compulsive disorder. The psychologist finds several case studies, a large correlational study linking symptom severity to coping skills, a pre-post study of a new therapy, and a large multi-site, double-blind randomized controlled trial (RCT) comparing a specific form of CBT to a placebo.

When considering the level of evidence for causality, why is the RCT considered superior to the other study designs listed?

  1. The double-blind procedure eliminates all potential threats to internal and external validity, ensuring the results are definitive.
  2. The multi-site nature of the trial ensures that the findings are generalizable to all clinical settings and patient populations.
  3. The large sample size increases statistical power to a level where any observed effect can be considered clinically significant.
  4. The random assignment of participants to conditions minimizes selection bias and helps to control for unknown confounding variables. (correct answer)
Explanation: The correct answer is D. The defining feature of an RCT and its primary strength for inferring causality is random assignment. This process helps ensure that the groups are equivalent on all variables (both known and unknown) before the intervention begins. Therefore, any post-treatment differences between groups can be more confidently attributed to the intervention itself, rather than to pre-existing differences. A is an overstatement; double-blinding controls for expectancy effects but not for other threats like attrition. B addresses external validity, but the core strength of the RCT for causality is internal validity derived from randomization. C is incorrect, as statistical significance does not automatically equal clinical significance, and this is a feature of sample size, not the core design for causality.

Question 11

To evaluate a new after-school tutoring program, a researcher compares the academic growth of students in a school that adopted the program to students in a neighboring school that did not. This is a non-equivalent control group design. At the end of the year, the students in the tutoring program show significantly greater gains. However, the school with the program serves a community with a higher average socioeconomic status (SES) than the control school.

Which of the following threats to internal validity is the most likely explanation for the observed difference in academic gains?

  1. History, as a unique local event may have affected one school but not the other during the study period.
  2. Selection-maturation interaction, where the groups differed at baseline and had different natural development rates. (correct answer)
  3. Instrumentation, as it is possible that teachers at the two schools used slightly different grading standards.
  4. Testing, because the pre-test may have prompted students in the treatment school to seek more academic support.
Explanation: The correct answer is B. A selection-maturation interaction is a specific threat in non-equivalent control group designs. It occurs when the two groups are different to begin with (selection bias, e.g., due to SES) and these initial differences lead them to develop or change at different rates over time (maturation), even without the intervention. Students from higher-SES backgrounds might have more resources and support at home, leading to faster academic growth regardless of the tutoring program. A is plausible, but the pre-existing SES difference points more specifically to a selection-based threat. C and D are possible but are less likely to account for a systematic group difference than the known, major difference in group composition.

Question 12

Researchers conduct a pilot study (N=20) on a new intervention for social anxiety. They compare their intervention group (n=10) to a control group (n=10). The results are not statistically significant (p = .21), but the mean anxiety score in the intervention group is lower than in the control group, and the calculated effect size is medium (Cohen's d = 0.55).

Based on a critical appraisal of these pilot results, what is the most appropriate scientific conclusion?

  1. The intervention has no effect on social anxiety, and further research on this specific approach is not warranted.
  2. The result can be interpreted as a 'trend' and provides preliminary evidence for the intervention's effectiveness.
  3. A Type II error is unlikely given the p-value, meaning the null hypothesis is probably true.
  4. The finding of a medium effect size despite non-significance suggests the study was likely underpowered. (correct answer)
Explanation: The correct answer is C. Statistical power is the ability to detect a true effect. A small sample size (N=20) results in low power. Finding a medium effect size (d=0.55 is considered medium) that is not statistically significant is a classic sign of an underpowered study. It suggests a potentially meaningful effect exists, but the study lacked the precision to detect it reliably. A is a premature conclusion. B uses the term 'trend' in a way that is statistically inappropriate; the result is non-significant. D is incorrect; a non-significant result in an underpowered study means a Type II error (failing to reject a false null hypothesis) is a very real possibility.

Question 13

To study team dynamics, a consulting psychologist receives permission to place video cameras in a company's conference rooms. The employees are informed that their meetings will be recorded for a study on corporate efficiency. The psychologist's analysis of the footage reveals that the teams are exceptionally cooperative, proactive, and stay on-task, with rates of positive communication that are much higher than industry benchmarks.

When interpreting these findings, the psychologist should be most concerned about which methodological issue?

  1. Observer bias, where the psychologist's hopes for a positive outcome influenced their coding of the video footage.
  2. Reactivity, where the employees' awareness of being observed and recorded altered their natural behavior. (correct answer)
  3. Selection bias, as the company that agreed to participate may have better-functioning teams than the average company.
  4. Low inter-rater reliability, suggesting that different coders would not agree on the communication ratings.
Explanation: The correct answer is B. Reactivity, also known as the Hawthorne effect, occurs when research participants modify their behavior because they are aware of being studied. In this scenario, the employees knew they were being recorded for an efficiency study, which likely motivated them to be on their best behavior, leading to the unusually positive results. A is possible, but reactivity affects the participants' actual behavior, which is a more fundamental issue than the researcher's interpretation. C is a threat to external validity, but reactivity is a threat to the internal validity of the conclusion that these behaviors are typical for this company. D is a measurement concern but not the most likely explanation for the unusually positive findings themselves.

Question 14

A study reports that a new web-based cognitive-behavioral therapy (CBT) program is as effective as traditional face-to-face CBT for treating panic disorder. In the methodology section, the authors note that participants were required to have access to a personal computer with high-speed internet and were excluded if they scored below a certain threshold on a measure of computer literacy.

When considering the generalizability of these findings, what is the most critical underlying assumption the study makes?

  1. Technological access and literacy are not significant barriers or moderators of treatment success in the general patient population. (correct answer)
  2. Participants in both treatment arms possessed a similar level of motivation to complete the therapy.
  3. The therapeutic alliance is not a critical component for effective treatment of panic disorder.
  4. The therapists providing the face-to-face CBT were all equally skilled and adhered strictly to the treatment manual.
Explanation: The correct answer is C. The study's exclusion criteria systematically removed individuals who lack technological resources or skills. By doing so, the study can only demonstrate efficacy within a technologically proficient population. Generalizing the finding (i.e., that web-based CBT is as effective as in-person CBT) to the broader population of individuals with panic disorder rests on the critical and likely incorrect assumption that these excluded factors (access and literacy) don't matter in the real world. A is an interesting theoretical question but not a direct assumption related to the methodology. B is a standard assumption controlled by random assignment. D relates to treatment fidelity and internal validity, but the core generalizability issue stems from the specific exclusion criteria.

Question 15

In a large-scale survey conducted in 2023, a researcher finds that individuals aged 65-75 report significantly higher levels of life satisfaction than individuals aged 25-35. The researcher concludes that as people age, their life satisfaction tends to increase.

The researcher's conclusion about aging is most challenged by which limitation of the study's methodology?

  1. The self-report measure of life satisfaction may not be a valid indicator of true happiness or well-being.
  2. The observed difference may be a cohort effect, reflecting different life experiences of the generations, rather than an aging effect. (correct answer)
  3. The sample may not be representative of the general population, limiting the external validity of the satisfaction levels.
  4. There could be a non-response bias, where happier people at all ages are more likely to complete surveys.
Explanation: The correct answer is B. This study uses a cross-sectional design, which compares different age groups at a single point in time. The major weakness of this design for drawing conclusions about development or aging is that it confounds age with cohort effects. The older group not only is older but also grew up in a different historical, social, and economic context than the younger group. The observed difference in life satisfaction could be due to these differing life experiences (a cohort effect) rather than the process of aging itself. A, C, and D are all potential limitations of survey research, but they do not represent the primary confound that specifically challenges the conclusion about aging.

Question 16

Researchers develop a novel, intensive, 12-session psychotherapy for social anxiety disorder. They test its efficacy in a randomized controlled trial (RCT) with a waitlist control group. Participants are recruited via flyers posted on a university campus and must be full-time students to be included. The results show a large and statistically significant reduction in social anxiety symptoms for the treatment group compared to the control group.

A psychologist working in a community mental health center is considering adopting this therapy. What is the primary limitation they should consider when applying these research findings to their client population?

  1. The lack of an active control group makes it difficult to determine if the therapy is superior to existing treatments for social anxiety.
  2. The use of a waitlist control group may have introduced expectancy effects that inflated the perceived effectiveness of the novel therapy.
  3. The reliance on a university student convenience sample limits the generalizability of the findings to a more diverse, non-student population. (correct answer)
  4. The internal validity may be compromised because participants were not blinded to their treatment condition, which is a flaw in the RCT design.
Explanation: The correct answer is C. The study's sample consists of university students, who may differ systematically from the general population served by a community mental health center in terms of age, socioeconomic status, education level, and severity or chronicity of symptoms. This sampling method threatens the external validity, or generalizability, of the findings. A is a valid critique of the study's design in terms of comparative effectiveness, but the primary limitation in applying it to a new population is generalizability. B and D both point to potential internal validity threats, but random assignment controls for many of these, and the most critical issue for the community psychologist is whether the results will apply to their specific clients (external validity).

Question 17

A pharmaceutical company conducts a study on a new medication for anxiety. The researchers measure 25 different outcomes, including various physiological markers, self-report scales, and behavioral assessments. They conduct 25 separate t-tests comparing the medication group to a placebo group. The results show no significant differences on 24 of the outcomes, but a significant difference (p = .04) is found for one subscale of a self-report measure of worry.

What is the most pressing concern a psychologist should have when critically appraising this single significant finding?

  1. The use of multiple comparisons without statistical correction has inflated the probability of a Type I error. (correct answer)
  2. The study may have been underpowered to detect effects on the other 24 outcome measures.
  3. The construct validity of the single 'worry' subscale may be poor, as it is only one component of anxiety.
  4. A ceiling effect may have occurred, preventing the detection of real effects on the other outcome measures.
Explanation: The correct answer is C. When multiple statistical tests are conducted, the overall probability of finding at least one 'significant' result just by chance (a Type I error) increases. With 25 independent tests using an alpha of .05, the probability of at least one false positive is very high (1 - .9525.95^25 โ‰ˆ 72%). This practice is often called 'fishing' or 'p-hacking.' Without a statistical correction (like a Bonferroni correction), the single significant finding is highly suspect and may be spurious. A, B, and D are all possible methodological issues, but the multiple comparisons problem is the most direct and severe statistical flaw that challenges the validity of the reported significant finding.

Question 18

A psychologist conducts a 12-month longitudinal study of a behavioral intervention for weight loss. The study begins with 200 participants. At the 12-month follow-up, only 120 participants remain. An analysis of the remaining participants shows a significant average weight loss. However, an examination of the attrition data reveals that the 80 participants who dropped out had, on average, a higher starting BMI and had lost less weight at the 3-month check-in compared to those who completed the study.

How does this pattern of attrition most likely impact the study's conclusion about the intervention's effectiveness?

  1. It strengthens the conclusion, as the final sample represents participants who were most engaged with the intervention.
  2. It weakens external validity but does not affect the internal validity of the findings for the completers.
  3. It likely inflates the apparent effectiveness of the intervention by systematically removing less successful participants. (correct answer)
  4. It reduces the statistical power of the study, making the significant finding even more robust and meaningful.
Explanation: The correct answer is C. This is a case of differential attrition, which is a major threat to internal validity. The participants who dropped out were systematically different from those who remained, specifically in ways related to the outcome (higher BMI, less initial success). By losing the participants who were faring poorly, the final sample is biased towards those who were more successful. This makes the intervention appear more effective than it actually was for the entire group that started it. A is incorrect because this is a biased interpretation that ignores the methodological flaw. B is incorrect because differential attrition is a classic threat to internal validity. D is incorrect because while attrition does reduce power, the key issue here is the systematic bias introduced, which compromises the validity of the conclusion, rather than making it more robust.

Question 19

A rigorously designed, large-scale (N=2000) randomized controlled trial is conducted to test the efficacy of a popular over-the-counter herbal supplement for improving mood. The study is double-blind and placebo-controlled, with low attrition. The results show no statistically significant difference between the supplement and placebo groups on the primary mood outcome (p = .82). In their marketing, the supplement manufacturer states, 'This study is inconclusive. The absence of evidence is not evidence of absence.'

How should a scientifically-oriented psychologist evaluate the manufacturer's claim in the context of this specific study?

  1. The manufacturer is correct; since the null hypothesis cannot be proven, the study offers no useful information about the supplement's efficacy.
  2. The claim is technically true but misleading; a well-powered study that fails to find an effect provides strong evidence for the absence of a meaningful effect. (correct answer)
  3. The claim is invalid because a p-value greater than .50 is strong statistical evidence in favor of the null hypothesis being true.
  4. The manufacturer is likely correct, as smaller, preliminary studies had previously shown positive effects for the supplement.
Explanation: The correct answer is B. The phrase 'absence of evidence is not evidence of absence' is typically invoked when a study is underpowered and a null result is found. In that case, the study may have simply missed a real effect. However, this study is described as large-scale and rigorously designed, implying it had high statistical power to detect even a small effect. In a high-powered study, a null result is much more informative. It provides strong evidence that if an effect exists at all, it is too small to be of clinical or practical significance. A is incorrect because the study is very useful. C misinterprets p-values; frequentist statistics do not provide evidence for the null hypothesis. D is incorrect as small, preliminary studies are more prone to bias and Type I errors than large, rigorous RCTs.

Question 20

A large clinical trial (N=1000) compares a new cognitive therapy (NCT) to treatment-as-usual (TAU) for depression. The results indicate a statistically significant difference between the groups in favor of NCT (p = .04). However, the calculated effect size for this difference is very small (Cohen's d = 0.12).

When appraising this research for clinical practice, what is the most appropriate interpretation of these findings?

  1. The large sample size ensured that even a small, clinically trivial difference between the groups would be statistically significant. (correct answer)
  2. The p-value near the .05 threshold suggests the finding is marginal and likely a Type I error that should be disregarded.
  3. The statistical significance indicates that NCT is a meaningfully superior treatment that should be adopted over TAU.
  4. The study was likely underpowered, and the small effect size does not accurately reflect the true potential of the NCT.
Explanation: The correct answer is A. This question requires distinguishing between statistical significance and clinical (or practical) significance. A p-value indicates the probability of observing the data (or more extreme data) if the null hypothesis were true. With a very large sample size, even a tiny, clinically meaningless effect can become statistically significant. The small effect size (d=0.12 is considered 'very small') suggests that while the difference is unlikely due to chance, it is probably not large enough to be meaningful in a clinical setting. B misinterprets p-values; the p-value is the probability of the data given the null, not the probability of an error. C incorrectly equates statistical significance with clinical importance. D is incorrect; a statistically significant result means the study had sufficient power to detect an effect of this magnitude. If anything, the study was very well-powered.