All questions
Question 1
A researcher is studying an intervention and wants to know if its effectiveness differs for clients with and without a comorbid substance use disorder. In this research question, the presence of a comorbid substance use disorder serves as a potential:
- mediator.
- dependent variable.
- confounding variable.
- moderator. (correct answer)
Explanation: A moderator is a variable that influences the strength or direction of the relationship between an independent variable (the intervention) and a dependent variable (the outcome). The question asks if the treatment's effectiveness differs for certain subgroups. This is a classic moderation question: the relationship between intervention and outcome is hypothesized to be different depending on the level of the third variable (presence/absence of a substance use disorder).
A, A mediator explains how the intervention works.
B, The dependent variable is the outcome being measured (e.g., symptom reduction).
C, A confounding variable would be an extraneous factor related to both the intervention and outcome that could explain the relationship.
Question 2
A research study on a new intervention for panic disorder reports a statistically significant result (p < .01) but a small effect size (Cohen's d = 0.25). How should a clinician best interpret this finding for their practice?
- The intervention is highly effective and should be prioritized due to the high level of statistical significance.
- The intervention is not effective because a small effect size indicates a trivial or meaningless therapeutic change.
- The intervention produces a reliable but likely small average change, which may not be clinically meaningful for all patients. (correct answer)
- The statistical significance is more important than the effect size, suggesting the results are robust and generalizable.
Explanation: This question requires differentiating between statistical and clinical significance. A significant p-value (p < .01) indicates that the observed effect is unlikely due to chance, but it does not describe the magnitude of the effect. The effect size (Cohen's d = 0.25) provides that information. A Cohen's d of 0.25 is conventionally considered a small effect. Therefore, the most accurate interpretation is that the intervention produces a real (statistically reliable) effect, but its magnitude is small on average, which may or may not translate to a meaningful clinical improvement for a given individual.
A is incorrect because it conflates statistical significance with a large effect.
B is incorrect because a small effect size is not necessarily trivial; it may still be important, especially at a population level or for specific individuals, but it is not a large effect.
D is incorrect because in clinical practice, the magnitude of the effect (effect size) is often more important for decision-making than the p-value, which is heavily influenced by sample size.
Question 3
A clinic launches a new anger management program for individuals court-ordered for treatment. All participants are assessed at intake and have extremely high scores on an anger expression inventory. After the 8-week program, the mean score for the group is significantly lower. The program has no control group.
Besides the program's effectiveness, which of the following is the most significant threat to the internal validity of the conclusion that the program caused the improvement?
- Placebo effect.
- History.
- Regression to the mean. (correct answer)
- Instrumentation change.
Explanation: Regression to the mean is a statistical phenomenon where extreme scores on a first measurement tend to be closer to the average on a second measurement. Because the participants were selected based on their extremely high anger scores, it is statistically probable that their scores would decrease upon re-testing, even in the absence of any intervention. This makes it a major threat to internal validity in this single-group pre-post design with an extreme group.
A, Placebo effect, is a possible threat, but regression is more salient given the selection of an extreme group.
B, History, refers to external events that could affect the outcome, which is possible but not as directly implied by the study design as regression.
D, Instrumentation change, would involve a change in the measurement tool or procedure, which is not mentioned in the scenario.
Question 4
A psychologist is evaluating a client who has met the criterion for reliable change on a measure of PTSD symptoms, but their post-treatment score remains well above the established clinical cutoff. What is the most accurate interpretation of this outcome?
- The client has recovered and is now in the functional range.
- The client's improvement is statistically meaningful but they still meet criteria for a clinical diagnosis. (correct answer)
- The change was likely due to measurement error and is not clinically important.
- The client has not achieved clinically significant change and the treatment was ineffective.
Explanation: This scenario tests the two components of Jacobson and Truax's model of clinically significant change. The client has met the first criterion (reliable change), meaning their improvement is statistically real and not just measurement error. However, they have failed to meet the second criterion (crossing the clinical cutoff), meaning their score is still in the dysfunctional range. Therefore, the most accurate interpretation is that while a meaningful change occurred, the client remains clinically symptomatic.
A is incorrect because the client is still above the clinical cutoff.
C is incorrect because meeting the reliable change criterion means the change is unlikely to be due to measurement error.
D is an overstatement; while full clinically significant change was not met, the treatment did produce a statistically reliable improvement, so it was not entirely ineffective.
Question 5
A study found that a new therapy for depression is effective. Further analysis revealed that the therapy's success is largely explained by its ability to increase clients' self-efficacy. That is, the therapy boosts self-efficacy, which in turn reduces depressive symptoms. In this relationship, self-efficacy is best described as a:
- moderator.
- mediator. (correct answer)
- confounder.
- covariate.
Explanation: A mediator is a variable that explains the mechanism or process by which an independent variable (the therapy) influences a dependent variable (depressive symptoms). The causal chain described is Therapy -> Self-Efficacy -> Symptom Reduction. Self-efficacy is the intervening variable that explains how the therapy works.
A, A moderator is a variable that affects the strength or direction of the relationship between two other variables (e.g., the therapy works better for women than men).
C, A confounder is an extraneous variable that is related to both the independent and dependent variables, potentially causing a spurious association.
D, A covariate is a variable that is potentially related to the outcome and is statistically controlled for in an analysis.
Question 6
A community clinic wants to evaluate its new treatment protocol for anxiety. To do so, they compare their clients' outcomes (e.g., effect size, recovery rates) to the outcomes reported in the large-scale randomized controlled trials (RCTs) that originally established the treatment's efficacy. This evaluation strategy is known as:
- formative evaluation.
- meta-analysis.
- needs assessment.
- benchmarking. (correct answer)
Explanation: Benchmarking is a quality improvement process that involves comparing one's own processes and outcomes to an external standard of excellence. In clinical psychology, this often means comparing the results of a treatment delivered in a real-world setting to the results achieved in highly controlled efficacy trials (RCTs). The goal is to see if the clinic's performance is 'on par' with the gold standard.
A, Formative evaluation is conducted during program development to make improvements.
C, Needs assessment is done before a program begins to determine the need for it.
D, Meta-analysis is a statistical technique for combining the results of multiple studies; it is what might be used to establish the benchmark, but it is not the act of comparing one's own results to it.
Question 7
Two empirically supported treatments for OCD (Treatment A and Treatment B) have been shown in meta-analyses to produce equivalent average outcomes. Treatment A requires 20 sessions with a highly specialized therapist and is very costly. Treatment B requires 10 sessions with a master's-level therapist and is less than half the cost. To help a healthcare system decide which treatment to offer, the most relevant type of program evaluation would be:
- a needs assessment.
- a summative evaluation.
- a formative evaluation.
- a cost-effectiveness analysis. (correct answer)
Explanation: When two interventions produce similar outcomes but differ in their resource requirements, a cost-effectiveness analysis is the appropriate evaluation method. This type of analysis compares the costs and outcomes of different interventions to determine which provides the most 'value for money'. Since outcomes are equivalent, the less costly option would be more cost-effective.
A, A needs assessment determines if a program is needed in the first place.
B, A summative evaluation assesses the overall effectiveness of a program, but the scenario states both are already known to be effective.
D, A formative evaluation is used to refine a program while it is being developed or implemented.
Question 8
A psychologist is designing a treatment plan for a client with complex trauma and highly specific, personal goals that are not well-captured by standardized anxiety or depression measures. To track progress, the psychologist and client collaboratively define five goals, each with a 5-point scale from 'much less than expected outcome' to 'much more than expected outcome'. This measurement strategy is best identified as:
- Benchmarking.
- Goal Attainment Scaling. (correct answer)
- Clinical Significance Cutoff.
- Nomothetic Assessment.
Explanation: Goal Attainment Scaling (GAS) is an idiographic (individualized) method for quantifying progress toward specific, personal goals. It involves defining a range of possible outcomes for each goal, typically on a -2 to +2 scale ('much less than expected' to 'much more than expected'), and then evaluating the client's status at a later time. This perfectly matches the description.
A, Benchmarking, involves comparing a group's outcomes to an external standard, such as outcomes from a clinical trial.
C, Clinical Significance Cutoff, is a feature of nomothetic measures used to determine if a score has moved from a clinical to a non-clinical range.
D, Nomothetic Assessment, refers to the use of standardized instruments with normative data to compare an individual to a larger group, which is the opposite of the approach described.
Question 9
A psychologist is evaluating the outcome of therapy for a client with social anxiety. The client's pre-treatment score on the Social Phobia Inventory (SPIN) was 55. After 12 weeks of treatment, the post-treatment score is 40. The SPIN has a reliability of .90 and a standard deviation of 10. Using the Reliable Change Index (RCI), what is the most accurate conclusion about the client's change?
- The change is not statistically reliable because the score difference does not exceed the standard error of measurement.
- The change is statistically reliable, as the 15-point decrease is significant when accounting for measurement error. (correct answer)
- The clinical significance of the change cannot be determined without comparing the score to a normative population mean.
- The change is likely due to regression to the mean, as the initial score was very high.
Explanation: The Reliable Change Index (RCI) determines if the observed change in a score is larger than what would be expected due to measurement error alone. The formula for the standard error of the difference (S_diff) is SD * sqrt(2) * sqrt(1-r_xx). Here, S_diff = 10 * sqrt(2) * sqrt(1-.90) ≈ 10 * 1.414 * 0.316 ≈ 4.47. The RCI is the change score divided by S_diff: (55-40)/4.47 = 15/4.47 ≈ 3.35. An RCI value greater than 1.96 is considered statistically reliable at the p < .05 level. Therefore, the change is statistically reliable.
A is incorrect because the change (15 points) is substantially larger than the standard error of the difference (≈4.47).
C is incorrect because this statement describes clinical significance, whereas the question asks about reliable change. Reliable change is about statistical reliability, not comparison to norms.
D is incorrect because while regression to the mean is a potential factor, the RCI calculation is specifically designed to account for measurement error and determine if the change exceeds this effect.
Question 10
During a course of cognitive-behavioral therapy for depression, a client's BDI-II scores were 34, 32, 33, and then 15 in consecutive weekly sessions. The score of 15 and similar low scores were maintained for the remainder of treatment.
This pattern of change is best characterized as a:
- measurement artifact due to poor test-retest reliability.
- sudden gain, which is often associated with positive treatment outcomes. (correct answer)
- flight into health, which is typically a transient and defensive maneuver.
- gradual treatment response that is typical of cognitive-behavioral therapy.
Explanation: A sudden gain is a large, abrupt, and stable improvement in symptoms that occurs between two consecutive therapy sessions. The pattern described (a large drop from 33 to 15 that is maintained) fits this definition perfectly. Research on sudden gains indicates they occur in a substantial minority of clients and are generally predictive of good long-term outcomes.
A is unlikely, as the BDI-II is a reliable measure, and the change was stable.
C, A 'flight into health' is a psychoanalytic concept describing a pseudo-improvement to avoid deeper therapeutic work; it is typically not stable, unlike a sudden gain.
D is incorrect because the change was abrupt, not gradual.
Question 11
A routine outcome monitoring system in a university counseling center flags a client whose scores on a distress measure have reliably increased over three consecutive sessions. What is the psychologist's most appropriate initial action?
- Refer the client to a more experienced therapist immediately.
- Discontinue the measure as it is likely causing distress for the client.
- Discuss the trend collaboratively with the client to understand its meaning and review the treatment plan. (correct answer)
- Wait for several more data points to confirm the deteriorating trend before taking any action.
Explanation: Routine outcome monitoring is intended to provide timely feedback to clinicians to improve care. When data indicates a client is deteriorating (getting worse), the first and most critical step is to use this information collaboratively. The psychologist should share the data with the client, solicit their perspective on the trend, and use the conversation as a basis for reviewing and potentially modifying the treatment approach.
A is a premature step that should only be considered after discussing the issue with the client and supervisor.
B is inappropriate; it 'shoots the messenger' and ignores important clinical information.
D abdicates the primary purpose of session-by-session monitoring, which is to allow for timely intervention when treatment is off-track.
Question 12
When defining 'successful outcomes' for a program aimed at reducing recidivism among formerly incarcerated individuals, a program evaluator insists on including the perspectives of community members, parole officers, and the clients themselves, in addition to statistical measures of re-arrest. This approach best reflects the importance of:
- stakeholder perspectives. (correct answer)
- statistical power.
- internal validity.
- longitudinal design.
Explanation: A 'stakeholder' is any person or group with a vested interest in a program or its evaluation. Modern program evaluation emphasizes that a successful outcome can be defined differently by different stakeholders (e.g., clients, staff, funders, community members). Including these diverse perspectives leads to a more comprehensive and meaningful definition of program success than relying on a single metric.
A, Internal validity concerns the causal inference of the program's effect.
B, Statistical power relates to the ability to detect an effect if one exists.
D, Longitudinal design refers to collecting data over time. While these are all important concepts, the emphasis on including multiple viewpoints directly relates to incorporating stakeholder perspectives.
Question 13
A clinic director reports that the average client improvement on a 0-100 scale was 10 points. However, an analysis of individual trajectories reveals that approximately 50% of clients improved by 30 points, while 50% of clients worsened by 10 points. This discrepancy highlights the primary limitation of:
- using non-standardized outcome measures.
- relying on mean group differences to represent individual outcomes. (correct answer)
- failing to include a no-treatment control group.
- evaluating change over an insufficient period of time.
Explanation: This scenario illustrates a key problem in outcome research: the group average can mask significant and meaningful individual differences. The average improvement of 10 points (0.5×30 + 0.5×(-10) = 15 - 5 = 10) suggests a generally positive outcome. However, the reality is that the treatment is having strong but divergent effects: it helps half the clients significantly, but it makes the other half worse. Relying solely on the mean obscures this critical information.
A, C, and D are all potential methodological issues, but the core problem demonstrated by the data provided is the misleading nature of the group average.
Question 14
According to the two-part model of clinically significant change proposed by Jacobson and Truax, a client's outcome is considered clinically significant only when:
- the magnitude of change is statistically reliable and the client's post-treatment score has crossed a cutoff into a functional distribution. (correct answer)
- the client reports a subjective sense of improvement and the therapist observes a corresponding behavioral change.
- the post-treatment score is at least two standard deviations below the mean of the dysfunctional population.
- the effect size of the individual's change is large (d > 0.80) when compared to a waitlist control group.
Explanation: The Jacobson and Truax (1991) model defines clinically significant change (CSC) as having two distinct components. First, the change must be statistically reliable, meaning it is greater than what can be attributed to measurement error (often assessed with the Reliable Change Index). Second, the client's post-treatment score must move from the 'dysfunctional' range to the 'functional' range. This is typically determined by the post-treatment score crossing a pre-determined cutoff point that places it closer to the mean of a healthy population than to the mean of the dysfunctional population.
B describes subjective and observational data, which are important but not the formal criteria of this specific model.
C describes one method for setting a cutoff point but does not include the first criterion of reliable change.
D describes a group-level effect size, whereas CSC is focused on individual-level change.
Question 15
A psychologist computes a change score for a client and finds that the 95% confidence interval for this change is [-2.5, 10.5]. What is the most accurate conclusion based on this confidence interval?
- The client has shown a statistically significant improvement.
- There is a 95% probability that the true change is between -2.5 and 10.5.
- The observed change is not statistically significant at the p < .05 level. (correct answer)
- The measurement tool is unreliable and should not be used for this client.
Explanation: A confidence interval (CI) for a difference score (or change score) indicates a range of plausible values for the true change. If the CI contains zero, it means that a change of zero is a plausible value. Therefore, we cannot reject the null hypothesis that no change has occurred. In statistical terms, if the 95% CI contains zero, the result is not statistically significant at the p < .05 level.
A is incorrect because the CI includes zero.
B is a common misinterpretation of confidence intervals; the CI refers to the procedure's likelihood of capturing the true parameter over many repetitions, not the probability of a single interval containing the true value.
D is an over-interpretation; while wide CIs can reflect measurement error, this single result is not sufficient to invalidate the tool.
Question 16
Which of the following describes the primary advantage of session-by-session outcome monitoring compared to a pre-test/post-test only design?
- It allows for timely clinical adjustments if a client is not progressing or is deteriorating. (correct answer)
- It provides higher statistical power for detecting overall group effects.
- It requires less time and fewer resources for data collection.
- It offers a more definitive conclusion about the long-term efficacy of the treatment.
Explanation: The key advantage of session-by-session monitoring (also known as measurement-based care) is the provision of rapid, ongoing feedback. This feedback loop allows the therapist and client to see if the treatment is on track, and if not, to make timely adjustments to the treatment plan. This is a significant advantage over pre-post designs, where a lack of progress is only identified after the treatment is complete.
A is incorrect; session-by-session monitoring requires more resources.
B is not its primary advantage; while more data points can help, the main purpose is clinical utility.
D is incorrect; long-term efficacy requires follow-up assessments, which are separate from the frequency of measurement during treatment.
Question 17
To establish that a client's post-treatment score on a depression scale has moved into a 'functional' or 'healthy' range, which piece of information is most essential?
- Normative data from a non-distressed or general population sample. (correct answer)
- The client's pre-treatment score on the same scale.
- The reliability coefficient of the depression scale.
- The standard deviation of change scores in a treated sample.
Explanation: Determining whether a score falls within a 'functional' range requires a comparison point or standard. This is typically derived from normative data from a healthy, non-distressed population. Cutoff scores (like those used in the Jacobson & Truax model) are often set based on the mean and standard deviation of such a normative sample (e.g., a score less than two standard deviations above the mean of the healthy population).
A and B are necessary for calculating if the change was reliable, but not for determining if the endpoint is in the functional range.
D is related to the variability of improvement but doesn't define what constitutes a healthy endpoint.
Question 18
The standard error of measurement (SEM) is a critical component in the calculation of the Reliable Change Index (RCI). What is the primary function of the SEM in this context?
- It establishes the cutoff score for determining clinical significance in a population.
- It adjusts the client's score to account for demographic differences from the normative sample.
- It is used to calculate the effect size of the intervention for an individual client.
- It provides an estimate of the amount of change expected due to random measurement error. (correct answer)
Explanation: The Standard Error of Measurement (SEM) is an index of how much an individual's score is expected to vary on repeated testing due to the unreliability of the test. In the context of the RCI, the SEM is used to calculate the standard error of the difference between two scores. This value represents the expected amount of change due to random error alone. The RCI then compares the client's actual change to this value to see if the change was larger than what chance would produce.
A is incorrect; clinical cutoffs are based on normative data, not the SEM.
C is incorrect; effect size is typically calculated using standard deviations, not the SEM.
D is incorrect; the SEM relates to reliability, not normative adjustments.
Question 19
When selecting an outcome measure to evaluate the effectiveness of a brief, focused intervention for insomnia, which psychometric property is most crucial for detecting change over a short period?
- High internal consistency (alpha > .90).
- Strong convergent validity with personality traits.
- High sensitivity to change. (correct answer)
- High test-retest reliability over a one-year interval.
Explanation: Sensitivity to change refers to a measure's ability to detect changes in a construct when they have actually occurred. For an outcome measure used to track progress in a brief intervention, this is the most critical property. A measure could be reliable and valid for diagnosis but not sensitive enough to pick up on the incremental improvements that happen during treatment.
A, High internal consistency is desirable, but extremely high alpha can sometimes indicate redundancy and may not be as important as sensitivity to change.
B, Convergent validity with stable traits is important for construct validity but less relevant for measuring state changes.
D, High test-retest reliability over a long interval is desirable for measuring stable constructs (like traits), but for a state-like construct like insomnia symptoms, you would expect scores to change with treatment, so lower long-term stability is expected.