All questions
Question 1
A physical therapist suspects a patient has a scaphoid fracture, estimating the pre-test probability to be approximately 20%. The therapist performs the anatomical snuff box tenderness test, which is positive. According to a systematic review, this test has a positive likelihood ratio (LR+) of 3.2. Using this information, what is the patient's approximate post-test probability of having a scaphoid fracture?
- 20%
- 46% (correct answer)
- 64%
- 82%
Explanation: To calculate post-test probability from pre-test probability and a likelihood ratio, one must first convert probability to odds, multiply by the LR, and then convert odds back to probability. Step 1: Pre-test odds = probability / (1 - probability) = 0.20 / (1 - 0.20) = 0.20 / 0.80 = 0.25. Step 2: Post-test odds = Pre-test odds * LR+ = 0.25 * 3.2 = 0.80. Step 3: Post-test probability = Post-test odds / (1 + Post-test odds) = 0.80 / (1 + 0.80) = 0.80 / 1.80 ≈ 0.444, or 44%. The closest answer is 46%.
Question 2
A therapist reviews a meta-analysis on the effectiveness of a specific exercise for chronic low back pain, measured by the Oswestry Disability Index (ODI). The pooled results show a mean difference of -4.5 points (95% CI [-8.0, -1.0]). The test for heterogeneity yields an I² statistic of 82% (p < 0.001). The accepted minimal clinically important difference (MCID) for the ODI is 10 points. How should the therapist BEST interpret these findings for their clinical practice?
- The exercise is highly effective and should be used for all patients, as the results are statistically significant.
- The exercise shows a statistically significant benefit, but its clinical importance is questionable and results varied widely across studies. (correct answer)
- The findings are invalid due to the high heterogeneity, and therefore no clinical conclusion can be drawn from the review.
- The exercise is not effective because the entire 95% confidence interval is below the minimal clinically important difference.
Explanation: The 95% CI of [-8.0, -1.0] does not cross zero, indicating a statistically significant result favoring the intervention. However, the entire range of the CI, including the point estimate of -4.5, is below the MCID of 10 points, suggesting the effect may not be clinically important. The high I² value (82%) indicates substantial heterogeneity, meaning the results of the individual studies were very inconsistent. Therefore, the most accurate interpretation is that while a statistical effect exists, its clinical relevance is uncertain, and the high variability suggests the average effect may not apply to every patient.
Question 3
A 50-year-old patient involved in a motor vehicle accident presents to physical therapy with neck pain. The patient was the driver in a simple rear-end collision, is ambulatory, has no midline cervical spine tenderness, and reports delayed onset of pain. However, the patient is unable to actively rotate their neck 40 degrees to the right due to stiffness. According to the Canadian C-Spine Rule, what is the MOST appropriate immediate action?
- Initiate gentle range of motion exercises as the patient meets all low-risk criteria.
- Refer the patient for radiographic imaging of the cervical spine. (correct answer)
- Clear the cervical spine from fracture risk and proceed with a full musculoskeletal evaluation.
- Apply moist heat to the neck and re-attempt active rotation assessment in 15 minutes.
Explanation: The Canadian C-Spine Rule is a multi-step algorithm. The patient has no high-risk factors. The patient meets several low-risk factors (simple rear-end MVA, ambulatory, delayed onset, no midline tenderness), which allows the clinician to proceed to the final step: active range of motion assessment. The rule requires the patient to be able to actively rotate the neck 45 degrees to the left and right. Since the patient cannot rotate more than 40 degrees, they fail this step of the algorithm. A failure at any step necessitates radiographic imaging.
Question 4
In a randomized controlled trial for a fall prevention program, 20 out of 100 participants in the control group experienced a fall over one year, while 8 out of 100 participants in the intervention group experienced a fall. What is the number needed to treat (NNT) to prevent one additional person from falling?
- 5
- 8 (correct answer)
- 12
- 60
Explanation: The Number Needed to Treat (NNT) is the reciprocal of the Absolute Risk Reduction (ARR). First, calculate the event rates: Control Event Rate (CER) = 20/100 = 0.20. Experimental Event Rate (EER) = 8/100 = 0.08. Next, calculate the ARR: ARR = CER - EER = 0.20 - 0.08 = 0.12. Finally, calculate the NNT: NNT = 1 / ARR = 1 / 0.12 ≈ 8.33. The closest whole number is 8, meaning approximately 8 people need to be treated with the intervention to prevent one additional fall.
Question 5
A study aims to assess the inter-rater reliability of the 5-point Manual Muscle Test (MMT) scale (0-5) for the quadriceps muscle between two physical therapists. The therapists will grade the strength of 50 patients. Which statistical method is MOST appropriate for analyzing the reliability of these measurements?
- Kappa coefficient
- Intraclass Correlation Coefficient (ICC) (correct answer)
- Pearson product-moment correlation
- Percent agreement
Explanation: The Intraclass Correlation Coefficient (ICC) is the most appropriate statistic for assessing reliability when data is ordinal (like the MMT scale), interval, or ratio, and involves two or more raters. The Kappa coefficient is used for nominal (categorical) data. Pearson correlation assesses the strength of a linear relationship, not agreement. Percent agreement is a less robust measure as it does not account for agreement that could occur by chance. The MMT scale is ordinal data, for which the ICC is the preferred method.
Question 6
A study investigating a new exercise program for shoulder pain randomly assigns participants to an intervention or control group. The outcome assessor is unaware of group assignments when measuring shoulder range of motion. However, the treating physical therapists and the patients are aware of their group assignments. The lack of blinding for the therapists and patients creates a risk for which type of bias?
- Selection bias
- Detection bias
- Attrition bias
- Performance bias (correct answer)
Explanation: Performance bias occurs when there are systematic differences in the care provided to the groups apart from the intervention being studied. When therapists and patients know the group assignment, they may behave differently. For example, therapists might unintentionally provide more encouragement to the intervention group, or patients in the control group might seek other treatments (co-intervention). Blinding of the outcome assessor mitigates detection bias, and randomization mitigates selection bias.
Question 7
A physical therapist uses the Phalen test to screen for carpal tunnel syndrome. A high-quality diagnostic accuracy study reports that the Phalen test has a sensitivity of 75% and a specificity of 95%. The test result is positive for the patient. What is the MOST appropriate clinical interpretation of this finding?
- The positive result strongly increases the probability that the patient has carpal tunnel syndrome. (correct answer)
- The test is not very useful because its sensitivity for detecting the condition is only moderate.
- A negative result on this test would have been more effective at ruling out carpal tunnel syndrome.
- The positive result confirms the diagnosis of carpal tunnel syndrome beyond any doubt.
Explanation: Specificity refers to a test's ability to correctly identify those without the condition (true negative rate). A test with high specificity is good for 'ruling in' a condition when the test is positive (SpPin mnemonic). With a specificity of 95%, there is a low rate of false positives, so a positive result significantly increases the post-test probability of the condition. While sensitivity is only moderate, which limits the value of a negative test, the high specificity makes a positive test clinically very useful.
Question 8
A study is published with the conclusion that a new bracing technique significantly reduces pain in patients with patellofemoral pain syndrome (p = 0.04). However, a later, much larger, and more rigorously conducted study finds that the brace has no effect. Assuming the larger study is correct, what type of error was made in the conclusion of the original, smaller study?
- Type I error (correct answer)
- Type II error
- Random measurement error
- Systematic review error
Explanation: A Type I error (alpha error or false positive) occurs when a researcher incorrectly rejects a true null hypothesis. In this case, the original study concluded there was an effect (rejected the null hypothesis of 'no effect') when, in reality, no effect existed. This is a false positive finding. A Type II error would be the opposite: concluding there is no effect when one truly exists.
Question 9
A physical therapist is working in an acute care setting and wants to know the likely clinical course and long-term functional outcome for a patient who has just sustained a severe traumatic brain injury. What type of clinical question is the therapist asking?
- Intervention
- Diagnosis
- Etiology
- Prognosis (correct answer)
Explanation: Prognostic questions are about predicting a patient's likely future. This therapist is asking about the expected clinical course and outcomes, which is a question of prognosis. An intervention question would ask about the effectiveness of a treatment. A diagnosis question would ask about the accuracy of a test to identify a condition. An etiology question would ask about the cause or origin of a disease.
Question 10
A retrospective case-control study investigated the link between playing overhead sports in adolescence and developing shoulder impingement later in life. The study reports an odds ratio of 4.2 (95% CI [2.1, 8.4]). Which statement is the MOST accurate interpretation of this finding?
- Playing overhead sports in adolescence increases the risk of shoulder impingement by 4.2 times.
- The finding is not statistically significant because it is from a retrospective study design.
- For every 4.2 people who play overhead sports, one will develop shoulder impingement.
- The odds of having played overhead sports were 4.2 times higher for individuals with impingement than for controls. (correct answer)
Explanation: When you encounter odds ratios in research questions, you need to understand what they actually measure and how the study design affects their interpretation. Case-control studies work backwards—they start with people who have the outcome (cases) and compare them to people without it (controls), then look back to see who had the exposure.
The correct interpretation is D because this odds ratio tells us the odds of having the exposure (playing overhead sports) were 4.2 times higher in the impingement group compared to the control group. The 95% confidence interval [2.1, 8.4] doesn't include 1.0, confirming statistical significance.
A is wrong because it confuses odds ratio with relative risk. While the language sounds similar, odds ratios from case-control studies don't directly give you risk increases—that requires knowing the actual incidence rates, which case-control studies don't provide.
B incorrectly suggests retrospective design affects statistical significance. The confidence interval, not the study design, determines significance. Since the CI doesn't include 1.0, this finding is statistically significant.
C completely misinterprets what 4.2 represents. This isn't a rate or proportion—it's a ratio comparing odds between two groups.
NPTE Strategy: When you see odds ratios from case-control studies, remember they compare the odds of exposure between cases and controls, not the risk of developing disease. Always check if the confidence interval includes 1.0 to assess significance, regardless of study design.
Question 11
A physical therapist is reviewing a meta-analysis forest plot where the outcome is pain on a 10-point scale. The summary diamond, representing the overall pooled effect of an intervention compared to a control, is located entirely to the left of the vertical line of no effect (mean difference = 0). What does this graphical representation indicate?
- The overall effect of the intervention is not statistically significant.
- The intervention is less effective than the control for reducing pain.
- The overall effect of the intervention is statistically significant and favors the intervention. (correct answer)
- There is significant heterogeneity among the included studies, making the result unreliable.
Explanation: In a forest plot, the vertical line represents the line of no effect (e.g., a mean difference of 0). The summary diamond represents the pooled effect from all studies. If the diamond does not touch or cross the vertical line, the result is statistically significant. For a pain scale, a lower score is better. An effect to the left of 0 (a negative mean difference) indicates that the intervention group had lower scores (less pain) than the control group. Therefore, the plot shows a statistically significant effect that favors the intervention.
Question 12
A physical therapist is treating a 55-year-old patient recovering from a rotator cuff repair. The patient's score on the Shoulder Pain and Disability Index (SPADI) improves from 60 (severe disability) to 51 (moderate disability) over a three-week period. A high-quality study established that the minimal detectable change with 95% confidence (MDC95) for the SPADI in this population is 12.1 points, and the minimal clinically important difference (MCID) is 14.5 points.
- The improvement is neither statistically meaningful nor clinically important. (correct answer)
- The improvement is statistically meaningful but not yet clinically important.
- The improvement is clinically important but not statistically meaningful.
- The improvement is both statistically meaningful and clinically important.
Explanation: The patient's score changed by 9 points (60 - 51). For the change to be considered statistically meaningful (i.e., a true change beyond measurement error), it must exceed the MDC95. Since 9 is less than 12.1, the change is not statistically meaningful. For the change to be clinically important, it must exceed the MCID. Since 9 is less than 14.5, the change is not clinically important. Therefore, the observed improvement is within the range of measurement error and has not reached the threshold for a meaningful clinical benefit.
Question 13
A researcher wishes to explore the perspectives, beliefs, and experiences of individuals with spinal cord injury regarding community reintegration. The plan is to use semi-structured interviews and focus groups, then analyze the textual data for dominant themes and patterns. This study design is BEST described as:
- Quantitative
- A randomized controlled trial
- A case-control study
- Qualitative (correct answer)
Explanation: Research methodology questions on the NPTE often test your ability to distinguish between different study designs based on their key characteristics. When you encounter a question describing a research plan, focus on the data collection methods and analysis approach to identify the study type.
This study design is clearly qualitative research (D). The key indicators are the use of semi-structured interviews and focus groups to gather data, along with the plan to analyze "textual data for dominant themes and patterns." Qualitative research explores subjective experiences, beliefs, and perspectives through open-ended data collection methods, then identifies recurring themes in non-numerical data. The researcher's goal of understanding personal experiences with community reintegration perfectly aligns with qualitative methodology.
Option A (Quantitative) is incorrect because quantitative research uses numerical data and statistical analysis to test hypotheses or measure relationships between variables. This study isn't collecting numerical measurements or using statistical tests.
Option B (A randomized controlled trial) is wrong because RCTs involve randomly assigning participants to different intervention groups to test cause-and-effect relationships. This study has no intervention or control groups.
Option C (A case-control study) is incorrect because case-control studies compare groups with and without a specific condition to identify risk factors. This research doesn't involve comparing different groups based on outcomes.
Study tip: Remember that qualitative studies use words (interviews, focus groups, observations) while quantitative studies use numbers (surveys with numerical scales, measurements, statistical analysis). The mention of "themes and patterns" from textual data is a dead giveaway for qualitative research.
Question 14
A research team is developing a new clinical test to screen for risk of falls in community-dwelling older adults. The team's primary goal is to ensure the test can accurately identify individuals who will go on to have a fall in the next six months. Which type of measurement validity is MOST critical for the team to establish for this specific purpose?
- Content validity
- Concurrent validity
- Predictive validity (correct answer)
- Construct validity
Explanation: Predictive validity, a subtype of criterion validity, assesses how well a test's results predict a future outcome. In this scenario, the goal is to use the new test to predict future falls. Therefore, establishing that the test scores are strongly associated with the occurrence of falls over the next six months (a future event) is the most critical form of validity. Content validity (test content relevance), concurrent validity (correlation with a gold standard at the same time), and construct validity (measuring the intended abstract concept) are all important, but predictive validity directly addresses the stated primary goal.
Question 15
A physical therapist is evaluating the evidence for using a specific manual therapy technique. Which of the following sources would provide the highest level of evidence regarding the effectiveness of this intervention?
- A case report published by a leading expert in manual therapy describing a successful outcome.
- A large, well-designed prospective cohort study tracking outcomes of patients receiving the technique.
- A systematic review of high-quality randomized controlled trials (RCTs) on the technique. (correct answer)
- A single, multi-center randomized controlled trial with a low risk of bias.
Explanation: According to the hierarchy of evidence for intervention questions, systematic reviews and meta-analyses of high-quality RCTs are considered the highest level of evidence. They synthesize the results from multiple studies, providing a more robust and generalizable estimate of the intervention's effect than any single study. A single RCT is the next highest level, followed by cohort studies, with case reports being a much lower level of evidence.
Question 16
A patient with acute low back pain meets 4 of the 5 criteria of the clinical prediction rule for the success of lumbar thrust manipulation. The clinical practice guideline from the APTA gives a strong recommendation for using thrust manipulation in patients who are positive on this rule. The patient has no contraindications. What is the therapist's BEST course of action?
- Avoid manipulation because the patient does not meet all five criteria of the rule.
- Search for additional primary research before deciding to use the manipulation.
- Incorporate lumbar thrust manipulation into the plan of care as part of a multimodal approach. (correct answer)
- Inform the patient that manipulation is the only treatment that will be effective for their condition.
Explanation: Clinical practice guidelines synthesize the best available evidence. A strong recommendation indicates that the benefits clearly outweigh the risks. The patient is positive on the clinical prediction rule (often defined as meeting a certain number of criteria, e.g., ≥4/5), has no contraindications, and the guideline strongly supports the intervention for this subgroup. Therefore, the most evidence-based action is to incorporate the recommended treatment. This should be done within a broader, multimodal plan of care.
Question 17
An extremely large randomized controlled trial (n=5,000) compares a new stretching technique to a standard technique for improving hamstring length. The study finds that the new technique results in 1.5 degrees more improvement in active knee extension range of motion than the standard technique. This difference was found to be statistically significant (p = 0.02), with a 95% confidence interval of [0.3, 2.7] degrees. The minimal clinically important difference (MCID) for this measure is considered to be 5 degrees. What is the MOST valid conclusion?
- The new technique is clinically superior to the standard technique due to the statistically significant finding.
- The study's findings are likely invalid because the p-value is very close to the alpha level of 0.05.
- The new technique provides a statistically significant benefit, but the magnitude of the effect is not clinically meaningful. (correct answer)
- The results are inconclusive because the effect size is small, suggesting the study was underpowered.
Explanation: The result is statistically significant because the p-value (0.02) is less than 0.05 and the 95% CI does not cross zero. However, the point estimate of the effect (1.5 degrees) and the entire range of the confidence interval (0.3 to 2.7 degrees) are well below the established MCID of 5 degrees. This means that while a true difference between the groups likely exists, it is too small to be considered meaningful to the patient. The large sample size allowed the study to have high power to detect a very small effect.
Question 18
A physical therapist reads a study with a 95% confidence interval reported for the mean improvement in a functional outcome score as [10.5, 12.5]. Which of the following is the MOST accurate interpretation of this confidence interval?
- 95% of the patients in the study improved their scores by an amount between 10.5 and 12.5 points.
- There is a 95% probability that the true population mean improvement is between 10.5 and 12.5 points. (correct answer)
- If the study were repeated many times, the sample mean would fall between 10.5 and 12.5 in 95% of the trials.
- The sample mean for improvement was 11.5, and the margin of error was 1.0 point.
Explanation: A 95% confidence interval for a mean provides a range of plausible values for the true population mean. The correct interpretation is that one can be 95% confident that this interval contains the true mean of the population from which the sample was drawn. It does not describe the range of individual scores (A) or the probability of future sample means (C). While D is mathematically correct in this specific symmetric case (mean is the midpoint, margin of error is half the width), B provides the correct conceptual interpretation of what a confidence interval represents.
Question 19
A research team conducts a pilot study (n=30) to test a novel rehabilitation protocol. The results show a positive trend favoring the new protocol, but the difference between groups is not statistically significant (p = 0.18). The researchers calculate that their study had only 40% power to detect the minimum clinically important difference. What is the MOST likely reason for the non-significant finding?
- The intervention has no true effect on the outcome.
- The study had insufficient statistical power. (correct answer)
- The alpha level was set too stringently at 0.05.
- The study suffered from significant performance bias.
Explanation: Statistical power is the probability of detecting a true effect if one exists. A power of 40% means that even if the intervention was truly effective, the study had only a 40% chance of producing a statistically significant result. This is considered very low (typically 80% is the minimum standard). Given the small sample size of a pilot study and the explicitly stated low power, this is the most likely reason for failing to find a statistically significant effect, representing a potential Type II error.