National Physical Therapy Examination (NPTE) Quiz: Research Interpretation
20 questions · exam conditions
0:00
Research InterpretationQuestion 1 of 20

Improvement=6 points; MDC=5, MCID=8. Best interpretation?

Reliable and below MCID
Clinical and below MDC
Not reliable; not clinical
Both reliable and clinical
← Back to quizzes

National Physical Therapy Examination (NPTE) Quiz

National Physical Therapy Examination (NPTE) Quiz: Research Interpretation

Practice Research Interpretation in National Physical Therapy Examination (NPTE) with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Research Interpretation, giving you a quick way to practice the rules, question types, and explanations that matter most for National Physical Therapy Examination (NPTE).

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Improvement=6 points; MDC=5, MCID=8. Best interpretation?

  1. Reliable and below MCID (correct answer)
  2. Clinical and below MDC
  3. Not reliable; not clinical
  4. Both reliable and clinical
Explanation: Your 6-point improvement passes the MDC of 5, so the change is reliable and not just measurement error. It falls below the MCID of 8, so it isn't clinically meaningful. The tempting mistake is choosing 'both reliable and clinical,' because reliability alone doesn't make the change clinically important.

Question 2

OR=2.1 (95% CI 0.9-4.8). Best conclusion?

  1. CI indicates a moderate effect
  2. Not statistically significant (correct answer)
  3. Treatment effect equals zero
  4. Wide CI shows high precision
Explanation: The 95% confidence interval from 0.9 to 4.8 includes 1, the null value for an odds ratio, so the result is not statistically significant. A wide interval that crosses 1 means there is insufficient evidence for an effect; the point estimate alone doesn't establish one. The tempting error is calling it a moderate effect: an OR of 2.1 looks moderate, but the interval includes no effect, so you can't claim one.

Question 3

A test has sensitivity 85%, specificity 70%, prevalence 20%. PPV?

  1. 85%
  2. 41% (correct answer)
  3. 20%
  4. 70%
Explanation: With 20% prevalence, 20 of 100 people have disease. Sensitivity 85% detects 17 true positives. Specificity 70% means 30% false positives among the 80 healthy people, so 24 false positives. PPV = 17 / (17 + 24) = 17/41 = 41%. The tempting wrong answer is 85% sensitivity, because it ignores prevalence and false positives.

Question 4

New balance test: r=0.85 with balance questionnaire, r=0.30 with depression. Supports?

  1. Construct validity (correct answer)
  2. Test-retest reliability
  3. Criterion validity
  4. Internal consistency
Explanation: A new test correlating highly with an established balance measure (convergent) and weakly with depression (discriminant) shows scores behave according to the theoretical construct. That pattern is construct validity. Criterion validity is tempting because of the r=0.85, but criterion validity requires comparison with a gold-standard outcome, not a pattern of convergent and discriminant correlations.

Question 5

What measurement property was highlighted when the tool demonstrated a ceiling effect in 28% of participants?

  1. Limited ability to detect improvement in higher-functioning patients because many started near maximum scores. (correct answer)
  2. Enhanced sensitivity to change because many participants achieved the highest possible score.
  3. High internal consistency, meaning items measured unrelated constructs across multiple domains.
  4. Strong criterion validity, meaning the measure predicted future falls with high accuracy.
Explanation: This question tests the ability to interpret research findings, measurement properties, and statistical results relevant to physical therapy practice. Understanding research involves analyzing statistical outcomes, measurement properties like reliability and validity, and their clinical implications. In this specific study, the tool showed a ceiling effect in 28% of participants, highlighting limited detection of improvement in high-functioners. Choice A is correct because it accurately reflects the study's findings on measurement properties, confirming sensitivity issues. Choice B is incorrect due to inverting ceiling effect benefits, a common mistake in scale limitations. To enhance interpretation skills: Encourage critical analysis of outcomes, consider the role of confounding variables, and emphasize understanding of measurement properties through practical examples.

Question 6

Which finding had the most significant clinical implication when effect size was moderate but CI was wide due to small sample?

  1. The intervention might be beneficial, but imprecision suggests clinicians should interpret magnitude cautiously. (correct answer)
  2. A wide CI proved the intervention was definitively superior for all patient subgroups.
  3. Moderate effect size guaranteed long-term benefit regardless of adherence or follow-up duration.
  4. Small samples eliminate random error, so the estimate was more precise than large trials.
Explanation: This question tests the ability to interpret research findings, measurement properties, and statistical results relevant to physical therapy practice. Understanding research involves analyzing statistical outcomes, measurement properties like reliability and validity, and their clinical implications. In this specific study, the moderate effect size with wide CI due to small sample highlighted cautious interpretation of benefits. Choice A is correct because it accurately reflects the study's findings on clinical implications, confirming imprecision in estimates. Choice B is incorrect due to overstating definitiveness, a common mistake with wide intervals. To enhance interpretation skills: Encourage critical analysis of outcomes, consider the role of confounding variables, and emphasize understanding of measurement properties through practical examples.

Question 7

A study evaluating a new balance training program for older adults with a history of falls reports a statistically significant improvement in the Berg Balance Scale (BBS) scores for the intervention group (p = 0.02). The study used a large sample of 500 participants. The mean improvement in the intervention group was 2.1 points. Previous research has established the Minimal Detectable Change at 90% confidence (MDC₉₀) for the BBS as 4 points and the Minimal Clinically Important Difference (MCID) as 5 points.

Based on these results, what is the MOST accurate conclusion about the balance training program?

  1. The program is highly effective because the large sample size produced a statistically significant result.
  2. The program's effectiveness is questionable because the mean improvement did not exceed the threshold for real or clinically meaningful change. (correct answer)
  3. The program produced a true change in balance because the improvement was found to be statistically significant.
  4. The results are invalid because the mean improvement of 2.1 points is less than the established MDC₉₀ and MCID.
Explanation: The correct interpretation requires integrating statistical significance with measures of clinical relevance. While the p-value is significant (p=0.02), likely due to the large sample size, the mean improvement of 2.1 points is less than both the MDC₉₀ (4 points) and the MCID (5 points). This means the observed change is not large enough to be confident it's a 'real' change beyond measurement error, nor is it large enough to be considered clinically meaningful to the patient. Therefore, the program's effectiveness is questionable.

Question 8

A study aims to create a clinical prediction rule (CPR) to identify patients with low back pain likely to respond to a spinal manipulation technique. After analyzing data from a derivation cohort of 100 patients, the researchers report that the CPR has a sensitivity of 0.92 and a specificity of 0.55 for predicting treatment success. The positive likelihood ratio (+LR) is 2.04 and the negative likelihood ratio (-LR) is 0.15.

A physical therapist is considering using this CPR with a new patient who has a pre-test probability of success estimated at 40%. Which feature of the CPR is MOST useful for this therapist in clinical practice?

  1. The high sensitivity, which is useful for ruling in patients who are likely to succeed with the intervention.
  2. The positive likelihood ratio of 2.04, which provides a large shift in post-test probability for a positive finding.
  3. The negative likelihood ratio of 0.15, which significantly decreases the probability of success if the patient is negative on the rule. (correct answer)
  4. The low specificity, which indicates the rule is effective at confirming patients who will respond well.
Explanation: The clinical utility of a CPR is determined by its likelihood ratios. An -LR between 0.1 and 0.2 generates large and often conclusive shifts in post-test probability. An -LR of 0.15 is considered moderately to highly useful for ruling out a condition (or in this case, the likelihood of success). The +LR of 2.04 is small and only slightly increases the post-test probability. High sensitivity is useful for ruling OUT, not ruling in (SnNout). Low specificity is poor for confirming a condition (SpPin). Therefore, the most powerful feature of this CPR is its ability to use a negative finding to confidently rule out the patient as a good candidate for the manipulation.

Question 9

A study is designed to compare the effectiveness of three different physical therapy interventions (A, B, and C) for shoulder impingement syndrome. The primary outcome is shoulder function, measured on a continuous scale using the SPADI score. The researchers want to know if there is a difference among the three groups after an 8-week intervention period. The data for the SPADI scores are found to be severely skewed and do not meet the assumptions for parametric testing.

Which statistical test is MOST appropriate to analyze the difference between the three intervention groups?

  1. Independent t-test
  2. Analysis of Variance (ANOVA)
  3. Mann-Whitney U test
  4. Kruskal-Wallis test (correct answer)
Explanation: The Kruskal-Wallis test is the non-parametric equivalent of a one-way ANOVA. It is used to determine if there are statistically significant differences between two or more independent groups on a continuous or ordinal dependent variable when the assumptions of ANOVA (e.g., normality of data) are not met. An independent t-test (A) is used for only two groups. An ANOVA (B) is inappropriate due to the skewed, non-normal data. A Mann-Whitney U test (C) is the non-parametric equivalent of an independent t-test and is also used for only two groups.

Question 10

A new patient-reported outcome measure for kinesiophobia is administered to a group of patients with chronic low back pain. The scale has a maximum possible score of 68. In a validation study, the mean score for the sample was 62 with a standard deviation of 4. The researchers report that 45% of the participants achieved the maximum score. The intervention study that follows fails to show a significant improvement in kinesiophobia for a new, promising intervention.

What measurement property MOST likely explains the intervention study's failure to detect an effect?

  1. Low test-retest reliability
  2. A significant ceiling effect (correct answer)
  3. A significant floor effect
  4. Poor concurrent validity
Explanation: A ceiling effect occurs when a high proportion of subjects in a study have scores at or near the maximum possible score on a measure. In this case, with 45% of participants scoring the maximum of 68, the scale cannot detect any further improvement in these individuals because they have already 'maxed out' the scale. This makes it very difficult to demonstrate a treatment effect, as any real improvement in this large subgroup of patients cannot be measured, thus masking the intervention's potential benefit.

Question 11

A physical therapist is evaluating a new special test for diagnosing subacromial pain syndrome. A study reports the test's performance using a Receiver Operating Characteristic (ROC) curve analysis. The Area Under the Curve (AUC) is reported as 0.78 (95% CI [0.69, 0.87]).

How should the therapist interpret the diagnostic utility of this test based on the AUC value?

  1. The test has excellent accuracy in distinguishing between patients with and without the syndrome.
  2. The test has fair accuracy, performing better than chance but not at a high level for definitive diagnosis. (correct answer)
  3. The test is no better than a coin flip for diagnosing the syndrome because the confidence interval is wide.
  4. The test has poor accuracy because the AUC value is less than the ideal value of 1.0.
Explanation: The Area Under the Curve (AUC) from an ROC analysis represents the overall accuracy of a diagnostic test. An AUC of 1.0 is a perfect test, and an AUC of 0.5 is equivalent to chance. General guidelines for interpreting AUC are: 0.90-1.00 = excellent; 0.80-0.90 = good; 0.70-0.80 = fair; 0.60-0.70 = poor; <0.60 = fail. An AUC of 0.78 falls into the 'fair' category. It indicates that the test has some ability to discriminate but is not accurate enough to be used as a standalone diagnostic tool.

Question 12

A study is conducted at a specialized sports medicine clinic to evaluate the effectiveness of a new post-operative ACL reconstruction protocol. The study uses a pre-test/post-test design with a single group of 50 consecutive patients. During the 6-month study period, a major local professional sports team has a highly publicized successful season, leading to increased community-wide interest and motivation in sports rehabilitation.

The study finds significant improvements in all outcome measures. Which threat to internal validity is MOST likely to confound these results?

  1. Selection bias
  2. Maturation
  3. History (correct answer)
  4. Instrumentation
Explanation: History refers to external events that occur during the course of a study that are not part of the intervention but could affect the dependent variable. The local sports team's success is a historical event that could have influenced the participants' motivation and engagement in rehab, independent of the new protocol. Selection bias (A) relates to non-random assignment. Maturation (B) refers to internal changes in subjects over time (e.g., natural healing), which is also a threat here, but the external event is a more distinct 'history' threat. Instrumentation (D) refers to changes in the measurement process.

Question 13

A study is published on the effects of a new exercise program for patients with chronic obstructive pulmonary disease (COPD). The inclusion criteria were: age 50-75, confirmed COPD diagnosis, and ability to walk independently. The exclusion criteria were: unstable cardiac disease, severe musculoskeletal comorbidities, and cognitive impairment. A physical therapist is treating a 78-year-old patient with COPD and stable angina who uses a rolling walker.

What is the MOST significant challenge when applying the results of this study to the therapist's patient?

  1. Internal validity
  2. Statistical conclusion validity
  3. Construct validity
  4. External validity (correct answer)
Explanation: External validity refers to the generalizability of study findings to other populations, settings, or times. The therapist's patient differs from the study sample in several ways that match the exclusion criteria: age (78 vs. 50-75), cardiac comorbidity (stable angina vs. no unstable cardiac disease), and use of an assistive device (walker vs. independent walking). Because the patient would have been excluded from the study, it is uncertain whether the study's results are applicable or generalizable to them. This is a direct challenge to the study's external validity.

Question 14

A randomized controlled trial compares a novel manual therapy technique plus exercise to exercise alone for chronic neck pain. The primary outcome is the Neck Disability Index (NDI), measured at baseline, 4 weeks, and 12 weeks. The researchers use a repeated measures ANOVA to analyze the data and report a significant time-by-group interaction effect (p = 0.01). They also report non-significant main effects for time (p = 0.08) and group (p = 0.15).

What is the MOST appropriate interpretation of these statistical findings?

  1. Neither the manual therapy nor the exercise was effective because the main effects for time and group were not significant.
  2. The rate of change in NDI scores over the 12-week period was significantly different between the two groups. (correct answer)
  3. Both groups improved significantly over time, but there was no overall difference between the manual therapy and exercise-only groups.
  4. The manual therapy group had significantly better NDI scores than the exercise-only group when averaged across all time points.
Explanation: A significant time-by-group interaction is the most important finding in a repeated measures ANOVA. It indicates that the effect of time (the change in the dependent variable, NDI) is different for the different levels of the grouping variable (manual therapy vs. exercise only). In other words, the two groups changed differently over time. When a significant interaction is present, the main effects (for time and group) are often not interpretable on their own and should be disregarded in favor of interpreting the interaction through post-hoc tests.

Question 15

A study evaluates the effectiveness of a workplace ergonomic intervention on the incidence of carpal tunnel syndrome (CTS) among office workers. The results are analyzed using logistic regression. The analysis yields an odds ratio of 0.60 (95% CI [0.45, 0.80]) for the intervention group compared to the control group, with the dependent variable being the diagnosis of CTS.

What is the correct interpretation of the odds ratio in this context?

  1. The odds of developing CTS in the intervention group are 60% of the odds in the control group. (correct answer)
  2. The intervention reduces the risk of developing CTS by 60%.
  3. The intervention is not effective because the odds ratio is greater than zero.
  4. For every 10 people receiving the intervention, 6 will be prevented from developing CTS.
Explanation: When you encounter odds ratios in research studies, remember that they represent the ratio of odds between two groups, not percentages or absolute risk reductions. An odds ratio of 0.60 means the odds of developing CTS in the intervention group are 0.60 times (or 60% of) the odds in the control group. Since this value is less than 1.0, it indicates the intervention is protective. The 95% confidence interval [0.45, 0.80] doesn't include 1.0, confirming statistical significance. Choice A correctly interprets this relationship - the intervention group has 60% of the odds compared to the control group. Choice B misinterprets the odds ratio as a risk reduction percentage. A 40% risk reduction would be calculated as (1 - 0.60) × 100% = 40%, not 60%. The intervention reduces odds by 40%, not 60%. Choice C demonstrates a fundamental misunderstanding. An odds ratio greater than zero doesn't indicate ineffectiveness - you need to compare it to 1.0. Values less than 1.0 (like 0.60) indicate protective effects, while values greater than 1.0 suggest increased risk. Choice D incorrectly treats the odds ratio as an absolute measure. Odds ratios are relative measures comparing two groups; they don't directly translate to "X out of 10 people prevented." This would require additional information like baseline risk rates. For NPTE success, remember that odds ratios compare relative odds between groups: less than 1.0 = protective effect, greater than 1.0 = increased risk, and exactly 1.0 = no difference. Never interpret them as percentages or absolute numbers.

Question 16

A validation study for a new 2-minute walk test (2MWT) in patients with multiple sclerosis (MS) compares its results to a standard 6-minute walk test (6MWT). The study reports a Pearson correlation coefficient of r = 0.85 between the two tests. It also reports that the 2MWT has excellent test-retest reliability (ICC = 0.95).

The high correlation between the 2MWT and the 6MWT provides evidence for which type of measurement validity?

  1. Predictive validity
  2. Content validity
  3. Construct validity
  4. Concurrent validity (correct answer)
Explanation: Concurrent validity, a subtype of criterion-related validity, is assessed by correlating a new measure with a well-established 'gold standard' measure administered at approximately the same time. In this case, the new 2MWT is being compared to the established 6MWT. A high correlation (r = 0.85) indicates that the 2MWT measures a similar construct to the 6MWT, providing strong evidence of its concurrent validity. Predictive validity would involve seeing if the 2MWT score predicts a future outcome. Content validity relates to how well the test's items represent the construct. Construct validity is a broader term that encompasses concurrent validity.

Question 17

Researchers conducted a case-control study to investigate the association between regular participation in high-impact sports during adolescence and the development of hip osteoarthritis (OA) later in life. They identified 200 individuals with diagnosed hip OA (cases) and 200 age- and sex-matched individuals without hip OA (controls). The study found an odds ratio of 3.5 (95% CI [1.8, 6.7]) for developing hip OA associated with a history of high-impact sports.

Based on this study's design and results, which conclusion is MOST appropriate?

  1. Participating in high-impact sports during adolescence increases the risk of developing hip OA by 3.5 times.
  2. The odds of having a history of high-impact sports participation are 3.5 times greater for individuals with hip OA compared to those without. (correct answer)
  3. High-impact sports participation during adolescence is the primary cause of hip OA in this population.
  4. For every 3.5 individuals who participated in high-impact sports, one will develop hip OA.
Explanation: An odds ratio (OR) from a case-control study represents the odds of exposure (high-impact sports) among cases (hip OA) compared to the odds of exposure among controls (no hip OA). It does not directly state the relative risk. Therefore, the correct interpretation is that individuals with hip OA were 3.5 times more likely to have a history of high-impact sports. Stating it as a relative risk (A) is technically incorrect for a case-control design, though OR can approximate RR if the disease is rare. Causation (C) cannot be inferred from an observational study. (D) is an incorrect interpretation of the statistic.

Question 18

A researcher conducts a pilot study with 20 participants to test a new intervention for improving gait speed in patients with Parkinson's disease. The results show a trend towards improvement, but the difference between the intervention and control groups is not statistically significant (p = 0.12). The observed effect size (Cohen's d) was 0.45. The researcher concludes that the intervention is ineffective.

Which of the following is the MOST likely reason for the non-significant finding?

  1. The intervention has no real effect on gait speed, as indicated by the p-value greater than 0.05.
  2. A Type I error occurred, leading to the incorrect conclusion that there was no effect.
  3. The study was likely underpowered, resulting in a Type II error despite a potentially meaningful effect size. (correct answer)
  4. The effect size of 0.45 is too small to be considered clinically relevant, so statistical significance was not achieved.
Explanation: A Type II error occurs when a study fails to detect a real effect that exists. This is often due to insufficient statistical power, which is highly dependent on sample size. A pilot study with only 20 participants is very likely to be underpowered. The effect size of 0.45 is considered small-to-medium and could be clinically relevant, but the small sample size prevented the study from reaching statistical significance. Therefore, the most probable explanation is a Type II error due to low power, and the researcher's conclusion of 'ineffective' may be premature.

Question 19

Researchers investigate the diagnostic accuracy of a new clinical test for identifying meniscal tears. They test a cohort of 150 patients with knee pain scheduled for arthroscopic surgery. The test is positive in 70 patients and negative in 80. Surgery confirms 75 meniscal tears. Among those with confirmed tears, the test was positive for 65. The pre-test probability of a meniscal tear in this population is 50%.

A patient from this population has a negative test result. What is the approximate probability that this patient still has a meniscal tear?

  1. 10%
  2. 13% (correct answer)
  3. 87%
  4. 25%
Explanation: This requires calculating the post-test probability after a negative result, which is the false negative rate (1 - Negative Predictive Value). First, create a 2x2 table. Total with disease = 75, Total without = 75. True Positives (TP) = 65. Therefore, False Negatives (FN) = 75 - 65 = 10. True Negatives (TN) = Total without disease - False Positives (FP). FP = Total Positive Tests - TP = 70 - 65 = 5. So, TN = 75 - 5 = 70. Total Negative Tests = FN + TN = 10 + 70 = 80. The probability of having a tear despite a negative test is FN / (FN + TN) = 10 / 80 = 0.125, or approximately 13%. This value is also known as (1 - NPV).

Question 20

A research team develops a new functional mobility scale for patients post-stroke. To establish inter-rater reliability, they recruit three experienced physical therapists from a large rehabilitation hospital to independently score 30 patients. The research protocol specifies that for clinical use, the average score from any two available therapists will be used. The researchers state that the therapists involved are representative of typical expert clinicians.

Given the study design and intended clinical application, which Intraclass Correlation Coefficient (ICC) model provides the MOST appropriate reliability estimate?

  1. ICC (2,k), because the raters are a random sample and the mean of k raters' scores is the unit of interest for clinical use. (correct answer)
  2. ICC (3,1), because a fixed group of raters was used and the reliability of a single rater's score is most important.
  3. ICC (2,1), because the raters are considered a random sample but clinical decisions will be based on a single therapist's rating.
  4. ICC (1,k), because each patient was rated by a different set of raters and the average score is used.
Explanation: The correct model is ICC (2,k). The '2' indicates a two-way random effects model, which is appropriate because the raters are considered a random sample of a larger population of similar therapists. The 'k' indicates that the reliability of the mean of 'k' raters (in this case, k=2) is being assessed, which aligns with the intended clinical use of averaging two therapists' scores. This model accounts for random error from both subjects and raters.