Psychology Quiz: Replication And Research Practices
20 questions · exam conditions
0:00
Replication And Research PracticesQuestion 1 of 20

A pharmaceutical company sponsors ten independent studies on its new anxiolytic drug. Two studies find a statistically significant benefit, while the other eight find no significant effect. The company only submits the two successful studies for publication.

It prevents other researchers from being able to replicate the successful studies accurately.
It invalidates the statistical significance found within the two published studies themselves.
It constitutes data fabrication, which is illegal in the context of clinical trials.
It leads to a biased overestimation of the drug's true efficacy in the published literature.
← Back to quizzes

Psychology Quiz

Psychology Quiz: Replication And Research Practices

Practice Replication And Research Practices in Psychology with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

What this quiz covers

This quiz focuses on Replication And Research Practices, giving you a quick way to practice the rules, question types, and explanations that matter most for Psychology.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

A pharmaceutical company sponsors ten independent studies on its new anxiolytic drug. Two studies find a statistically significant benefit, while the other eight find no significant effect. The company only submits the two successful studies for publication.

  1. It prevents other researchers from being able to replicate the successful studies accurately.
  2. It invalidates the statistical significance found within the two published studies themselves.
  3. It constitutes data fabrication, which is illegal in the context of clinical trials.
  4. It leads to a biased overestimation of the drug's true efficacy in the published literature. (correct answer)
Explanation: This scenario describes the "file drawer problem," a form of publication bias. When only studies with statistically significant results are published, the public record becomes skewed. A meta-analysis of the published literature would suggest the drug is effective, substantially overestimating its true effect size because the null findings from the other eight studies are hidden. This practice is a QRP, not data fabrication (C), and while it makes replication context difficult (A), the primary consequence is the distortion of the evidence base (D).

Question 2

A researcher analyzes a large public dataset without a specific hypothesis. She discovers that individuals who report eating breakfast daily also report higher levels of life satisfaction. Which of the following subsequent actions would be the clearest example of HARKing (Hypothesizing After the Results are Known)?

  1. Conducting a new, pre-registered study to experimentally test if providing breakfast increases life satisfaction.
  2. Reporting the finding transparently as the result of an exploratory analysis of a large dataset.
  3. Writing the manuscript's introduction as if the study was originally designed to test a theory about breakfast's role in well-being. (correct answer)
  4. Searching the existing literature to find other studies that have examined the relationship between diet and well-being.
Explanation: HARKing is the act of presenting a post-hoc finding as if it were an a priori prediction. Actions A, B, and D represent good scientific practice. (A) is the proper way to follow up on an exploratory finding. (B) is the transparent way to report the finding itself. (D) is standard practice for contextualizing a finding. Action (C) is the only one that involves misrepresenting the discovery process by creating a new hypothesis to fit the existing data and pretending it came first.

Question 3

A 1970s study found that feeling anonymous (by wearing hoods and robes) increased participants' willingness to administer fake electric shocks. A modern researcher conducts a study finding that participants who are part of an anonymous online forum are more likely to use aggressive language. If the modern study is successful, it is best described as a conceptual replication.

  1. It proves that the original 'hoods and robes' study was not a Type I error.
  2. It provides an exact estimate of the effect size of anonymity on aggression.
  3. It increases confidence in the generalizability of the theory linking anonymity to deindividuation and aggression. (correct answer)
  4. It shows that online aggression is psychologically identical to physical aggression.
Explanation: A conceptual replication tests the same underlying hypothesis as an original study but uses different methods, populations, or operationalizations. The primary value of a successful conceptual replication is that it enhances the generalizability of the underlying theory—in this case, that anonymity promotes aggressive or antisocial behavior. It shows the principle is robust across different contexts. It does not directly test whether the original study was a false positive (A), nor does it provide a comparable effect size due to the different methods (B).

Question 4

In a study on decision-making, a researcher predicts that participants will be more risk-averse after viewing a sad film clip. The initial analysis shows a non-significant result (p = .12). The researcher notices high variance in the data and decides to exclude the 15% of participants with the slowest reaction times, justifying this by arguing they were inattentive. After this exclusion, the result is significant (p = .04).

  1. The procedure is sound because inattentive participants add noise and reduce the power to detect a true effect.
  2. The procedure is questionable because the exclusion rule was implemented after seeing its impact on the p-value. (correct answer)
  3. The procedure is problematic because excluding 15% of the sample is an arbitrary and unjustifiably large proportion.
  4. The procedure is invalid because reaction time is not a legitimate basis for excluding participants in a decision-making study.
Explanation: While it can be legitimate to exclude data based on pre-defined criteria (e.g., 'we will exclude participants with RTs 3 SDs above the mean'), the problem here is that the decision was made post-hoc and was clearly motivated by the desire to achieve significance. This is a form of p-hacking. The researcher is capitalizing on chance by finding a justification to remove data points that are inconvenient for their hypothesis. The core issue is not the criterion itself (D) or the proportion (C), but the data-contingent way in which the criterion was applied.

Question 5

A pharmaceutical company sponsors ten independent studies on its new anxiolytic drug. Two studies find a statistically significant benefit, while the other eight find no significant effect. The company only submits the two successful studies for publication.

  1. It prevents other researchers from being able to replicate the successful studies accurately.
  2. It invalidates the statistical significance found within the two published studies themselves.
  3. It constitutes data fabrication, which is illegal in the context of clinical trials.
  4. It leads to a biased overestimation of the drug's true efficacy in the published literature. (correct answer)
Explanation: This scenario describes the "file drawer problem," a form of publication bias. When only studies with statistically significant results are published, the public record becomes skewed. A meta-analysis of the published literature would suggest the drug is effective, substantially overestimating its true effect size because the null findings from the other eight studies are hidden. This practice is a QRP, not data fabrication (C), and while it makes replication context difficult (A), the primary consequence is the distortion of the evidence base (D).

Question 6

Consider the following sequence of studies: Study 1, an initial experiment, finds a large effect of a 'growth mindset' intervention on academic performance. Study 2, a large, multi-site, pre-registered direct replication, finds no effect on academic performance. Study 3, a conceptual replication, finds the intervention does increase students' self-reported persistence on difficult tasks.

  1. The direct replication (Study 2) was likely flawed, as Study 3 confirmed the original finding from Study 1.
  2. The original effect on academic performance is likely not reliable, but the intervention might influence a related psychological process. (correct answer)
  3. The original finding (Study 1) and the conceptual replication (Study 3) are strong evidence that the direct replication (Study 2) is a Type II error.
  4. The entire line of research is contradictory and should be disregarded until a single, definitive study is conducted.
Explanation: This scenario requires synthesizing multiple, seemingly conflicting, results. The high-quality direct replication (Study 2) failing to find an effect on performance casts serious doubt on the reliability of Study 1's main finding. However, Study 3 suggests the intervention is not inert; it affects a related construct (persistence). The most nuanced conclusion is that the originally claimed outcome (performance) may not be replicable, but the intervention has a different, more subtle effect (on persistence). This integrates all three findings without dismissing any of them outright.

Question 7

Dr. Anya Sharma publishes a surprising finding that a specific cognitive training exercise improves creativity. Dr. Ben Carter's lab conducts a high-powered, pre-registered direct replication of Dr. Sharma's study, following the original methodology precisely. Dr. Carter's study finds no significant effect.

  1. Dr. Sharma must have committed scientific fraud in the original experiment.
  2. The theory that cognitive training can improve creativity is definitively proven false.
  3. Dr. Carter's replication must have been flawed, as it failed to find the original effect.
  4. The evidence for the original finding is substantially weakened, suggesting it could be a false positive or context-dependent. (correct answer)
Explanation: A single failed direct replication does not prove fraud (A) or definitively falsify a broad theory (B). Given that the replication was high-powered and pre-registered, it is unlikely to be simply flawed (C). The most cautious and appropriate scientific interpretation is that the evidence for the original effect is now much weaker. The original finding might have been a Type I error (a false positive), or the effect might depend on subtle contextual variables not captured in the methods section and therefore not reproduced in the replication attempt.

Question 8

A researcher conducts an experiment on a new attention-enhancing supplement. They measure performance on five different cognitive tasks: sustained attention, selective attention, task switching, working memory, and response inhibition. The supplement only shows a statistically significant improvement (p = .04) on the task-switching measure. The researcher then publishes a paper titled "A New Supplement Specifically Enhances Cognitive Flexibility."

  1. The research fails to establish the causal mechanism of the supplement's effect.
  2. The research demonstrates a clear case of scientific fraud by fabricating data.
  3. The research practice dramatically increases the probability of committing a Type I error. (correct answer)
  4. The five cognitive tasks selected lack sufficient construct validity for a robust conclusion.
Explanation: This scenario describes a form of p-hacking where a researcher tests multiple dependent variables but only reports the one that yields a significant result. By conducting multiple statistical tests (in this case, five), the overall probability of finding at least one significant result by chance alone (a Type I error) increases well above the nominal alpha level of .05. The primary distortion is statistical, not an issue of causal mechanism, fraud, or construct validity.

Question 9

A researcher conducts an exploratory study on personality and career choice, collecting data on hundreds of variables. They discover an unexpected, strong correlation between conscientiousness and preference for jobs with clear hierarchies. In the introduction to their manuscript, they present a detailed theory of how conscientiousness predisposes individuals to seek structured environments, framing their analysis as a direct test of this theory.

  1. It inappropriately presents an exploratory finding as if it were a confirmatory test of a prior hypothesis. (correct answer)
  2. It relies on correlational data, which is an invalid basis for the development of new psychological theories.
  3. It violates the principle of parsimony by creating an unnecessarily complex theory to explain a single finding.
  4. It selectively reports only one of the hundreds of correlations that were calculated during the exploratory phase.
Explanation: This is a classic example of HARKing (Hypothesizing After the Results are Known). The core problem is the misrepresentation of the research process. An exploratory finding, discovered without an a priori hypothesis, is presented as if it were the result of a confirmatory test of a pre-existing theory. This practice distorts the evidence and makes the finding seem much stronger than it is. While selective reporting (D) might also have occurred, HARKing specifically refers to the retrofitting of a hypothesis.

Question 10

If questionable research practices such as p-hacking and selective reporting of positive results are widespread in a field, what is the most direct consequence for the published literature of that field?

  1. The average statistical power of published studies will be unacceptably low.
  2. The rate of false positives will be substantially higher than the nominal alpha level. (correct answer)
  3. The external validity of findings will be severely compromised, but internal validity will be unaffected.
  4. The complexity of theories will increase as researchers explain away null findings.
Explanation: The nominal alpha level (usually .05) is the acceptable rate of Type I errors (false positives) under the assumption that the null hypothesis is true and all assumptions are met. QRPs like p-hacking and selective reporting systematically violate these assumptions, specifically by creating a system where non-significant results are either converted into significant ones or hidden from view. This directly inflates the proportion of published findings that are false positives, meaning the true Type I error rate in the literature is much higher than the stated 5%.

Question 11

An initial study on a new therapy reports a very large effect size (Cohen's d = 0.9). A subsequent, large-scale replication study confirms the therapy has a statistically significant effect, but finds a much smaller effect size (d = 0.25). Assuming both studies were conducted with high integrity, what is the most likely explanation for this discrepancy?

  1. The replication study must have used a less reliable outcome measure, causing the effect size to shrink.
  2. The initial, large effect size was likely an overestimation due to random sampling error combined with publication bias. (correct answer)
  3. The difference between d = 0.9 and d = 0.25 is negligible as long as both findings are statistically significant.
  4. The original study must have had much lower statistical power than the replication study.
Explanation: This pattern is often explained by the "winner's curse" and publication bias. Initial studies that get published are often those that, by chance, found a larger-than-average effect. Studies that by chance found a smaller (or null) effect are less likely to be published. Therefore, the first published effect size for a phenomenon is often an overestimate of the true effect. A large, high-powered replication provides a more precise estimate, which is typically smaller and closer to the true value. This is a systemic issue, not necessarily a flaw in either study.

Question 12

A researcher tests a new educational intervention with three different variations (A, B, and C) against a control group. The results show that variation A significantly outperforms the control group, but variations B and C do not. In the final manuscript, the researcher omits all mention of variations B and C and presents the study as a simple, successful comparison between intervention A and a control.

  1. This practice decreases the internal validity of the comparison between variation A and the control group.
  2. This practice prevents an accurate calculation of the effect size for the comparison between variation A and control.
  3. This practice is a form of data falsification and is considered scientific fraud.
  4. This practice misleads readers by creating a cleaner, more compelling narrative than the full data would support. (correct answer)
Explanation: This question tests your understanding of research ethics and publication practices, specifically focusing on selective reporting and its impact on scientific integrity. The researcher's decision to omit unsuccessful variations B and C creates a misleading narrative about the research process. By presenting only the successful comparison between variation A and control, the study appears as a straightforward success story rather than revealing the full picture: that two-thirds of the tested variations actually failed. This selective reporting gives readers an overly optimistic view of the intervention's effectiveness and suggests the research was more targeted and successful than it actually was. The correct answer is D because this practice fundamentally distorts the scientific narrative. Let's examine why the other options are incorrect. Choice A is wrong because internal validity refers to whether the study design properly controls for confounding variables—omitting failed variations doesn't affect the quality of the A versus control comparison itself. Choice B is incorrect because effect size calculations for the A versus control comparison remain mathematically valid regardless of whether other variations are mentioned. Choice C overstates the severity—while this practice is ethically problematic, it's selective reporting rather than data falsification, since no actual data was changed or fabricated. When you encounter research ethics questions, focus on how practices affect transparency and reader understanding. Selective reporting is a key issue in psychology research—it creates publication bias and gives an incomplete picture of what actually works, even when the reported results themselves are accurate.

Question 13

An initial study reports that a specific gene variant, GENE-A, is associated with superior memory (p = .02). A second, larger study attempts to replicate this and fails (p = .35). A subsequent massive genome-wide association study (GWAS) with 100,000 participants finds no link for GENE-A, but finds a robust association between a different variant, GENE-B, and memory.

  1. The effect of GENE-A is real but very small, requiring massive studies like the GWAS to detect reliably.
  2. The link between GENE-A and memory is likely moderated by GENE-B, explaining the inconsistent findings.
  3. The original finding was most likely a Type I error, and the cumulative evidence weighs against the claim. (correct answer)
  4. The original study and the replication study were both underpowered, so only the GWAS result is trustworthy.
Explanation: This question requires integrating three pieces of evidence. The initial finding was significant but was followed by a failed replication. The massive GWAS, which has immense statistical power, not only failed to find the GENE-A link but found a different one. The most parsimonious conclusion is that the original, small study produced a false positive (Type I error). The idea that the effect is real but tiny (A) is directly contradicted by the powerful GWAS. A moderation effect (B) is a post-hoc complication that is less likely than the simpler explanation of a false positive.

Question 14

A social psychologist finds a non-significant correlation between social media use and depression (r = .15, p = .20). The psychologist then re-runs the analysis, this time controlling for age and socioeconomic status. This second model yields a significant partial correlation (r = .22, p = .04). The final paper only reports the significant result from the second model, arguing it is more precise.

  1. The conclusion is strengthened because potential confounding variables have been appropriately controlled for.
  2. The reported p-value is likely misleading because the analytical strategy was chosen after seeing the data. (correct answer)
  3. The finding is likely a Type II error because the initial non-significant result was the more accurate one.
  4. The addition of covariates invalidates the use of a p-value for hypothesis testing in this context.
Explanation: This is an example of p-hacking through analytical flexibility. While controlling for covariates can be a valid approach, doing so only after an initial analysis fails to yield significance—and then reporting only the significant model—is a QRP. The choice of analysis was contingent on the outcome, which inflates the false positive rate. The reported p-value of .04 does not reflect the actual probability of a Type I error, which is higher due to the undisclosed analytical flexibility.

Question 15

Dr. Anya Sharma publishes a surprising finding that a specific cognitive training exercise improves creativity. Dr. Ben Carter's lab conducts a high-powered, pre-registered direct replication of Dr. Sharma's study, following the original methodology precisely. Dr. Carter's study finds no significant effect.

  1. Dr. Sharma must have committed scientific fraud in the original experiment.
  2. The theory that cognitive training can improve creativity is definitively proven false.
  3. Dr. Carter's replication must have been flawed, as it failed to find the original effect.
  4. The evidence for the original finding is substantially weakened, suggesting it could be a false positive or context-dependent. (correct answer)
Explanation: A single failed direct replication does not prove fraud (A) or definitively falsify a broad theory (B). Given that the replication was high-powered and pre-registered, it is unlikely to be simply flawed (C). The most cautious and appropriate scientific interpretation is that the evidence for the original effect is now much weaker. The original finding might have been a Type I error (a false positive), or the effect might depend on subtle contextual variables not captured in the methods section and therefore not reproduced in the replication attempt.

Question 16

A team of researchers publicly posts their study plan before collecting any data. The plan specifies their primary hypothesis, the target sample size with a power analysis, and the exact statistical models they will use to test the hypothesis. This practice is known as pre-registration.

  1. It primarily guards against the unintentional introduction of experimenter bias during data collection.
  2. It primarily prevents HARKing and p-hacking by constraining analytical flexibility. (correct answer)
  3. It primarily ensures that the study's findings will be correctly interpreted by the news media.
  4. It primarily guarantees that the study will be published regardless of the outcome.
Explanation: Pre-registration is an open science practice designed to increase transparency and reduce bias. By committing to a hypothesis and analysis plan before seeing the data, researchers cannot engage in HARKing (Hypothesizing After the Results are Known) or p-hacking (trying multiple analyses and reporting only the significant one). It draws a clear line between confirmatory and exploratory research. While it can help with media interpretation or publication (e.g., in a Registered Report), its primary purpose is to prevent QRPs that distort statistical conclusions.

Question 17

Consider the following sequence of studies: Study 1, an initial experiment, finds a large effect of a 'growth mindset' intervention on academic performance. Study 2, a large, multi-site, pre-registered direct replication, finds no effect on academic performance. Study 3, a conceptual replication, finds the intervention does increase students' self-reported persistence on difficult tasks.

  1. The direct replication (Study 2) was likely flawed, as Study 3 confirmed the original finding from Study 1.
  2. The original effect on academic performance is likely not reliable, but the intervention might influence a related psychological process. (correct answer)
  3. The original finding (Study 1) and the conceptual replication (Study 3) are strong evidence that the direct replication (Study 2) is a Type II error.
  4. The entire line of research is contradictory and should be disregarded until a single, definitive study is conducted.
Explanation: This scenario requires synthesizing multiple, seemingly conflicting, results. The high-quality direct replication (Study 2) failing to find an effect on performance casts serious doubt on the reliability of Study 1's main finding. However, Study 3 suggests the intervention is not inert; it affects a related construct (persistence). The most nuanced conclusion is that the originally claimed outcome (performance) may not be replicable, but the intervention has a different, more subtle effect (on persistence). This integrates all three findings without dismissing any of them outright.

Question 18

A researcher tests a new educational intervention with three different variations (A, B, and C) against a control group. The results show that variation A significantly outperforms the control group, but variations B and C do not. In the final manuscript, the researcher omits all mention of variations B and C and presents the study as a simple, successful comparison between intervention A and a control.

  1. This practice decreases the internal validity of the comparison between variation A and the control group.
  2. This practice prevents an accurate calculation of the effect size for the comparison between variation A and control.
  3. This practice is a form of data falsification and is considered scientific fraud.
  4. This practice misleads readers by creating a cleaner, more compelling narrative than the full data would support. (correct answer)
Explanation: This question tests your understanding of research ethics and publication practices, specifically focusing on selective reporting and its impact on scientific integrity. The researcher's decision to omit unsuccessful variations B and C creates a misleading narrative about the research process. By presenting only the successful comparison between variation A and control, the study appears as a straightforward success story rather than revealing the full picture: that two-thirds of the tested variations actually failed. This selective reporting gives readers an overly optimistic view of the intervention's effectiveness and suggests the research was more targeted and successful than it actually was. The correct answer is D because this practice fundamentally distorts the scientific narrative. Let's examine why the other options are incorrect. Choice A is wrong because internal validity refers to whether the study design properly controls for confounding variables—omitting failed variations doesn't affect the quality of the A versus control comparison itself. Choice B is incorrect because effect size calculations for the A versus control comparison remain mathematically valid regardless of whether other variations are mentioned. Choice C overstates the severity—while this practice is ethically problematic, it's selective reporting rather than data falsification, since no actual data was changed or fabricated. When you encounter research ethics questions, focus on how practices affect transparency and reader understanding. Selective reporting is a key issue in psychology research—it creates publication bias and gives an incomplete picture of what actually works, even when the reported results themselves are accurate.

Question 19

Even when researchers are acting in good faith and not engaging in QRPs, conducting a successful direct replication of a previous study can be challenging. Which of the following is a primary methodological reason for this difficulty?

  1. Original authors are often uncooperative and refuse to share their precise study materials and raw data.
  2. Funding agencies historically have been reluctant to provide grants for replication studies over novel research.
  3. Method sections in papers often omit crucial procedural details or information about the study's context. (correct answer)
  4. The statistical analyses used in psychological research are often too complex for other labs to reproduce accurately.
Explanation: While institutional factors like funding (B) and author cooperation (A) are real challenges, a core methodological reason for replication difficulty is the problem of "tacit knowledge." Method sections, due to space limitations and convention, cannot possibly describe every single detail of a procedure, the demeanor of the experimenter, or subtle characteristics of the sample or setting. A replication attempt may fail not because the original finding was false, but because the replication did not faithfully reproduce a critical, unstated element of the original procedure.

Question 20

In a study on decision-making, a researcher predicts that participants will be more risk-averse after viewing a sad film clip. The initial analysis shows a non-significant result (p = .12). The researcher notices high variance in the data and decides to exclude the 15% of participants with the slowest reaction times, justifying this by arguing they were inattentive. After this exclusion, the result is significant (p = .04).

  1. The procedure is sound because inattentive participants add noise and reduce the power to detect a true effect.
  2. The procedure is questionable because the exclusion rule was implemented after seeing its impact on the p-value. (correct answer)
  3. The procedure is problematic because excluding 15% of the sample is an arbitrary and unjustifiably large proportion.
  4. The procedure is invalid because reaction time is not a legitimate basis for excluding participants in a decision-making study.
Explanation: While it can be legitimate to exclude data based on pre-defined criteria (e.g., 'we will exclude participants with RTs 3 SDs above the mean'), the problem here is that the decision was made post-hoc and was clearly motivated by the desire to achieve significance. This is a form of p-hacking. The researcher is capitalizing on chance by finding a justification to remove data points that are inconvenient for their hypothesis. The core issue is not the criterion itself (D) or the proportion (C), but the data-contingent way in which the criterion was applied.