Home

Tutoring

Subjects

Live Classes

Study Coach

Essay Review

On-Demand Courses

Colleges

Games


Sign up

Log in

Opening subject page...

Loading your content

Practice

  • All Subjects
  • Algebra Flashcards
  • SAT Math Practice Tests
  • Math Question of the Day
  • Live Classes
  • On-Demand Courses

Varsity Tutors

  • Find a Tutor
  • Test Prep
  • Online Classes
  • K-12 Learning
  • College Search
  • VarsityTutors.com

© 2026 Varsity Tutors. All rights reserved.

← Back to quizzes

Genetics Quiz

Genetics Quiz: Gwas Genome Wide Association Studies

Practice Gwas Genome Wide Association Studies in Genetics with focused quiz questions that help you check what you know, review explanations, and build confidence with test-style prompts.

Question 1 / 20

0 of 20 answered

Analyze the Quantile-Quantile (QQ) plot from a GWAS shown. The genomic inflation factor is noted to be λ = 1.45. What is the most appropriate conclusion for the research team to draw?

Select an answer to continue

What this quiz covers

This quiz focuses on Gwas Genome Wide Association Studies, giving you a quick way to practice the rules, question types, and explanations that matter most for Genetics.

How to use this quiz

Try each quiz question before looking at the correct answer. Use the explanations to review missed ideas, then come back to similar questions until the pattern feels familiar.

All questions

Question 1

Analyze the Quantile-Quantile (QQ) plot from a GWAS shown. The genomic inflation factor is noted to be λ = 1.45. What is the most appropriate conclusion for the research team to draw?

  1. The study is well-calibrated and the observed inflation is due to the highly polygenic nature of the trait.
  2. The results are invalid due to severe, uncorrected systematic bias, likely population stratification. (correct answer)
  3. Genotyping quality is poor for the most significant SNPs, leading to an artificially strong signal.
  4. The sample size is too small, resulting in a lack of power and a noisy distribution of p-values.

Explanation: The QQ plot shows a strong, early, and consistent deviation of the observed p-values from the null expectation (the y=x line). This pattern, coupled with a high genomic inflation factor (λ = 1.45), is the classic signature of systemic bias. A well-controlled study should have λ close to 1.0 (e.g., 1.0-1.05). A value of 1.45 indicates widespread inflation of test statistics across the entire genome, most commonly caused by uncorrected population stratification or extensive cryptic relatedness. While high polygenicity (A) can cause a tail-end deviation, it does not explain this level of early and uniform inflation. Poor genotyping of specific SNPs (C) would affect individual points, not the entire distribution. A small sample size (D) would lead to low power, meaning the observed p-values would be less significant and less likely to show such strong inflation.

Question 2

A GWAS on a cohort of 10,000 individuals identifies three novel loci associated with asthma at p < 5x10^-8. According to best practices in the field, what is the most crucial immediate next step for these findings?

  1. Initiating development of a drug that targets the proteins encoded by the genes in these loci.
  2. Conducting a replication study by testing the three lead SNPs in a large, independent cohort. (correct answer)
  3. Performing deep sequencing of the three loci in the original 10,000 individuals to find the causal variant.
  4. Calculating a polygenic risk score for asthma using these three SNPs and testing its clinical utility.

Explanation: The gold standard for validating a GWAS finding is replication in an independent sample. Before investing significant resources in follow-up studies like sequencing (C) or drug development (A), it is essential to confirm that the initial association is not a false positive specific to the discovery cohort (due to chance or subtle biases). A successful replication provides strong evidence that the association is robust. Calculating a PRS (D) is a potential application, but it is premature before the loci are independently validated.

Question 3

Why do researchers primarily perform meta-analyses of multiple GWAS for the same trait?

  1. To increase the overall sample size, which enhances the statistical power to detect variants with small effect sizes. (correct answer)
  2. To identify population-specific genetic effects by comparing results from different ancestries.
  3. To correct for batch effects and other technical artifacts that may have affected the individual studies.
  4. To combine different phenotyping methods for the trait in order to create a more robust disease definition.

Explanation: When you encounter questions about meta-analyses in genetics, focus on the fundamental principle: combining data to overcome the limitations of individual studies. GWAS (genome-wide association studies) face a classic challenge in genetics research—most trait-associated variants have very small effect sizes that require enormous sample sizes to detect reliably. Answer A is correct because meta-analysis directly addresses this power problem. By pooling data from multiple GWAS of the same trait, researchers dramatically increase their total sample size, which provides the statistical power needed to identify variants that might be missed in smaller individual studies. Think of it as turning up the volume on a weak genetic signal until it becomes detectable above the noise. Answer B is incorrect because comparing population-specific effects isn't the primary purpose of meta-analysis—that would be the goal of population stratification or ancestry-specific analyses. Answer C misunderstands the methodology; meta-analyses typically combine summary statistics rather than raw data, so they don't directly correct technical artifacts from individual studies. Answer D confuses meta-analysis with phenotype harmonization—while consistent phenotyping is important for meta-analysis, the goal isn't to create new disease definitions but to leverage existing data. Remember this pattern: when you see questions about combining multiple genetic studies, the answer usually relates to statistical power and sample size. The genetics field constantly battles small effect sizes, making "more data = more power" a central theme in study design questions.

Question 4

In a GWAS with 5,000 cases and 5,000 controls, a common variant (MAF=0.3) has a true odds ratio of 1.5 for a disease. However, the result is not genome-wide significant (p=1x10^-6). What is the most likely reason for this outcome?

  1. The effect size is too small to be detected in any humanly achievable sample size.
  2. The presence of population stratification completely masked an otherwise strong genetic signal.
  3. The variant is in linkage equilibrium with the true causal variant, obscuring the association signal.
  4. The study had insufficient statistical power to detect this moderate effect size at the stringent p < 5x10^-8 threshold. (correct answer)

Explanation: When you encounter GWAS questions, focus on the relationship between statistical power, effect size, sample size, and significance thresholds. This question tests whether you understand when studies fail to reach genome-wide significance despite having a real genetic effect. The study has a moderate effect size (OR = 1.5) and a reasonably large sample (10,000 total participants), but fails to reach the stringent genome-wide significance threshold of p < 5×10⁻⁸. With only 10,000 participants, this study likely lacks sufficient statistical power to detect a moderate effect at such a strict significance level. GWAS require enormous sample sizes (often 100,000+ participants) to reliably detect common variants with modest effects while controlling for multiple testing across millions of SNPs. Option A is incorrect because an odds ratio of 1.5 is definitely detectable with sufficiently large samples - many GWAS have successfully identified variants with similar or smaller effect sizes. Option B overstates population stratification's impact; while stratification can cause problems, it typically wouldn't completely mask a strong signal, and modern GWAS use principal components and other methods to control for this. Option C contains a logical error - variants in linkage equilibrium are independently inherited, so this wouldn't obscure an association signal. You're thinking of linkage disequilibrium, where the tested variant tags the causal variant. Remember that GWAS statistical power depends heavily on the combination of effect size, allele frequency, sample size, and significance threshold. When a study fails to reach genome-wide significance despite reasonable effect sizes, insufficient power is often the culprit.

Question 5

The 'missing heritability' problem describes the observation that genome-wide significant SNPs identified by GWAS explain only a fraction of the heritability estimated from family studies. Which of the following is NOT considered a likely contributor to this gap?

  1. The cumulative effect of many common variants with effect sizes too small to reach genome-wide significance.
  2. The contribution of rare variants that are not well-tagged by standard genotyping arrays.
  3. Systematic overestimation of heritability from twin and family studies due to shared environmental factors.
  4. The existence of a few undiscovered variants with very large, Mendelian-like effects on the trait. (correct answer)

Explanation: The 'missing heritability' puzzle has several proposed solutions. Plausible contributors include many common variants of tiny effect that fall below the stringent significance threshold (A), rare variants not captured by GWAS arrays (B), and potential overestimation of heritability from classical methods (C). However, the existence of major undiscovered Mendelian-like variants (D) is considered unlikely for most common, complex traits. GWAS is well-powered to detect common variants with large effects; if such variants existed, they would have been among the first and easiest to find.

Question 6

How are the results of a large-scale GWAS typically used to create a polygenic risk score (PRS) for an individual?

  1. By summing the dosages of risk alleles carried by the individual, each weighted by its GWAS-derived effect size. (correct answer)
  2. By counting how many of the genome-wide significant risk alleles the individual carries.
  3. By sequencing all genes identified in the GWAS to find the total number of rare pathogenic mutations.
  4. By measuring the linkage disequilibrium between the top associated SNPs in that individual's genome.

Explanation: When you encounter questions about polygenic risk scores (PRS), remember that these scores quantify an individual's genetic predisposition to a trait or disease based on many genetic variants across the genome, each contributing a small effect. A polygenic risk score is calculated by taking each risk allele an individual carries and weighting it by the effect size discovered in the GWAS study, then summing all these weighted contributions. The effect size reflects how much each variant increases disease risk, so variants with larger effects contribute more to the final score. This approach captures both which risk alleles you carry and how important each one is based on the research findings. Looking at the incorrect options: Option B oversimplifies PRS calculation by just counting significant risk alleles without considering their different effect sizes—a variant that doubles disease risk shouldn't be weighted the same as one that increases risk by only 5%. Option C confuses PRS methodology with rare variant analysis; PRS typically uses common variants identified in GWAS, not rare pathogenic mutations found through sequencing. Option D describes a population genetics measurement rather than an individual risk calculation—linkage disequilibrium patterns don't directly translate to personal risk scores. The correct answer is A because it captures the essential PRS formula: summing risk allele dosages weighted by their GWAS-derived effect sizes. Study tip: Remember that PRS questions often test whether you understand the difference between simply counting risk variants versus properly weighting them by their biological importance—always look for the weighted approach in PRS calculations.

Question 7

Following a GWAS that identifies a 250kb region of association for a disease, researchers initiate a 'fine-mapping' study. What is the primary objective of this follow-up study?

  1. To test for gene-environment interactions involving the identified region and lifestyle factors.
  2. To replicate the initial association signal in a larger, more diverse population cohort.
  3. To increase the density of genotyped or imputed variants in the region to narrow down the set of potential causal SNPs. (correct answer)
  4. To estimate the total proportion of disease heritability that is explained by this single genomic region.

Explanation: Fine-mapping is the process that follows the discovery of an associated locus. The goal is to move from a large region of correlated SNPs (an LD block) to a much smaller, credible set of variants that are most likely to be the true causal variant(s). This is achieved by using sequencing or dense imputation to get information on all variants in the region and then applying statistical methods to prioritize them based on their association strength and LD patterns. Replication (B) precedes fine-mapping. Testing for interactions (A) and estimating regional heritability (D) are other types of follow-up analyses but do not define fine-mapping.

Question 8

A large genome-wide association study (GWAS) for Crohn's disease identifies a SNP with a p-value of 1x10^-15. This SNP is located in a 100kb region of high linkage disequilibrium (LD) that contains three genes. Subsequent fine-mapping and functional studies fail to identify a causal role for this specific SNP. Which of the following is the most likely explanation for the initial GWAS result?

  1. The initial association was a Type I error due to insufficient multiple testing correction.
  2. The identified SNP is a 'tag SNP' that is strongly correlated with a true, ungenotyped causal variant elsewhere in the LD block. (correct answer)
  3. The three genes in the region must act together epistatically to cause the disease, making the effect of a single SNP undetectable.
  4. The study was confounded by severe population stratification, creating a spurious association signal in this genomic region.

Explanation: The most fundamental concept in interpreting GWAS hits is linkage disequilibrium. A highly significant SNP (the 'lead SNP' or 'tag SNP') is often not the biologically functional variant itself. Instead, it is in strong non-random association (LD) with the true causal variant, which might not have been included on the genotyping array. The entire LD block shows a signal because all variants within it are correlated. Distractor A is unlikely given the extremely low p-value, which far surpasses the standard genome-wide significance threshold (e.g., 5x10^-8). Distractor C proposes epistasis, which is possible but not the most direct or common explanation for a single strong locus. Distractor D is a possible issue in any GWAS, but the scenario describes a very strong, localized signal, which is more characteristic of a true locus in high LD than a genome-wide artifact.

Question 9

Standard GWAS designs using common SNP arrays are known to have limitations. Which of the following genetic association types are they least equipped to detect?

  1. Associations with common SNPs (MAF > 5%) that have very small additive effects on a trait.
  2. Associations for polygenic traits that are influenced by hundreds or thousands of loci.
  3. Associations located within non-coding, intergenic regions of the genome.
  4. Associations with structural variants, such as large copy number variations (CNVs). (correct answer)

Explanation: When approaching GWAS limitation questions, consider what standard SNP arrays can and cannot detect based on their technical design. These arrays rely on common SNPs to tag genomic regions through linkage disequilibrium. Standard GWAS using SNP arrays struggle most with structural variants like large CNVs because these arrays weren't designed to detect them. SNP arrays identify single nucleotide changes at predetermined positions, but structural variants involve large insertions, deletions, or duplications that span thousands to millions of base pairs. The sparse spacing of SNPs on arrays means they often miss the breakpoints of these variants entirely, and even when nearby SNPs exist, they're poorly correlated with structural variant presence. This makes option D correct. Option A is wrong because GWAS can detect small-effect common variants—that's exactly what they're designed for. With large enough sample sizes (hundreds of thousands of individuals), even tiny effects become statistically detectable. Option B is incorrect because polygenic traits are precisely what GWAS excel at studying. The method can identify multiple independent associations across many loci contributing to complex traits. Option C is wrong because SNP arrays cover both coding and non-coding regions extensively. Many GWAS hits actually fall in intergenic regions, and these associations are readily detected when the trait-causing variant is in linkage disequilibrium with array SNPs. For genetics exams, remember that GWAS limitations typically involve what the technology cannot "see"—rare variants, structural variants, and poorly tagged regions—rather than what statistical approaches can handle.

Question 10

A GWAS for educational attainment identifies 1,271 independent genome-wide significant loci. A critic argues the study is flawed because 'no single gene can determine educational attainment.' How does the design and typical result of a GWAS address this criticism?

  1. The criticism is valid; GWAS is only suitable for simple Mendelian traits, not complex behavioral ones.
  2. GWAS corrects for this by focusing only on the single SNP with the lowest p-value as the sole determinant of the trait.
  3. The study identifies causal genes, and the large number found proves the trait is determined by many individual genes.
  4. GWAS assumes a polygenic model, where each significant locus contributes a very small, additive effect to the overall phenotype. (correct answer)

Explanation: When you encounter GWAS (Genome-Wide Association Studies) questions, remember that these studies are specifically designed to identify genetic variants associated with complex traits that are influenced by many genes, each with small effects. The critic's argument reflects a fundamental misunderstanding of how GWAS works. GWAS operates under a polygenic model, which assumes that complex traits like educational attainment result from the combined effects of many genetic variants, each contributing a tiny amount to the overall phenotype. The discovery of 1,271 independent loci actually supports this model perfectly—it shows that educational attainment is influenced by hundreds of genetic variants scattered across the genome, each with a very small individual effect size (typically explaining less than 1% of trait variance). Option A is wrong because GWAS was specifically developed for complex traits, not simple Mendelian ones. Simple traits typically involve single genes with large effects, which don't require genome-wide scanning. Option B misrepresents GWAS methodology—these studies examine all identified significant variants together, not just the most statistically significant one. Option C incorrectly claims that GWAS identifies causal genes when it actually identifies associated genetic variants; establishing causation requires additional functional studies. The large number of significant loci found in this study demonstrates exactly what GWAS is designed to detect: polygenic architecture where many variants of small effect combine to influence a complex trait. Study tip: For GWAS questions, remember the key principle: many variants, small effects, complex traits. This distinguishes GWAS from single-gene disease studies and explains why finding hundreds or thousands of associated loci is actually expected, not problematic.

Question 11

A GWAS finds a strong association between a SNP in the CYP1A2 gene region and daily coffee consumption. CYP1A2 is known to be the primary enzyme for caffeine metabolism. The same SNP is also strongly associated with the number of cigarettes smoked per day. What is the most significant challenge in interpreting the coffee consumption association?

  1. The association may be confounded by smoking behavior, which is correlated with both the SNP and coffee drinking. (correct answer)
  2. The effect size of the SNP is likely too small to be of any biological importance for caffeine metabolism.
  3. The true causal variant is likely a rare mutation that was not genotyped on the SNP array.
  4. The result is probably a false positive because behavioral traits are not strongly influenced by genetics.

Explanation: This scenario describes potential confounding. Smokers tend to drink more coffee, and smoking induces the CYP1A2 enzyme, leading to faster caffeine metabolism. If the SNP is associated with smoking, it will appear to be associated with coffee consumption through this behavioral link, even if it has no direct biological effect on coffee preference. This confounding effect must be statistically adjusted for (e.g., by including smoking status as a covariate) before one can claim a direct genetic association with coffee consumption. B is a statement about effect size, not interpretation. C is a general point about LD, but confounding is the more immediate issue here. D is an incorrect generalization; many behavioral traits have a genetic component.

Question 12

A Quantile-Quantile (QQ) plot generated from a GWAS of hypertension shows that the observed p-values begin to deviate from the expected diagonal line almost immediately and continue to deviate uniformly across the entire distribution. The genomic inflation factor (λ) is calculated to be 1.3. What is the most probable cause of this pattern?

  1. The trait has a highly polygenic architecture with thousands of true associations.
  2. The study is well-controlled, and the deviation represents a set of strong, true-positive signals.
  3. Systematic bias from unaccounted-for population stratification is inflating test statistics. (correct answer)
  4. A single, rare variant of large effect is driving the association signal for the disease.

Explanation: An early, uniform deviation of observed p-values from the expected null distribution on a QQ plot, reflected by a genomic inflation factor (λ) substantially greater than 1, is the classic sign of systemic bias. The most common cause is uncorrected population stratification, where allele frequency differences between subpopulations in cases and controls create spurious associations across the genome. While polygenicity (A) can cause deviation, it typically manifests as a tail of inflation at the most significant end, not a uniform shift from the beginning. A well-controlled study (B) would show points hugging the diagonal until a late tail departure. A single rare variant (D) would not cause a genome-wide inflation of p-values.

Question 13

The threshold for declaring genome-wide significance in a GWAS is typically set at p < 5x10^-8. What is the primary justification for using such a stringent threshold?

  1. To ensure that any identified associations have a large biological effect size and clinical relevance.
  2. To account for the high probability of finding chance associations when performing millions of statistical tests. (correct answer)
  3. To correct for confounding variables such as population stratification and cryptic relatedness.
  4. To increase the statistical power of the study to detect associations with rare genetic variants.

Explanation: The stringent p-value threshold is a direct consequence of the multiple testing problem. A typical GWAS tests ~1 million independent common variants. A Bonferroni correction for this would be 0.05 / 1,000,000 = 5x10^-8. This threshold is designed to keep the family-wise error rate (the probability of making at least one Type I error) at an acceptable level (e.g., 5%). It does not relate to effect size (A), as a very significant p-value can be associated with a tiny effect size. Confounding variables (C) are addressed through study design and statistical adjustments (like including principal components as covariates), not by the significance threshold itself. A more stringent threshold actually decreases statistical power (D), making it harder to detect true associations.

Question 14

What is the primary advantage of a genome-wide association study (GWAS) compared to a candidate gene study for investigating the genetics of a complex disease?

  1. GWAS requires a much smaller sample size to achieve statistical significance for an association.
  2. GWAS is a hypothesis-free approach that can identify novel genes and biological pathways. (correct answer)
  3. GWAS directly identifies the causal functional variant rather than just a correlated marker.
  4. GWAS is less susceptible to confounding by factors like population stratification.

Explanation: The key strength of GWAS is its 'hypothesis-free' nature. It surveys the entire genome without a priori assumptions about which genes might be involved. This allows for the discovery of completely novel associations and biological pathways that would be missed by a candidate gene study, which is limited to testing genes already suspected of involvement. GWAS requires a much larger, not smaller, sample size (A) due to the massive multiple testing burden. It does not directly identify causal variants (C), but rather regions of LD that are associated. GWAS is highly susceptible to population stratification (D), which must be carefully controlled for.

Question 15

A researcher is examining a Manhattan plot from a GWAS. The y-axis of the plot is labeled '-log10(p-value)'. A SNP with a p-value of 1x10^-5 would be plotted at what value on this y-axis?

  1. 5 (correct answer)
  2. -5
  3. 0.00001
  4. 8

Explanation: The y-axis of a Manhattan plot displays the negative base-10 logarithm of the p-value for each SNP's association test. This transformation converts small p-values (indicating strong evidence of association) into large, positive numbers that are easier to visualize. For a p-value of 1x10^-5, the calculation is -log10(10^-5) = -(-5) = 5. Distractor B makes a sign error. Distractor C is the p-value itself, not its transformed value. Distractor D corresponds to a p-value of 1x10^-8, which is the typical threshold for genome-wide significance, a common point of confusion.

Question 16

A GWAS identifies a significant association between an intronic SNP within the FTO gene and obesity. Which of the following represents the most plausible molecular mechanism for this association?

  1. The SNP introduces a premature stop codon, resulting in a truncated, non-functional FTO protein.
  2. The SNP alters the protein-coding sequence, leading to an FTO enzyme with modified metabolic activity.
  3. The SNP disrupts a binding site for a regulatory protein within the intron, altering the expression of FTO or a distant gene like IRX3. (correct answer)
  4. The SNP is a benign marker and the true causal variant must be a coding SNP in a different gene located on another chromosome.

Explanation: Since the SNP is located in an intron (a non-coding region), it cannot directly alter the protein sequence by creating a stop codon (A) or changing an amino acid (B). The most likely mechanism for intronic variants is that they affect gene regulation. Introns can contain regulatory elements like enhancers or silencers. The SNP could disrupt the binding of a transcription factor or other regulatory protein, thereby altering the expression levels of the host gene (FTO) or even distant genes that are brought into physical proximity through chromatin looping (as is the case for FTO variants affecting IRX3). While the causal variant could be elsewhere in the LD block, it's very plausible for an intronic SNP to be functional itself; therefore, D is a less direct explanation of the SNP's potential mechanism.

Question 17

A researcher is examining a Manhattan plot from a GWAS. The y-axis of the plot is labeled '-log10(p-value)'. A SNP with a p-value of 1x10^-5 would be plotted at what value on this y-axis?

  1. 5 (correct answer)
  2. -5
  3. 0.00001
  4. 8

Explanation: The y-axis of a Manhattan plot displays the negative base-10 logarithm of the p-value for each SNP's association test. This transformation converts small p-values (indicating strong evidence of association) into large, positive numbers that are easier to visualize. For a p-value of 1x10^-5, the calculation is -log10(10^-5) = -(-5) = 5. Distractor B makes a sign error. Distractor C is the p-value itself, not its transformed value. Distractor D corresponds to a p-value of 1x10^-8, which is the typical threshold for genome-wide significance, a common point of confusion.

Question 18

A research team conducts a GWAS for a rare autoimmune disease. They recruit cases from a specialized clinic in Northern Europe and controls from a general population database in Southern Europe. The study identifies numerous loci with highly significant associations. Which action is most critical before interpreting these loci as being disease-related?

  1. Sequencing the associated loci to identify the precise causal mutations in the case subjects.
  2. Replicating the associations in an independent cohort with carefully matched ancestry for cases and controls. (correct answer)
  3. Performing functional studies on the genes nearest to the most significant SNPs to understand their mechanism.
  4. Increasing the p-value threshold to 1x10^-5 to account for the rarity of the disease being studied.

Explanation: The study design has a major flaw: cases and controls are drawn from genetically distinct populations (Northern vs. Southern Europe). This creates massive potential for confounding by population stratification. Any genetic variants that differ in frequency between these two populations will appear to be associated with the disease, regardless of any true biological link. Therefore, the most critical step is to attempt to replicate the findings in a new, independent study where cases and controls are carefully matched for genetic ancestry. This will determine if the signals are real or simply artifacts of the poor initial design. Functional studies (A, C) are premature until the associations are validated. Relaxing the p-value threshold (D) would only increase the number of false positives.

Question 19

What is the primary advantage of a genome-wide association study (GWAS) compared to a candidate gene study for investigating the genetics of a complex disease?

  1. GWAS requires a much smaller sample size to achieve statistical significance for an association.
  2. GWAS is a hypothesis-free approach that can identify novel genes and biological pathways. (correct answer)
  3. GWAS directly identifies the causal functional variant rather than just a correlated marker.
  4. GWAS is less susceptible to confounding by factors like population stratification.

Explanation: The key strength of GWAS is its 'hypothesis-free' nature. It surveys the entire genome without a priori assumptions about which genes might be involved. This allows for the discovery of completely novel associations and biological pathways that would be missed by a candidate gene study, which is limited to testing genes already suspected of involvement. GWAS requires a much larger, not smaller, sample size (A) due to the massive multiple testing burden. It does not directly identify causal variants (C), but rather regions of LD that are associated. GWAS is highly susceptible to population stratification (D), which must be carefully controlled for.

Question 20

A GWAS on a cohort of 10,000 individuals identifies three novel loci associated with asthma at p < 5x10^-8. According to best practices in the field, what is the most crucial immediate next step for these findings?

  1. Initiating development of a drug that targets the proteins encoded by the genes in these loci.
  2. Conducting a replication study by testing the three lead SNPs in a large, independent cohort. (correct answer)
  3. Performing deep sequencing of the three loci in the original 10,000 individuals to find the causal variant.
  4. Calculating a polygenic risk score for asthma using these three SNPs and testing its clinical utility.

Explanation: The gold standard for validating a GWAS finding is replication in an independent sample. Before investing significant resources in follow-up studies like sequencing (C) or drug development (A), it is essential to confirm that the initial association is not a false positive specific to the discovery cohort (due to chance or subtle biases). A successful replication provides strong evidence that the association is robust. Calculating a PRS (D) is a potential application, but it is premature before the loci are independently validated.