All questions
Question 1
The human reference genome sequence (e.g., GRCh38) shows the allele 'T' at a chromosomal position known to be a common C/T SNP. In a global population survey, the 'C' allele is found to have a frequency of 75%. Which statement accurately explains this situation?
- The reference genome always contains the ancestral allele, and 'T' must therefore be the original human allele.
- This indicates a systematic error in either the reference sequence or the population frequency data.
- The 'T' allele must be the one associated with normal function, and the 'C' allele is pathogenic.
- The reference genome was constructed from individuals who happened to carry the minor 'T' allele at this position. (correct answer)
Explanation: When you encounter questions about reference genomes and allele frequencies, remember that reference sequences are constructed from specific individuals, not from population-wide consensus data. This creates situations where the reference allele may not be the most common variant in the global population.
The correct explanation is D: the reference genome contains the 'T' allele simply because the individuals whose DNA was used to construct that particular genomic region happened to carry 'T' at this position. Reference genomes are mosaics built from multiple donors, and each position reflects whatever allele those specific donors carried, regardless of global frequency patterns.
Option A is incorrect because reference genomes don't necessarily contain ancestral alleles—they contain whatever alleles the donor individuals had. Ancestral state determination requires separate evolutionary analysis. Option B wrongly assumes there must be an error when reference and population data don't align. This misunderstanding stems from thinking the reference should represent the most common allele, which isn't how reference genomes are constructed. Option C makes an unfounded assumption about functional significance. Allele frequency doesn't directly indicate pathogenicity—many common variants are neutral, and some rare variants are actually beneficial.
This scenario illustrates why modern genomics increasingly uses population-specific reference panels rather than relying solely on single reference sequences. For genetics exams, remember that reference genomes are historical snapshots from specific individuals, not population summaries. Always distinguish between reference sequences, population frequencies, ancestral states, and functional significance—these are independent concepts that students often conflate.
Question 2
A researcher observes a significant deficit of heterozygotes at a SNP locus in a large, randomly collected human population sample, compared to Hardy-Weinberg expectations. Which of the following is the most plausible biological or technical explanation for this deviation?
- The SNP provides a strong heterozygous advantage (overdominance) in this population.
- The population sample is inadvertently a mix of two or more distinct subpopulations (Wahlund effect). (correct answer)
- The mutation rate at this locus is unusually high, constantly creating new homozygotes.
- The genotyping assay used has a high false-negative rate for calling homozygous genotypes.
Explanation: The correct answer is B. A deficit of heterozygotes is a classic sign of the Wahlund effect, which occurs when distinct subpopulations with different allele frequencies are pooled and analyzed as a single group. Within each subpopulation, the SNP may be in HWE, but when they are mixed, the overall sample shows a heterozygote deficit. This is a form of non-random mating. Choice A, heterozygote advantage, would lead to an excess of heterozygotes, not a deficit. Choice C is incorrect; mutation rates are far too low to cause a detectable deviation from HWE in a single generation. Choice D describes a technical artifact, but a high false-negative rate for homozygotes would lead to them being miscalled as heterozygotes, causing an apparent excess of heterozygotes.
Question 3
Two linked SNPs, rsA (alleles A/G) and rsB (alleles C/T), have allele frequencies of p(A)=0.6 and p(C)=0.5. If the frequency of the A-C haplotype in the population is observed to be 0.42, what can be concluded about these two SNPs?
- They are in perfect linkage disequilibrium (D' = 1.0).
- They are in linkage equilibrium, and their inheritance is independent.
- They are in positive linkage disequilibrium, with A and C alleles associated more often than expected. (correct answer)
- They are in negative linkage disequilibrium, with A and C alleles associated less often than expected.
Explanation: The correct answer is C. This requires a two-step analysis. First, calculate the expected frequency of the A-C haplotype if the SNPs were in linkage equilibrium (i.e., independent). Expected P(A-C) = p(A) * p(C) = 0.6 * 0.5 = 0.30. Second, compare the observed frequency (0.42) to the expected frequency (0.30). Since the observed frequency is greater than the expected frequency, the alleles A and C are found together on the same chromosome more often than expected by chance. This is defined as positive linkage disequilibrium. The disequilibrium coefficient, D, would be P(AC) - p(A)p(C) = 0.42 - 0.30 = 0.12. Since D is positive, the LD is positive. Choice B is incorrect because observed and expected frequencies differ. Choice D is incorrect because the association is more frequent, not less. Choice A is an overstatement; while they are in LD, we cannot conclude D'=1.0 without knowing the other haplotype frequencies.
Question 4
A researcher identifies a genetic variant where some individuals in a population have two copies of a 25kb gene-containing segment, while others have three or four copies. This type of variation is best classified as a:
- Single Nucleotide Polymorphism (SNP), because it is a heritable change in the DNA.
- Copy Number Variant (CNV), because it involves a change in the dosage of a large genomic segment. (correct answer)
- Frameshift mutation, because the addition of DNA alters the reading frame of the gene.
- Haplotype, because it represents a combination of alleles found on the same chromosome.
Explanation: The correct answer is B. A Copy Number Variant (CNV) is defined as a variation in the number of copies of a specific DNA segment that is 1kb or larger. The scenario describes variation in the dosage of a 25kb segment, which fits the definition of a CNV perfectly. Choice A is incorrect because a SNP is a change at a single nucleotide base, not a large segment. Choice C is incorrect; a frameshift is typically caused by an insertion or deletion of a number of nucleotides not divisible by three within a coding sequence, while a CNV involves the duplication of an entire gene or region. Choice D is incorrect; a haplotype refers to the specific combination of linked alleles (like SNPs) on a chromosome, not the copy number of a large segment.
Question 5
A genome-wide association study (GWAS) for hypertension identifies a single nucleotide polymorphism (SNP), rs12345, located within an intron of the ACE2 gene, as having the strongest association signal (p = 1 x 10⁻¹⁵). Subsequent functional studies fail to show any effect of this SNP on ACE2 splicing or expression. What is the most likely explanation for the strong association signal?
- The association is a false positive (Type I error) due to population stratification in the study cohort.
- The intronic SNP is in strong linkage disequilibrium with a nearby, ungenotyped causal variant that alters gene function. (correct answer)
- The SNP directly alters the tertiary structure of the ACE2 protein, leading to increased blood pressure.
- The finding is irrelevant because intronic variations are phenotypically silent and do not contribute to complex traits.
Explanation: The correct answer is B. GWAS identifies statistical associations, not causal variants. The 'lead' SNP with the lowest p-value is often not the functional variant itself but is instead strongly correlated (in linkage disequilibrium) with the true causal variant, which may not have been included on the genotyping array. This causal variant could be in a regulatory region or a coding region of a nearby gene. Choice A is a possibility for any GWAS hit, but a p-value this low makes it less likely to be the most likely explanation compared to LD. Choice C is incorrect because an intronic SNP does not alter the protein sequence. Choice D is a common misconception; intronic and other non-coding variants can have significant regulatory effects and are major contributors to complex traits.
Question 6
Consider two genomic regions of equal physical length (100 kb). Region A has a high meiotic recombination rate, while Region B is a recombination 'coldspot' with a very low rate. Assuming both regions have a similar underlying SNP density, which statement accurately compares their haplotype structures?
- Region A will have more distinct haplotypes, each spanning a shorter genetic distance, than Region B. (correct answer)
- Region B will have more distinct haplotypes due to the accumulation of mutations that are not shuffled.
- Both regions will have an identical number of haplotypes since they have the same SNP density.
- Region A will exhibit stronger linkage disequilibrium across the 100 kb distance than Region B.
Explanation: The correct answer is A. Recombination acts to shuffle existing genetic variation. In a region with a high recombination rate (Region A), associations between SNPs are broken down frequently. This creates a large number of different combinations of alleles, or distinct haplotypes, but these haplotypes are short because the correlation (LD) decays quickly with distance. In contrast, in a recombination coldspot (Region B), SNPs are rarely separated, so they are inherited together in large blocks. This results in fewer, longer haplotypes and strong LD across the region. Therefore, Region A has more, shorter haplotypes. Choice B is incorrect; low recombination leads to fewer haplotypes. Choice C is incorrect because recombination, not just SNP density, determines haplotype diversity. Choice D is incorrect; high recombination leads to weaker, not stronger, LD.
Question 7
In a trio analysis (mother, father, child), a specific locus is genotyped. The mother is homozygous A/A, and the father is homozygous A/A. The child, however, is heterozygous A/G. Assuming no sample mix-up or genotyping error, what is the most likely genetic explanation for the child's genotype?
- The G allele shows incomplete penetrance in the parents.
- The locus is subject to genomic imprinting from the paternal allele.
- The G allele is a recessive mutation that was masked in the parents.
- A de novo mutation occurred in the germline of one of the parents. (correct answer)
Explanation: When analyzing inheritance patterns in trio studies, you should always consider whether the observed genotypes can be explained by standard Mendelian inheritance before exploring more complex mechanisms.
In this case, both parents are homozygous A/A, meaning they can only contribute A alleles to their offspring. Under normal Mendelian inheritance, their child should be A/A. However, the child is A/G, which means the G allele had to come from somewhere. Since neither parent carries a G allele, the most straightforward explanation is that a de novo mutation occurred in the germline of one parent, converting an A allele to a G allele during gamete formation. This makes option D correct.
Let's examine why the other options don't fit. Option A incorrectly applies penetrance—a concept about gene expression—to explain the presence of an entirely different allele. Penetrance affects whether a genotype produces its expected phenotype, not whether an allele exists. Option B suggests genomic imprinting, but this mechanism affects gene expression patterns, not the fundamental presence or absence of alleles in the parents. Option C misunderstands recessiveness; even if G were recessive, the parents would still need to carry the G allele (as A/G heterozygotes) to pass it to their child, which contradicts their A/A genotypes.
Remember: when you see unexpected alleles in offspring that neither parent carries, think de novo mutation first. This is a key pattern in medical genetics and explains many spontaneous genetic disorders.
Question 8
The primary rationale for using tag SNPs in a genome-wide association study is to:
- directly identify the causal variants responsible for the phenotype of interest.
- increase the statistical power by focusing only on SNPs with high minor allele frequency.
- reduce genotyping costs by selecting a subset of SNPs that efficiently capture the haplotype diversity in a genomic region. (correct answer)
- ensure that only evolutionarily conserved, functionally significant SNPs are included in the analysis.
Explanation: The correct answer is C. Due to linkage disequilibrium, alleles at nearby SNPs are inherited together in blocks (haplotypes). A tag SNP is a representative SNP within a region of high LD. By genotyping just the tag SNP, one can infer the alleles of the many other SNPs it is correlated with. This strategy allows researchers to capture most of the common genetic variation in a region without having to genotype every single SNP, thereby significantly reducing the cost and complexity of the study. Choice A is incorrect; tag SNPs are proxies for variation, not necessarily the causal variants themselves. Choice B is an oversimplification; while tag SNPs often have a sufficient MAF, their primary selection criterion is how well they represent other SNPs. Choice D is incorrect; tag SNPs are chosen based on LD patterns, not necessarily on evolutionary conservation or known function.
Question 9
A pharmacogenomic test reveals that a patient is homozygous for a loss-of-function SNP in the CYP2D6 gene. This enzyme is responsible for metabolizing codeine into its active form, morphine. What is the most likely clinical outcome if this patient is prescribed a standard dose of codeine for pain?
- The patient will experience a greatly exaggerated analgesic effect, risking morphine overdose.
- The patient will experience little to no pain relief because the active drug metabolite is not being produced. (correct answer)
- The patient will rapidly clear the codeine from their system, requiring a higher-than-standard dose.
- There will be no difference in effect, as metabolic pathways are highly redundant.
Explanation: The correct answer is B. This is a classic pharmacogenomics problem. Codeine is a pro-drug, meaning it is inactive until it is metabolized by the body into its active form, which is morphine. The enzyme CYP2D6 carries out this conversion. A patient with a loss-of-function variant in CYP2D6 is a 'poor metabolizer'. They cannot efficiently convert codeine to morphine. Therefore, even with a standard dose of codeine, very little of the active analgesic will be produced, and the patient will experience inadequate pain relief. Choice A describes an 'ultra-rapid metabolizer' who has multiple active copies of the gene. Choice C is the opposite of a poor metabolizer. Choice D is incorrect; while some redundancy exists, CYP2D6 is the primary pathway for codeine activation and its loss has significant clinical effects.
Question 10
An individual's polygenic risk score (PRS) for coronary artery disease is calculated to be in the 99th percentile of the population. Which of the following is the most accurate interpretation of this result?
- The individual has a 99% chance of developing coronary artery disease during their lifetime.
- The individual carries a single, rare genetic variant that confers a very high risk for the disease.
- The individual has inherited a large number of common, small-effect risk alleles for the disease. (correct answer)
- The individual will certainly develop coronary artery disease, regardless of lifestyle factors like diet and exercise.
Explanation: The correct answer is C. A polygenic risk score summarizes the combined effect of many (often thousands) of common genetic variants (SNPs) across the genome. Each variant contributes a small amount to the overall risk. A high PRS, such as being in the 99th percentile, indicates that the individual has inherited an aggregate of risk-conferring alleles that is greater than 99% of the reference population. Choice A is incorrect; a PRS indicates relative risk, not absolute probability. Choice B is incorrect; a PRS is based on the additive effects of many common variants, not a single rare one (which would be considered monogenic). Choice D is incorrect because a PRS represents genetic predisposition, not a deterministic outcome. Lifestyle and environmental factors play a crucial role in modifying the risk of complex diseases.
Question 11
A population is founded by a small number of individuals who are reproductively isolated. Over many generations, this population would be expected to exhibit which pattern of genetic variation compared to the larger, ancestral population?
- An increase in overall heterozygosity and a greater number of unique haplotypes.
- Allele frequencies identical to the ancestral population due to balancing selection.
- A reduction in overall genetic variation and an increase in the extent of linkage disequilibrium. (correct answer)
- A significantly higher rate of de novo mutation to compensate for the lack of gene flow.
Explanation: The correct answer is C. This scenario describes a founder effect, a form of genetic drift. The small number of founders likely carries only a subset of the genetic variation present in the ancestral population, leading to an immediate reduction in overall variation. Due to the small population size, genetic drift will continue to cause random fluctuations in allele frequencies, often leading to the loss of more alleles. Furthermore, because all individuals descend from a few common ancestors, specific combinations of alleles (haplotypes) can become much more common than in the ancestral population, leading to an increase in the extent of linkage disequilibrium. Choice A is the opposite of what is expected. Choice B is incorrect; drift will cause allele frequencies to diverge significantly from the ancestral population. Choice D is incorrect; the mutation rate is a fundamental biological property and is not affected by population size or gene flow.
Question 12
A SNP located 2 kb upstream of the transcription start site of the INS gene is strongly associated with type 1 diabetes. This SNP does not alter any protein sequence. What is its most plausible biological mechanism of action?
- It alters the binding affinity of a transcription factor, leading to dysregulated INS gene expression. (correct answer)
- It introduces a premature stop codon, resulting in a truncated, non-functional insulin protein.
- It changes the rate of post-translational modification of the proinsulin polypeptide.
- It causes a frameshift mutation during translation of the INS mRNA.
Explanation: The correct answer is A. A SNP located upstream of a transcription start site is in the promoter or a regulatory region. These regions contain binding sites for transcription factors that control the rate of gene expression. A change in this DNA sequence can alter the binding affinity of such a factor, either increasing or decreasing it, which in turn dysregulates the expression of the target gene (in this case, insulin). Choice B is incorrect because the SNP is not in the coding sequence and cannot create a stop codon. Choice D is incorrect because SNPs are single base substitutions and do not cause frameshifts (which are caused by insertions/deletions). Choice C is incorrect because post-translational modification is a process that occurs on the protein, and a DNA variant in a non-coding region would not directly affect it.
Question 13
A population's genotypes for a SNP are counted as follows: 490 individuals are C/C, 420 are C/T, and 90 are T/T. What is the minor allele frequency (MAF) for this SNP, and what does this value imply about its utility in an association study?
- The MAF is 0.30; the SNP is highly informative due to its common frequency. (correct answer)
- The MAF is 0.09; the SNP is of limited use as the minor allele is too rare.
- The MAF is 0.42; the SNP has high heterozygosity, but is not in Hardy-Weinberg equilibrium.
- The MAF is 0.60; the SNP is uninformative because the 'minor' allele is actually the major allele.
Explanation: The correct answer is A. First, calculate the total number of individuals: 490 + 420 + 90 = 1000. The total number of alleles is 2000. The number of T alleles is (2 * 90) + 420 = 180 + 420 = 600. The frequency of the T allele is 600 / 2000 = 0.30. The number of C alleles is (2 * 490) + 420 = 980 + 420 = 1400. The frequency of the C allele is 1400 / 2000 = 0.70. The minor allele is the one with the lower frequency, which is T (0.30). A MAF of 0.30 (or 30%) is considered common, making the SNP very informative for association studies as both alleles are well-represented in the population, which increases statistical power. Choice B incorrectly uses the frequency of the T/T genotype as the MAF. Choice C incorrectly identifies the heterozygote frequency as the MAF. Choice D misidentifies the major allele frequency as the MAF.
Question 14
An individual's polygenic risk score (PRS) for coronary artery disease is calculated to be in the 99th percentile of the population. Which of the following is the most accurate interpretation of this result?
- The individual has a 99% chance of developing coronary artery disease during their lifetime.
- The individual carries a single, rare genetic variant that confers a very high risk for the disease.
- The individual has inherited a large number of common, small-effect risk alleles for the disease. (correct answer)
- The individual will certainly develop coronary artery disease, regardless of lifestyle factors like diet and exercise.
Explanation: The correct answer is C. A polygenic risk score summarizes the combined effect of many (often thousands) of common genetic variants (SNPs) across the genome. Each variant contributes a small amount to the overall risk. A high PRS, such as being in the 99th percentile, indicates that the individual has inherited an aggregate of risk-conferring alleles that is greater than 99% of the reference population. Choice A is incorrect; a PRS indicates relative risk, not absolute probability. Choice B is incorrect; a PRS is based on the additive effects of many common variants, not a single rare one (which would be considered monogenic). Choice D is incorrect because a PRS represents genetic predisposition, not a deterministic outcome. Lifestyle and environmental factors play a crucial role in modifying the risk of complex diseases.
Question 15
The primary rationale for using tag SNPs in a genome-wide association study is to:
- directly identify the causal variants responsible for the phenotype of interest.
- increase the statistical power by focusing only on SNPs with high minor allele frequency.
- reduce genotyping costs by selecting a subset of SNPs that efficiently capture the haplotype diversity in a genomic region. (correct answer)
- ensure that only evolutionarily conserved, functionally significant SNPs are included in the analysis.
Explanation: The correct answer is C. Due to linkage disequilibrium, alleles at nearby SNPs are inherited together in blocks (haplotypes). A tag SNP is a representative SNP within a region of high LD. By genotyping just the tag SNP, one can infer the alleles of the many other SNPs it is correlated with. This strategy allows researchers to capture most of the common genetic variation in a region without having to genotype every single SNP, thereby significantly reducing the cost and complexity of the study. Choice A is incorrect; tag SNPs are proxies for variation, not necessarily the causal variants themselves. Choice B is an oversimplification; while tag SNPs often have a sufficient MAF, their primary selection criterion is how well they represent other SNPs. Choice D is incorrect; tag SNPs are chosen based on LD patterns, not necessarily on evolutionary conservation or known function.
Question 16
The human reference genome sequence (e.g., GRCh38) shows the allele 'T' at a chromosomal position known to be a common C/T SNP. In a global population survey, the 'C' allele is found to have a frequency of 75%. Which statement accurately explains this situation?
- The reference genome always contains the ancestral allele, and 'T' must therefore be the original human allele.
- This indicates a systematic error in either the reference sequence or the population frequency data.
- The 'T' allele must be the one associated with normal function, and the 'C' allele is pathogenic.
- The reference genome was constructed from individuals who happened to carry the minor 'T' allele at this position. (correct answer)
Explanation: When you encounter questions about reference genomes and allele frequencies, remember that reference sequences are constructed from specific individuals, not from population-wide consensus data. This creates situations where the reference allele may not be the most common variant in the global population.
The correct explanation is D: the reference genome contains the 'T' allele simply because the individuals whose DNA was used to construct that particular genomic region happened to carry 'T' at this position. Reference genomes are mosaics built from multiple donors, and each position reflects whatever allele those specific donors carried, regardless of global frequency patterns.
Option A is incorrect because reference genomes don't necessarily contain ancestral alleles—they contain whatever alleles the donor individuals had. Ancestral state determination requires separate evolutionary analysis. Option B wrongly assumes there must be an error when reference and population data don't align. This misunderstanding stems from thinking the reference should represent the most common allele, which isn't how reference genomes are constructed. Option C makes an unfounded assumption about functional significance. Allele frequency doesn't directly indicate pathogenicity—many common variants are neutral, and some rare variants are actually beneficial.
This scenario illustrates why modern genomics increasingly uses population-specific reference panels rather than relying solely on single reference sequences. For genetics exams, remember that reference genomes are historical snapshots from specific individuals, not population summaries. Always distinguish between reference sequences, population frequencies, ancestral states, and functional significance—these are independent concepts that students often conflate.
Question 17
A research group aims to identify novel, previously unknown genetic variants associated with a rare, complex disease in a small cohort of 200 patients. Which of the following technologies would be most appropriate for this discovery-oriented goal?
- A genome-wide SNP array that genotypes 700,000 common tag SNPs.
- Targeted Sanger sequencing of five well-established candidate genes.
- Whole-genome sequencing (WGS) to capture all types of genetic variation. (correct answer)
- An allele-specific PCR assay for a single, high-risk SNP.
Explanation: The correct answer is C. The goal is to discover novel variants, including rare SNPs, indels, and structural variants, without prior hypotheses about specific genes. Whole-genome sequencing (WGS) is the most comprehensive approach, as it reads the entire genome and can identify all types of variation, common and rare. Choice A is suboptimal because a SNP array is designed to assay common pre-selected variants and will miss rare or novel SNPs, which are often implicated in rare diseases. Choice B is a hypothesis-driven approach and is not suitable for discovering variants in unknown genes. Choice D is a screening tool for a single known variant and cannot be used for discovery.
Question 18
Consider two genomic regions of equal physical length (100 kb). Region A has a high meiotic recombination rate, while Region B is a recombination 'coldspot' with a very low rate. Assuming both regions have a similar underlying SNP density, which statement accurately compares their haplotype structures?
- Region A will have more distinct haplotypes, each spanning a shorter genetic distance, than Region B. (correct answer)
- Region B will have more distinct haplotypes due to the accumulation of mutations that are not shuffled.
- Both regions will have an identical number of haplotypes since they have the same SNP density.
- Region A will exhibit stronger linkage disequilibrium across the 100 kb distance than Region B.
Explanation: The correct answer is A. Recombination acts to shuffle existing genetic variation. In a region with a high recombination rate (Region A), associations between SNPs are broken down frequently. This creates a large number of different combinations of alleles, or distinct haplotypes, but these haplotypes are short because the correlation (LD) decays quickly with distance. In contrast, in a recombination coldspot (Region B), SNPs are rarely separated, so they are inherited together in large blocks. This results in fewer, longer haplotypes and strong LD across the region. Therefore, Region A has more, shorter haplotypes. Choice B is incorrect; low recombination leads to fewer haplotypes. Choice C is incorrect because recombination, not just SNP density, determines haplotype diversity. Choice D is incorrect; high recombination leads to weaker, not stronger, LD.
Question 19
A research group aims to identify novel, previously unknown genetic variants associated with a rare, complex disease in a small cohort of 200 patients. Which of the following technologies would be most appropriate for this discovery-oriented goal?
- A genome-wide SNP array that genotypes 700,000 common tag SNPs.
- Targeted Sanger sequencing of five well-established candidate genes.
- Whole-genome sequencing (WGS) to capture all types of genetic variation. (correct answer)
- An allele-specific PCR assay for a single, high-risk SNP.
Explanation: The correct answer is C. The goal is to discover novel variants, including rare SNPs, indels, and structural variants, without prior hypotheses about specific genes. Whole-genome sequencing (WGS) is the most comprehensive approach, as it reads the entire genome and can identify all types of variation, common and rare. Choice A is suboptimal because a SNP array is designed to assay common pre-selected variants and will miss rare or novel SNPs, which are often implicated in rare diseases. Choice B is a hypothesis-driven approach and is not suitable for discovering variants in unknown genes. Choice D is a screening tool for a single known variant and cannot be used for discovery.
Question 20
A SNP (rs7903146) in the TCF7L2 gene is one of the strongest common variant risk factors for type 2 diabetes. The risk allele 'T' has a frequency of approximately 30% in European populations. A person who is homozygous T/T has approximately double the risk of developing the disease compared to a person who is homozygous C/C. This mode of inheritance is best described as:
- Recessive, because the disease only manifests with two copies of the risk allele.
- Dominant, because heterozygotes (C/T) have the same risk as T/T homozygotes.
- Additive, because each copy of the 'T' allele contributes a quantifiable amount to the overall risk. (correct answer)
- X-linked, because the gene is known to interact with hormonal pathways.
Explanation: The correct answer is C. For most common variants associated with complex diseases, the effects are additive (or co-dominant). This means that individuals with one copy of the risk allele (heterozygotes) have a risk intermediate between the two homozygote groups. In this case, C/C is baseline risk (1x), C/T would have an intermediate risk (~1.5x), and T/T has the highest risk (~2x). Each 'T' allele adds a certain amount to the risk. Choice A is incorrect because heterozygotes also have an elevated risk, even if it's less than T/T homozygotes. Choice B is incorrect because the risk for C/T is lower than for T/T, so the T allele is not fully dominant. Choice D is incorrect; mode of inheritance is determined by risk patterns, not gene function, and TCF7L2 is on chromosome 10 (an autosome).