All questions
Question 1
A standard curve for a new qPCR primer set is generated using a 10-fold serial dilution of a plasmid DNA template. A plot of the Ct value versus the logarithm of the starting quantity yields a linear regression with a slope of -3.45. Given the formula: Efficiency = (10^(-1/slope)) - 1, what is the approximate amplification efficiency of this reaction?
- 85%
- 95% (correct answer)
- 105%
- 115%
Explanation: The amplification efficiency (E) is calculated from the slope of the standard curve. An ideal slope for 100% efficiency is -3.32. Using the provided formula:
- E = (10^(-1 / -3.45)) - 1
- E = (10^(1 / 3.45)) - 1
- E = (10^0.2898) - 1
- E = 1.949 - 1 = 0.949
- Expressed as a percentage, the efficiency is 0.949 * 100% ≈ 95%. This falls within the acceptable range for a reliable qPCR assay (90-110%).
Question 2
In a qPCR assay with 100% amplification efficiency, the Ct value for gene X is 19 in Sample 1 and 22 in Sample 2. Based solely on this information, which conclusion can be drawn about the initial quantity of the target sequence for gene X?
- Sample 1 contains 3 times more target sequence than Sample 2.
- Sample 2 contains 3 times more target sequence than Sample 1.
- Sample 1 contains 8 times more target sequence than Sample 2. (correct answer)
- Sample 2 contains 8 times more target sequence than Sample 1.
Explanation: The Ct (cycle threshold) value is inversely proportional to the logarithm of the initial amount of target template. At 100% efficiency, the amount of product doubles each cycle. A difference of 1 cycle (ΔCt = 1) therefore represents a 2-fold difference in initial template. Here, the ΔCt is 22 - 19 = 3. This corresponds to a 2^3 = 8-fold difference in starting material. Since a lower Ct value indicates a higher initial amount of template, Sample 1 (Ct=19) has 8 times more target sequence than Sample 2 (Ct=22).
Question 3
A clinical lab is developing a qPCR-based test to detect a specific pathogenic bacterial species in patient samples. The samples often contain other non-pathogenic, but genetically similar, commensal bacteria. Why is a TaqMan probe-based assay generally superior to a SYBR Green assay for this application?
- TaqMan assays are less expensive due to simpler primer design requirements.
- TaqMan assays provide an additional layer of specificity through the sequence-specific probe, reducing false positives. (correct answer)
- SYBR Green is unable to quantify DNA concentrations, it can only indicate presence or absence.
- TaqMan assays allow for post-run melt curve analysis to verify product specificity.
Explanation: The main challenge is distinguishing the pathogenic target from closely related non-targets. SYBR Green binds to any double-stranded DNA, so if the primers non-specifically amplify DNA from commensal bacteria, it will generate a signal, leading to a false positive. A TaqMan assay requires three oligonucleotides to be specific: two primers AND a probe that must bind between them. This third level of sequence recognition significantly increases the specificity of the assay and minimizes the risk of false positives from off-target amplification.
Question 4
The 5' to 3' exonuclease activity of Taq polymerase is essential for the function of which type of qPCR assay?
- SYBR Green-based assays, by degrading non-specific products.
- TaqMan probe-based assays, by cleaving the probe to release the reporter dye. (correct answer)
- Melt curve analysis, by ensuring amplicons dissociate at specific temperatures.
- Reverse transcription qPCR (RT-qPCR), by synthesizing cDNA from the RNA template.
Explanation: TaqMan assays use a sequence-specific probe with a reporter dye on the 5' end and a quencher on the 3' end. The probe binds to the target DNA between the primer sites. As Taq polymerase extends from the primer, its 5'->3' exonuclease activity degrades the probe, separating the reporter from the quencher and allowing fluorescence. This activity is not required for SYBR Green (A), which binds dsDNA directly. It is not involved in melt curve analysis (C). Reverse transcription (D) is performed by reverse transcriptase, not Taq polymerase.
Question 5
The ΔΔCt method for relative quantification in qPCR is based on several assumptions. Which of the following is the most critical assumption, the violation of which would render the 2^-(ΔΔCt) calculation mathematically invalid?
- The Ct values for the reference gene must be lower than the Ct values for the target gene.
- The amount of starting template RNA is the same for all samples being compared.
- The expression of the reference gene is constant across all experimental conditions and tissues.
- The amplification efficiencies of the target and reference genes are approximately equal and near 100%. (correct answer)
Explanation: When analyzing qPCR quantification methods, you need to understand what makes the mathematical foundation valid. The ΔΔCt method uses the formula 2−(ΔΔCt) to calculate relative gene expression, but this formula is only mathematically sound under specific conditions.
The correct answer is D because the ΔΔCt calculation assumes that both genes double with each PCR cycle (100% efficiency). If amplification efficiencies differ significantly between your target and reference genes, the 2−(ΔΔCt) formula becomes mathematically meaningless. For example, if your target gene amplifies at 90% efficiency while your reference amplifies at 100%, the fold-change calculations will be systematically skewed, making your results unreliable.
Here's why the other options, while important for accurate results, don't invalidate the mathematical calculation itself: A is incorrect because Ct values can be in any order - the math still works regardless of which gene amplifies first. B is wrong because unequal starting amounts affect accuracy but don't break the mathematical relationship - you're still measuring relative changes. C represents a biological assumption that affects interpretation quality, but even with variable reference gene expression, the ΔΔCt calculation remains mathematically valid.
The key distinction is between assumptions that affect accuracy versus those that make the calculation mathematically invalid. Only severely mismatched amplification efficiencies actually break the mathematical foundation of the method.
Study tip: Remember that qPCR method validity depends on the mathematical assumptions built into the formula, not just experimental best practices.
Question 6
When comparing gene expression levels across different RNA-seq samples, Transcripts Per Million (TPM) is often preferred over Fragments Per Kilobase of transcript per Million mapped reads (FPKM). What is the primary advantage of TPM that facilitates more accurate between-sample comparisons?
- TPM normalizes for gene length first, making the subsequent library size normalization more consistent across samples. (correct answer)
- TPM is an absolute quantification method, whereas FPKM provides only relative quantification.
- TPM incorporates a correction for PCR duplicates introduced during library preparation, unlike FPKM.
- TPM values are not dependent on the length of the gene being measured, simplifying calculations.
Explanation: The key difference between FPKM and TPM lies in the order of normalization. FPKM normalizes read counts to gene length first, then to sequencing depth. This can lead to situations where the sum of FPKM values differs between samples, complicating direct comparisons. TPM normalizes for gene length first, then normalizes for sequencing depth in a way that the sum of all TPM values in each sample is the same (1 million). This scaling factor makes the expression values more directly comparable across different libraries.
Question 7
A research group is studying the response of human macrophages to bacterial infection. Their primary hypothesis is that the infection specifically alters the expression of a small, well-defined panel of 12 cytokine genes. They have a large number of samples to process (over 100) and a moderate budget. Their goal is to precisely quantify the expression changes for only these 12 genes across all samples.
Given the research goal and constraints described in the passage, which gene expression measurement technique is the most appropriate choice?
- RNA-seq, because it provides a comprehensive, transcriptome-wide view of gene expression changes.
- Quantitative PCR (qPCR), because it is highly sensitive and cost-effective for interrogating a small number of target genes. (correct answer)
- Microarray analysis, because it can measure thousands of genes simultaneously with high throughput.
- Digital PCR (dPCR), because it provides absolute quantification without the need for a reference gene.
Explanation: The research goal is hypothesis-driven, focusing on a small, predefined set of 12 genes. qPCR is the ideal technique for this scenario because it is targeted, highly sensitive, quantitative, and relatively inexpensive on a per-sample basis when the number of target genes is small. RNA-seq (A) would be unnecessarily expensive and generate a large amount of unneeded data. Microarray (C) is also not ideal as it interrogates many more genes than necessary. While dPCR (D) is highly precise, it is generally more expensive and lower-throughput than qPCR, making it less suitable for analyzing 12 genes in over 100 samples on a moderate budget.
Question 8
A cancer researcher wants to detect and quantify a rare somatic mutation that is expected to be present in only 0.5% of the cells in a blood sample. Why is digital PCR (dPCR) often a more suitable technique than qPCR for this specific task?
- dPCR provides absolute quantification by partitioning the sample into thousands of reactions, allowing for precise counting of rare molecules. (correct answer)
- dPCR reactions have a much faster cycling time, allowing for higher throughput when screening for rare events.
- dPCR avoids the need for reverse transcription, making the workflow simpler and less prone to error when starting from RNA.
- qPCR is unable to detect targets at low concentrations, whereas dPCR has a fundamentally lower limit of detection.
Explanation: The key advantage of dPCR for rare allele detection is its ability to provide absolute quantification without a standard curve. By partitioning the sample into thousands of individual wells or droplets, such that each partition contains either zero or one (or very few) target molecules, the technique changes the problem from relative quantification (qPCR) to binary counting (positive vs. negative reactions). This partitioning allows for the very precise and sensitive detection of a small number of mutant molecules against a massive background of wild-type molecules, which is very difficult to achieve accurately with qPCR. qPCR can be very sensitive, but its quantification of rare events is less precise and subject to efficiency variations.
Question 9
A researcher aims to profile the complete transcriptome, including both coding mRNAs and long non-coding RNAs (lncRNAs), from a bacterial sample. Bacterial RNA lacks the polyadenylated (poly-A) tails common in eukaryotes. Which RNA-seq library preparation strategy should be used?
- Poly-A selection, to enrich for mature messenger RNA transcripts and increase sequencing depth for coding regions.
- Ribosomal RNA (rRNA) depletion, to remove the highly abundant rRNA molecules and sequence all other RNA classes. (correct answer)
- Small RNA sequencing, to specifically capture regulatory molecules like microRNAs and siRNAs.
- Exome capture, to selectively sequence only the protein-coding exons present in the sample.
Explanation: The goal is to sequence all RNA types except rRNA. Since bacterial mRNA lacks poly-A tails, poly-A selection (A) is not a viable method. Ribosomal RNA depletion (B) is the correct strategy; it uses probes to remove the highly abundant rRNA (which can constitute >90% of total RNA) and allows for the sequencing of all remaining transcripts, including mRNAs and lncRNAs. Small RNA sequencing (C) is designed for a different class of RNAs and would miss the longer transcripts of interest. Exome capture (D) is a DNA-based method, not an RNA-based one, and would not provide gene expression information.
Question 10
A researcher is studying gene expression in mouse brains following a specific diet. They choose Actb (β-actin) as their reference gene for qPCR. Later, literature review reveals that this specific diet is known to alter cytoskeletal organization in neurons, a process in which Actb is heavily involved. What is the most direct and significant consequence of using Actb as the reference gene in this study?
- The normalization will be invalid, leading to inaccurate and potentially misleading fold-change calculations for target genes. (correct answer)
- The qPCR reactions for Actb will likely fail due to the high variability in its expression levels.
- All calculated fold changes will be systematically underestimated, but the direction of change (up or down) will remain correct.
- The Ct values for the target genes will need to be adjusted using a statistical correction factor to account for Actb instability.
Explanation: The fundamental assumption of a reference gene in qPCR is that its expression is stable and unaffected by the experimental conditions. If the diet alters Actb expression, this assumption is violated. Using an unstable reference gene for normalization introduces significant error, making the calculated fold changes for the target genes unreliable. The direction and magnitude of the error depend on how Actb expression changes, so one cannot assume the error is systematic (C). While statistical corrections (D) exist, the primary consequence is the invalidation of the normalization itself. The reactions will not necessarily fail (B), but the data they produce will be misinterpreted.
Question 11
Following a SYBR Green qPCR run, the melt curve analysis for a target gene in one sample reveals two distinct peaks: one at the expected melting temperature (Tm) of 85°C and a smaller, secondary peak at 72°C. This secondary peak is also present in the no-template control (NTC). What is the most likely interpretation and the best course of action?
- The secondary peak indicates genomic DNA contamination; the RNA sample should be treated with DNase and the qPCR repeated.
- The secondary peak indicates the presence of a splice variant; the result is valid and reflects the sample's biology.
- The secondary peak indicates the formation of primer-dimers; results for this gene are only reliable if the sample Ct is well below the NTC Ct. (correct answer)
- The secondary peak indicates inefficient polymerase activity; a new master mix with a more robust enzyme should be used.
Explanation: A low-Tm peak (typically <80°C) that also appears in the no-template control (NTC) is the classic signature of primer-dimer formation. SYBR Green binds to this non-specific product, generating a signal. The presence of this artifact compromises quantification, especially in samples with low target expression where the Ct values might approach that of the NTC. While the data might be usable if the sample amplifies much earlier (e.g., Ct < 30) than the NTC, the best practice is to optimize the reaction (e.g., raise annealing temperature) or redesign primers to eliminate the artifact. Genomic DNA contamination (A) would likely have a different Tm and would not appear in the NTC. Splice variants (B) would not appear in the NTC. Inefficient polymerase (D) would likely result in poor or no amplification rather than a specific secondary peak.
Question 12
In designing an RNA-seq experiment, a primary consideration is sequencing depth. For a project whose main goal is to discover and quantify the expression of rare, low-abundance transcripts, which sequencing strategy would be most appropriate?
- Low depth (e.g., 5 million reads/sample) with long reads (250 bp) to improve transcript assembly.
- High depth (e.g., 100 million reads/sample) but using poly-A selection to reduce library complexity.
- Low depth (e.g., 5 million reads/sample) with paired-end reads to identify fusion transcripts.
- High depth (e.g., 100 million reads/sample) with short reads (50 bp) to maximize transcript sampling. (correct answer)
Explanation: When approaching RNA-seq experimental design, you need to match your sequencing strategy to your biological question. The key principle here is that detecting rare, low-abundance transcripts requires sufficient sampling depth to capture molecules that are present in very small quantities within the total RNA pool.
Answer D is correct because high sequencing depth (100 million reads) maximizes your chances of detecting rare transcripts by providing extensive sampling of the transcriptome. While 50 bp reads are shorter, they're perfectly adequate for quantification purposes, and the high read count compensates for any limitations in read length. More reads means better statistical power to detect transcripts expressed at low levels.
Answer A fails because 5 million reads provides insufficient sampling depth. Even with longer reads that help assembly, you simply won't capture enough rare transcript molecules to reliably detect and quantify them. Answer B is problematic because poly-A selection, while providing high depth, actually increases bias by only capturing polyadenylated mRNAs and missing other transcript types like long non-coding RNAs or processed transcripts. Answer C also suffers from inadequate depth (5 million reads), and while paired-end reads help identify fusion transcripts, this doesn't address the core challenge of detecting low-abundance molecules.
Remember this principle: for rare transcript discovery, depth trumps read length. You need enough reads to sample the rare molecules multiple times for reliable quantification. High-depth sequencing with shorter reads is more cost-effective than low-depth sequencing with longer reads when your goal is comprehensive transcript detection.
Question 13
A student performs an RT-qPCR experiment to measure mRNA levels of gene TP53. For one sample, they accidentally omit the reverse transcriptase enzyme during the cDNA synthesis step. The qPCR primers for TP53 are both located within Exon 4. Surprisingly, the qPCR still produces a product, with a Ct of 36. What is the most plausible explanation for this result?
- The Taq polymerase used in the qPCR step has sufficient reverse transcriptase activity to synthesize cDNA.
- This represents a baseline level of signal from spontaneous primer-dimer formation.
- The TP53 mRNA formed a secondary structure that could be directly amplified by Taq polymerase.
- The RNA sample was contaminated with a small amount of genomic DNA containing the TP53 gene. (correct answer)
Explanation: When you encounter RT-qPCR questions, always consider what could go wrong at each step and what alternative sources of DNA template might exist. RT-qPCR typically requires reverse transcriptase to convert mRNA into cDNA, which then serves as the template for qPCR amplification.
The key insight here is that a Ct of 36 indicates a weak but real signal - something is being amplified. Since both primers are within the same exon (Exon 4), they would successfully amplify either cDNA or genomic DNA, as introns are removed from mature mRNA anyway. The most logical explanation is that the RNA sample contained trace amounts of genomic DNA contamination, providing an alternative template for amplification despite the missing reverse transcriptase.
Let's examine why the other options don't work. Option A is incorrect because Taq polymerase lacks reverse transcriptase activity - it can only amplify existing DNA, not synthesize cDNA from RNA. Option B misunderstands primer-dimers, which occur when primers anneal to each other rather than to template, typically producing very short products with different melting characteristics than a legitimate gene product. Option C reflects a fundamental misunderstanding - Taq polymerase cannot directly amplify RNA regardless of secondary structure, as it requires a DNA template.
The high Ct value (36) supports contamination rather than intentional amplification, as contaminating DNA would be present in much smaller quantities than properly synthesized cDNA. Remember: whenever you see unexpected amplification in molecular biology experiments, always consider contamination as a primary explanation, especially DNA contamination in RNA samples.
Question 14
In a qPCR assay with 100% amplification efficiency, the Ct value for gene X is 19 in Sample 1 and 22 in Sample 2. Based solely on this information, which conclusion can be drawn about the initial quantity of the target sequence for gene X?
- Sample 1 contains 3 times more target sequence than Sample 2.
- Sample 2 contains 3 times more target sequence than Sample 1.
- Sample 1 contains 8 times more target sequence than Sample 2. (correct answer)
- Sample 2 contains 8 times more target sequence than Sample 1.
Explanation: The Ct (cycle threshold) value is inversely proportional to the logarithm of the initial amount of target template. At 100% efficiency, the amount of product doubles each cycle. A difference of 1 cycle (ΔCt = 1) therefore represents a 2-fold difference in initial template. Here, the ΔCt is 22 - 19 = 3. This corresponds to a 2^3 = 8-fold difference in starting material. Since a lower Ct value indicates a higher initial amount of template, Sample 1 (Ct=19) has 8 times more target sequence than Sample 2 (Ct=22).
Question 15
A researcher is studying gene expression in mouse brains following a specific diet. They choose Actb (β-actin) as their reference gene for qPCR. Later, literature review reveals that this specific diet is known to alter cytoskeletal organization in neurons, a process in which Actb is heavily involved. What is the most direct and significant consequence of using Actb as the reference gene in this study?
- The normalization will be invalid, leading to inaccurate and potentially misleading fold-change calculations for target genes. (correct answer)
- The qPCR reactions for Actb will likely fail due to the high variability in its expression levels.
- All calculated fold changes will be systematically underestimated, but the direction of change (up or down) will remain correct.
- The Ct values for the target genes will need to be adjusted using a statistical correction factor to account for Actb instability.
Explanation: The fundamental assumption of a reference gene in qPCR is that its expression is stable and unaffected by the experimental conditions. If the diet alters Actb expression, this assumption is violated. Using an unstable reference gene for normalization introduces significant error, making the calculated fold changes for the target genes unreliable. The direction and magnitude of the error depend on how Actb expression changes, so one cannot assume the error is systematic (C). While statistical corrections (D) exist, the primary consequence is the invalidation of the normalization itself. The reactions will not necessarily fail (B), but the data they produce will be misinterpreted.
Question 16
The 5' to 3' exonuclease activity of Taq polymerase is essential for the function of which type of qPCR assay?
- SYBR Green-based assays, by degrading non-specific products.
- TaqMan probe-based assays, by cleaving the probe to release the reporter dye. (correct answer)
- Melt curve analysis, by ensuring amplicons dissociate at specific temperatures.
- Reverse transcription qPCR (RT-qPCR), by synthesizing cDNA from the RNA template.
Explanation: TaqMan assays use a sequence-specific probe with a reporter dye on the 5' end and a quencher on the 3' end. The probe binds to the target DNA between the primer sites. As Taq polymerase extends from the primer, its 5'->3' exonuclease activity degrades the probe, separating the reporter from the quencher and allowing fluorescence. This activity is not required for SYBR Green (A), which binds dsDNA directly. It is not involved in melt curve analysis (C). Reverse transcription (D) is performed by reverse transcriptase, not Taq polymerase.
Question 17
In designing an RNA-seq experiment, a primary consideration is sequencing depth. For a project whose main goal is to discover and quantify the expression of rare, low-abundance transcripts, which sequencing strategy would be most appropriate?
- Low depth (e.g., 5 million reads/sample) with long reads (250 bp) to improve transcript assembly.
- High depth (e.g., 100 million reads/sample) but using poly-A selection to reduce library complexity.
- Low depth (e.g., 5 million reads/sample) with paired-end reads to identify fusion transcripts.
- High depth (e.g., 100 million reads/sample) with short reads (50 bp) to maximize transcript sampling. (correct answer)
Explanation: When approaching RNA-seq experimental design, you need to match your sequencing strategy to your biological question. The key principle here is that detecting rare, low-abundance transcripts requires sufficient sampling depth to capture molecules that are present in very small quantities within the total RNA pool.
Answer D is correct because high sequencing depth (100 million reads) maximizes your chances of detecting rare transcripts by providing extensive sampling of the transcriptome. While 50 bp reads are shorter, they're perfectly adequate for quantification purposes, and the high read count compensates for any limitations in read length. More reads means better statistical power to detect transcripts expressed at low levels.
Answer A fails because 5 million reads provides insufficient sampling depth. Even with longer reads that help assembly, you simply won't capture enough rare transcript molecules to reliably detect and quantify them. Answer B is problematic because poly-A selection, while providing high depth, actually increases bias by only capturing polyadenylated mRNAs and missing other transcript types like long non-coding RNAs or processed transcripts. Answer C also suffers from inadequate depth (5 million reads), and while paired-end reads help identify fusion transcripts, this doesn't address the core challenge of detecting low-abundance molecules.
Remember this principle: for rare transcript discovery, depth trumps read length. You need enough reads to sample the rare molecules multiple times for reliable quantification. High-depth sequencing with shorter reads is more cost-effective than low-depth sequencing with longer reads when your goal is comprehensive transcript detection.
Question 18
A cancer researcher wants to detect and quantify a rare somatic mutation that is expected to be present in only 0.5% of the cells in a blood sample. Why is digital PCR (dPCR) often a more suitable technique than qPCR for this specific task?
- dPCR provides absolute quantification by partitioning the sample into thousands of reactions, allowing for precise counting of rare molecules. (correct answer)
- dPCR reactions have a much faster cycling time, allowing for higher throughput when screening for rare events.
- dPCR avoids the need for reverse transcription, making the workflow simpler and less prone to error when starting from RNA.
- qPCR is unable to detect targets at low concentrations, whereas dPCR has a fundamentally lower limit of detection.
Explanation: The key advantage of dPCR for rare allele detection is its ability to provide absolute quantification without a standard curve. By partitioning the sample into thousands of individual wells or droplets, such that each partition contains either zero or one (or very few) target molecules, the technique changes the problem from relative quantification (qPCR) to binary counting (positive vs. negative reactions). This partitioning allows for the very precise and sensitive detection of a small number of mutant molecules against a massive background of wild-type molecules, which is very difficult to achieve accurately with qPCR. qPCR can be very sensitive, but its quantification of rare events is less precise and subject to efficiency variations.
Question 19
When comparing gene expression levels across different RNA-seq samples, Transcripts Per Million (TPM) is often preferred over Fragments Per Kilobase of transcript per Million mapped reads (FPKM). What is the primary advantage of TPM that facilitates more accurate between-sample comparisons?
- TPM normalizes for gene length first, making the subsequent library size normalization more consistent across samples. (correct answer)
- TPM is an absolute quantification method, whereas FPKM provides only relative quantification.
- TPM incorporates a correction for PCR duplicates introduced during library preparation, unlike FPKM.
- TPM values are not dependent on the length of the gene being measured, simplifying calculations.
Explanation: The key difference between FPKM and TPM lies in the order of normalization. FPKM normalizes read counts to gene length first, then to sequencing depth. This can lead to situations where the sum of FPKM values differs between samples, complicating direct comparisons. TPM normalizes for gene length first, then normalizes for sequencing depth in a way that the sum of all TPM values in each sample is the same (1 million). This scaling factor makes the expression values more directly comparable across different libraries.
Question 20
A standard curve for a new qPCR primer set is generated using a 10-fold serial dilution of a plasmid DNA template. A plot of the Ct value versus the logarithm of the starting quantity yields a linear regression with a slope of -3.45. Given the formula: Efficiency = (10^(-1/slope)) - 1, what is the approximate amplification efficiency of this reaction?
- 85%
- 95% (correct answer)
- 105%
- 115%
Explanation: The amplification efficiency (E) is calculated from the slope of the standard curve. An ideal slope for 100% efficiency is -3.32. Using the provided formula:
- E = (10^(-1 / -3.45)) - 1
- E = (10^(1 / 3.45)) - 1
- E = (10^0.2898) - 1
- E = 1.949 - 1 = 0.949
- Expressed as a percentage, the efficiency is 0.949 * 100% ≈ 95%. This falls within the acceptable range for a reliable qPCR assay (90-110%).