All questions
Question 1
A student is preparing a five-minute oral presentation on climate change for an advanced ESL class. She records herself and notices that she rushes through key statistical data, drops her voice at the end of every sentence regardless of whether it is a statement or a question, and pauses only between major topic sections rather than within them to allow listeners to process complex ideas.
Which of the following revisions to the student's delivery would MOST directly address the combined effect of her pacing, intonation, and pausing errors in extended speech?
- Slow down uniformly throughout the entire presentation so that every sentence receives equal time, add a brief pause after each period to signal sentence completion, and practice sustaining consistent volume so that statistical figures are not lost in transitions between sections.
- Increase vocal emphasis on statistical data, use rising intonation on questions rather than falling intonation on all sentence types, and insert micro-pauses after complex noun phrases and subordinate clauses rather than only at section boundaries. (correct answer)
- Memorize the presentation word for word to eliminate hesitation pauses, practice falling intonation on declarative sentences so the voice drop sounds intentional, and use a consistent medium pace throughout to prevent the rushes that occur specifically during data-heavy segments.
- Project a confident tone by keeping intonation relatively flat and consistent across sentence types, increase overall volume to ensure audibility, and reserve all pausing for moments when the speaker needs to recall the next point in the outline.
Explanation: When evaluating oral presentation delivery, you need to match each proposed revision to the specific errors identified — here, rushing through data, uniform falling intonation, and pausing only at section boundaries. The best revision must address all three problems precisely, not just partially or with new errors introduced.
B is the strongest choice because each element directly targets one identified problem. Increasing vocal emphasis on statistical data counters the rushing issue by naturally slowing the speaker down at critical moments rather than everywhere. Using rising intonation on questions corrects the inappropriate falling intonation pattern without overcorrecting declarative sentences. Inserting micro-pauses after complex noun phrases and subordinate clauses directly solves the within-sentence pausing deficit — the student's real gap — rather than just adding pauses between sentences.
A sounds reasonable but misses the mark in two ways: slowing down uniformly is inefficient and unnatural, and adding pauses only after periods doesn't solve the problem of listeners needing processing time within complex sentences. The revision is too blunt.
C introduces a new problem — memorization increases rigidity and can actually worsen delivery under pressure. It also accepts the falling intonation pattern as intentional rather than correcting the intonation error, and "consistent medium pace" ignores that data segments specifically need strategic slowing.
D makes the errors worse, not better. Flat intonation eliminates meaningful prosodic variation, and reserving all pausing for speaker recall turns a communication tool into a personal coping mechanism.
On exam questions like this, watch for answer choices that address the surface of a problem rather than its source — the correct answer will always target the underlying cause with precision.
Question 2
An advanced ESL speaker consistently places primary stress on the grammatical subject of every sentence, regardless of whether the subject represents new or given information. During a two-minute impromptu speech, a native-speaker listener reports that the speech 'sounds mechanical and makes it hard to follow what's new.' Which principle of English stress in extended speech best explains the listener's reaction?
- English stress rules require that content words always receive greater stress than function words, so stressing the subject — which is often a content word — should enhance, not impede, comprehension.
- In English extended speech, nuclear stress typically falls on the element that represents new or contrastive information rather than on a fixed grammatical position, so stressing given information (the subject) disrupts the listener's ability to track the information flow. (correct answer)
- English intonation units always begin with a stressed syllable, so placing stress on the subject is a natural prosodic convention that should align with the listener's expectations for sentence-initial prominence.
- The mechanical quality the listener perceives stems from the speaker's failure to vary sentence length, since English listeners rely on alternating short and long sentences to distinguish main ideas from supporting details.
Explanation: When you see a question about English prosody in extended speech, focus on how stress functions communicatively — not just grammatically. Native English speakers use stress to signal information structure: new or contrastive information gets emphasized, while given (already-known) information is typically de-stressed. This system helps listeners continuously track what's important and what's already established.
This is exactly why B is correct. In natural English, nuclear stress — the most prominent stress in an intonation unit — lands on new or contrastive information, regardless of its grammatical position. When the speaker mechanically stresses the subject every time, they're often emphasizing given information (subjects are frequently things already mentioned or assumed). The listener can't distinguish what's genuinely new, making the speech feel flat and hard to follow. That "mechanical" quality the listener describes is a direct symptom of misaligned information structure.
Choice A is tempting because content words do generally receive more stress than function words — that's a real rule. However, it doesn't mean stressing the subject is always appropriate. Grammatical position and information status are separate factors, and A ignores the crucial given/new distinction entirely.
Choice C misapplies a real phenomenon. While intonation units often have a stressed syllable, sentence-initial position doesn't automatically deserve primary nuclear stress. Conflating "some stress" with "maximum stress" is the trap here.
Choice D is a distraction that introduces sentence length as the culprit. Sentence length variation affects rhythm but is unrelated to the information-tracking problem described.
Study tip: Remember that English stress is discourse-driven, not just grammatically assigned — always ask yourself whether the stressed element is new, given, or contrastive.
Question 3
An advanced ESL speaker is delivering a five-minute persuasive speech. She uses correct word-level stress on every content word, but her instructor marks her down for 'flat, ineffective delivery that fails to guide the audience.' The speaker argues that since her word-stress is accurate, her pronunciation is adequate for extended speech. Which of the following best identifies the flaw in the speaker's reasoning?
- The speaker conflates word-level stress accuracy with discourse-level prosodic competence; extended speech also requires sentence-level stress, intonation contours, and rhythmic grouping that signal information structure across utterances — none of which are captured by word stress alone. (correct answer)
- The speaker is incorrect because word-level stress in English is not phonemically contrastive the way it is in tonal languages, meaning that accurate word stress contributes minimally to comprehensibility in any speech context, whether at the word or discourse level.
- The speaker overlooks the fact that persuasive speech specifically requires a wider pitch range than informational genres; therefore, intonation adjustments are only necessary in persuasion contexts and need not be considered when delivering academic or neutral extended speech.
- The speaker's argument would be valid if the audience were also non-native speakers, since research consistently shows that non-native listeners rely more heavily on word-level stress than on sentence-level intonation when processing extended speech for meaning.
Explanation: When a question asks you to evaluate a speaker's reasoning about their own pronunciation, you're being tested on your understanding of prosodic competence across different levels of language — specifically, the difference between word-level and discourse-level pronunciation features.
Word-level stress (knowing that PREsent differs from preSENT) is genuinely important, but it's only one layer of spoken English. Extended speech — especially persuasive speech — also demands sentence-level (nuclear) stress to highlight new information, intonation contours to signal questions, conclusions, or emphasis, and rhythmic grouping (chunking utterances into meaningful phrases). These features together guide the listener through the argument. A speaker who controls word stress but ignores these tools sounds flat and robotic because the audience receives no prosodic cues about what matters. Answer A correctly identifies this gap: the speaker conflates one narrow skill with the full system of prosodic competence required for effective extended speech.
Answer B is wrong because it misrepresents English phonology. Word stress is phonemically contrastive in English (e.g., record as noun vs. verb), and accurate word stress absolutely contributes to comprehensibility — this answer dismisses a real skill with a false claim. Answer C is wrong because it invents a false genre restriction; intonation and pitch variation are important across all extended speech, not only in persuasive contexts. Answer D is wrong because the research claim it cites is fabricated and the logic is flawed — audience background doesn't make the speaker's argument "valid," since prosodic features aid processing for all listeners.
A useful reminder: on questions about spoken language competence, always ask yourself whether the concept operates at the word level, sentence level, or discourse level — these are distinct, and confusing them is the most common trap.
Question 4
An ESL instructor evaluates two students on a three-minute oral argument. Student 1 speaks at approximately 160 words per minute with consistent rhythm, clear thought-group boundaries, and strategic pausing before key claims. Student 2 speaks at approximately 110 words per minute, pauses frequently but unpredictably, and elongates unstressed syllables to fill time. Both students are equally accurate in grammar and vocabulary. The instructor rates Student 1 significantly higher on 'extended speech clarity.'
Which of the following most precisely explains why Student 1's delivery is rated higher, given that Student 2 speaks more slowly — a feature often recommended for ESL learners?
- Slower speech rate alone does not guarantee clarity; Student 2's unpredictable pausing disrupts the listener's ability to anticipate thought-group boundaries, and elongating unstressed syllables distorts English stress-timed rhythm, both of which impede processing more than a faster but rhythmically coherent delivery. (correct answer)
- Student 1's higher rating reflects the fact that 160 words per minute is the documented optimal speaking rate for all formal English speech contexts, meaning Student 2 is simply speaking too slowly to be rated highly regardless of any other prosodic features present in the delivery.
- Student 1 is rated higher because native-speaker evaluators consistently prefer faster speech rates as a reliable proxy for fluency and confidence, regardless of the structural qualities of the delivery; the rating thus reflects an evaluator bias rather than a genuine difference in listener comprehensibility.
- Student 2's frequent pausing signals disfluency to the evaluator, and since disfluency in ESL speech is always caused by insufficient vocabulary retrieval speed, the lower rating reflects Student 2's weaker lexical access rather than any underlying prosodic or rhythmic shortcoming.
Explanation: When evaluating oral fluency, you need to think beyond surface-level features like speaking rate. This question tests whether you understand prosody — the rhythm, stress, and pacing patterns that make spoken English easy or difficult to process. A common misconception is that slower always means clearer, but that's only true when the slowness is structured.
English is a stress-timed language, meaning listeners subconsciously predict where emphasis and thought-group boundaries will fall. When a speaker delivers these cues consistently — even at a faster pace — the listener's brain can chunk and process information efficiently. Student 1 does exactly this: strategic pausing before key claims signals structure, and consistent rhythm preserves English's natural stress patterns. Answer A correctly identifies that Student 2's unpredictable pausing destroys anticipatory processing, and elongating unstressed syllables actively distorts the stress-timed rhythm listeners depend on. Slower speech without rhythmic coherence creates more cognitive load, not less — making A the precise, well-reasoned explanation.
Answer B is wrong because no single speaking rate is universally "optimal" for all formal contexts — this is an oversimplification that ignores prosodic quality entirely. Answer C misattributes the rating difference to evaluator bias rather than genuine comprehensibility factors, which contradicts the evidence in the passage and misrepresents how trained instructors assess delivery. Answer D makes an unfounded causal claim — that disfluency is always caused by slow vocabulary retrieval — when the passage clearly frames the issue as prosodic and rhythmic, not lexical.
Your takeaway: whenever a question contrasts two speakers with different rates, ask yourself how each speaker uses time, not just how much time they use.
Question 5
A student reads the following prepared statement aloud in class: 'Although the policy was initially controversial, most stakeholders eventually agreed that the long-term economic benefits outweighed the short-term social costs.' A rater notes: 'The concessive clause felt like the main point, and the main clause felt like an afterthought — the opposite of what you intended.'
Which prosodic adjustment would most directly correct the rater's perception while preserving the sentence's grammatical structure?
- Reorder the sentence so the main clause comes first, since listeners process information linearly and will always assign greater importance to whichever clause they hear first, regardless of how stress or intonation are distributed across the utterance.
- Stress every content word in both clauses equally so that neither clause receives disproportionate prominence, allowing the grammatical subordination marker 'although' to carry the full communicative weight of signaling the hierarchical relationship between the two clauses.
- Insert a longer pause after 'controversial' to separate the two clauses more distinctly; pausing is the primary prosodic mechanism by which listeners distinguish main clauses from subordinate clauses in spoken English, and a clear boundary pause will redirect prominence toward whichever clause follows it.
- Place the intonation peak (nuclear stress and highest pitch point) on a key word in the main clause — such as 'outweighed' or 'benefits' — and use a lower, flatter pitch and faster tempo through the concessive clause to mark it as background information. (correct answer)
Explanation: When speaking aloud, prosody — your use of pitch, stress, tempo, and pause — tells listeners which ideas are central and which are background. Grammatical structure alone doesn't communicate hierarchy in speech the way it does on the page. This question tests whether you understand how to use prosodic tools intentionally to match your spoken emphasis to your intended meaning.
The student's problem is that the concessive clause ("Although the policy was initially controversial") received too much prominence, making it sound like the main point. The fix is to use nuclear stress and pitch strategically. D is correct because placing the intonation peak — the highest pitch point and heaviest stress — on a key word like "outweighed" or "benefits" in the main clause signals to listeners that this is the core message. Simultaneously, delivering the "although" clause with lower pitch and faster tempo marks it as subordinate, background information. This is exactly how skilled speakers use prosody to mirror grammatical hierarchy.
A is wrong because reordering the sentence changes the grammatical structure, which the question explicitly prohibits. It also overstates a false rule — word order alone doesn't determine perceived importance; prosody does. B is wrong because flattening stress across both clauses doesn't resolve the hierarchy problem; it creates monotony and leaves listeners without clear guidance. The word "although" alone cannot carry that communicative burden. C is wrong because pausing is not the primary mechanism for signaling clause hierarchy — pitch and stress are. A pause merely separates clauses; it doesn't direct prominence toward either one.
Remember: in spoken English, pitch peak = main point. Use high stress and slow tempo for what matters most, and flatten your delivery for background information.
Question 6
An advanced ESL speaker consistently produces correct segmental sounds (individual vowels and consonants) but loses points on extended-speech evaluations for 'failure to signal discourse structure prosodically.' A peer suggests that the speaker simply needs to 'speak louder overall' to fix the problem. Why is the peer's advice insufficient, and what would more accurately address the evaluator's concern?
- The peer's advice is insufficient because evaluators of extended speech do not assess acoustic features such as volume or pitch; they assess only grammatical accuracy and lexical diversity, meaning the speaker should focus on expanding vocabulary range and reducing syntactic errors rather than modifying any aspect of delivery.
- The peer's advice is insufficient because loudness is a feature of individual phoneme production rather than of suprasegmental organization; since the speaker's segmental accuracy is already high, the only remaining issue is increasing the total number of words produced per minute to match native-speaker output norms for extended discourse.
- The peer's advice is insufficient because volume alone does not encode the pitch movements, rhythmic grouping, and stress contrasts that listeners use to identify topic shifts, list boundaries, and key arguments; the speaker needs to develop control of pitch range, nuclear stress placement, and intonation unit structure. (correct answer)
- The peer's advice is insufficient primarily because speaking louder strains the vocal cords and reduces speech rate, which indirectly causes the speaker to pause more frequently; while this may inadvertently improve thought-group chunking slightly, it does not address the underlying prosodic deficit the evaluator has identified.
Explanation: Whenever you see a question about pronunciation evaluation, distinguish between segmental features (individual sounds) and suprasegmental/prosodic features (pitch, stress, rhythm, intonation across stretches of speech). This question tests whether you understand that "sounding clear" in extended discourse requires far more than correct individual sounds or simple volume adjustments.
The evaluator's specific complaint — failure to signal discourse structure prosodically — means the speaker isn't using pitch movement, nuclear stress, and intonation phrasing to guide listeners through the organization of ideas: where topics shift, which words carry key information, where lists begin and end. Choice C correctly identifies this gap. Volume alone is a single acoustic dimension; it cannot substitute for the varied pitch contours, rhythmic grouping into thought chunks, and strategic stress placement that listeners rely on to follow extended speech. Addressing this requires targeted work on intonation unit structure and pitch range control.
Choice A is wrong because it dismisses acoustic/prosodic features entirely, claiming evaluators only care about grammar and vocabulary — a false claim that ignores the well-established role of prosody in discourse organization assessments. Choice B is wrong on two counts: it mischaracterizes loudness as a segmental feature (it isn't), and it incorrectly claims speaking faster would solve the problem — speech rate has no direct relationship to prosodic signaling of discourse structure. Choice D is technically creative but fundamentally misleading; it treats an accidental side effect (more pausing) as a partial fix while ignoring that the real deficit — pitch and stress patterning — remains completely unaddressed.
Your study tip: when you see "prosodic" or "suprasegmental" in a question, think pitch, stress, rhythm, and intonation — never just volume or speed.
Question 7
During a two-minute oral summary, an advanced ESL speaker uses a high-rising terminal (HRT) — a rising intonation contour — on every declarative statement. A native-speaker evaluator comments that the speech 'sounds uncertain and makes it difficult to tell when the speaker has finished a point.' Which explanation most accurately accounts for the evaluator's dual perception of uncertainty AND structural ambiguity?
- HRT on declaratives is pragmatically associated with seeking listener confirmation or signaling incompleteness; its consistent use converts statements into apparent checks for understanding, simultaneously projecting speaker uncertainty and preventing the listener from perceiving any statement as concluded. (correct answer)
- HRT is a feature found exclusively in Australian and New Zealand English varieties and is therefore unfamiliar to evaluators trained in American English norms; the evaluator's dual perception of uncertainty and ambiguity thus reflects a dialect bias rather than a genuine prosodic error in the speaker's extended speech.
- Rising intonation on declaratives is categorically incorrect in formal academic English because it violates the universal rule that all statements must end with falling intonation; the evaluator is simply enforcing this prescriptive standard, and any rising contour on a statement will always produce both uncertainty and structural confusion.
- The evaluator perceives uncertainty because HRT raises fundamental frequency at utterance end, which requires measurably greater vocal effort; listeners consistently associate any sustained increase in pitch range or vocal intensity with emotional distress or lack of confidence on the part of the speaker.
Explanation: When analyzing prosody questions like this one, focus on the pragmatic function of intonation patterns — what a rising or falling contour communicates to listeners, not just how it sounds physically.
The high-rising terminal (HRT) is a well-documented prosodic feature in which a speaker uses rising intonation on declarative (statement) sentences. In conversational English, rising intonation on a statement typically signals one of two things: the speaker is seeking listener confirmation ("You understand so far?") or the speaker hasn't finished their thought yet. When a speaker uses HRT consistently on every statement, both of these signals fire simultaneously and repeatedly. The evaluator therefore hears each statement as an unfinished check for understanding rather than a completed, confident assertion — producing the exact dual perception described: uncertainty and structural ambiguity. That makes A the correct answer, because it accurately names the pragmatic mechanism (seeking confirmation / signaling incompleteness) and explains how consistent use creates both problems at once.
B is wrong because HRT is not exclusive to Australian and New Zealand English — it appears in many varieties, including American English — and framing the evaluator's response as mere dialect bias sidesteps the real pragmatic issue entirely. C is wrong because there is no universal prescriptive rule requiring falling intonation on all statements; HRT can function appropriately in specific contexts, so calling it "categorically incorrect" is an overgeneralization. D is wrong because it misidentifies the mechanism — listener perception of uncertainty from HRT is not driven by increased vocal effort or pitch range associated with emotional distress; it's driven by the pragmatic meaning listeners assign to rising contours.
Your study tip: on prosody questions, always ask what meaning does this intonation pattern signal to listeners — function matters more than acoustics.
Question 8
A language instructor plays two recordings of the same paragraph for the class. In Recording A, the speaker pauses only at punctuation marks. In Recording B, the speaker pauses at punctuation marks AND at the boundaries of long prepositional phrases, relative clauses, and conditional clauses. The instructor asks which recording better exemplifies skilled extended speech and why.
Which of the following responses most accurately explains why one recording demonstrates more advanced extended-speech competence than the other?
- Recording A is superior because pausing only at punctuation marks reflects the written structure of the text, helping the listener follow the author's intended organization without introducing additional breaks that could fragment the rhetorical flow.
- Recording A is superior because experienced listeners of English rely on stress and intonation — not pausing — to signal syntactic boundaries; additional pauses at clause edges interrupt natural prosodic contours and can mislead listeners about where sentence-level emphasis falls.
- Recording B is superior primarily because more frequent pausing gives the speaker additional processing time, which reduces disfluencies such as 'um' and 'uh,' lowers the speaker's cognitive burden, and makes the overall delivery sound more fluent and prepared.
- Recording B is superior because inserting pauses at syntactic boundaries within long clauses reflects how proficient speakers chunk speech into listener-friendly thought groups, aiding real-time parsing and reducing cognitive load during complex structures. (correct answer)
Explanation: When evaluating spoken language competence, think about how skilled speakers organize speech for listeners, not just how they mirror written text. The key concept here is thought groups — the way proficient speakers chunk continuous speech into meaningful units that help listeners parse complex grammar in real time.
Recording B demonstrates stronger extended-speech competence precisely because it reflects this chunking behavior. When a speaker inserts brief pauses at syntactic boundaries — before relative clauses, after long prepositional phrases, at conditional clause edges — those pauses signal to the listener where one meaningful unit ends and another begins. This reduces the cognitive load of processing dense, multi-clause sentences on the fly. Listeners don't have to hold an entire complex sentence in working memory before making sense of it; the pauses do that organizational work for them. This is the reasoning behind D, which correctly identifies thought-group chunking and listener-friendly parsing as marks of advanced delivery.
Choice A misunderstands the relationship between written and spoken language. Punctuation guides the eye of a reader; skilled speech is governed by different prosodic principles. Choice B contains a partial truth — stress and intonation do carry syntactic information — but concludes incorrectly that clause-boundary pauses mislead listeners. In reality, those pauses reinforce prosodic signals rather than contradict them. Choice C focuses on the speaker's internal benefit (reduced disfluency, lower cognitive burden) rather than the communicative effect on the listener, which is what defines skilled extended speech.
As a study tip: on questions about spoken fluency and prosody, always ask yourself whether an answer centers on the listener's comprehension experience — that perspective usually points toward the correct response.
Question 9
During a graduate seminar, a non-native English speaker is asked to summarize a 10-page article in approximately three minutes. A classmate notes: 'I understood your main point, but I kept losing the thread when you explained the methodology. Your sentences there felt like one long blur, and I couldn't tell which part was most important.'
The classmate's feedback most likely reflects a breakdown in which combination of spoken-language features during the student's extended speech?
- Inadequate vocabulary range and overuse of passive constructions, which obscure agency and make the methodology section feel imprecise, causing the listener to work harder to reconstruct who did what and why it matters.
- Insufficient nuclear stress placement and absence of thought-group chunking, causing the listener to receive the methodology as an undifferentiated stream without clear informational hierarchy. (correct answer)
- Excessive use of discourse markers such as 'furthermore' and 'however,' which signal contrast and addition so frequently that the listener cannot distinguish primary from secondary information or identify where one idea ends and another begins.
- Overreliance on a monotone pitch range combined with speaking too quietly, so that the listener must concentrate on audibility rather than on processing the logical structure of the methodology section.
Explanation: When a listener says speech felt like "one long blur" with no sense of what was most important, that's a clue about prosodic and chunking features — the tools speakers use to organize spoken information through rhythm, stress, and pausing rather than through word choice or grammar.
Answer B correctly identifies the two features most directly responsible for this breakdown. Nuclear stress is the emphasis a speaker places on the most informationally significant word in a phrase — it tells listeners "this is the key idea." Thought-group chunking is the way pauses and intonation boundaries divide continuous speech into digestible units. Without these, even grammatically perfect sentences collapse into an undifferentiated stream, which is exactly what the classmate described.
Answer A points to vocabulary and passive constructions — real academic challenges, but the classmate said they understood the main point, suggesting vocabulary wasn't the core issue. The problem was delivery structure, not word-level precision.
Answer C focuses on overuse of discourse markers. In reality, markers like "furthermore" and "however" help listeners track logical relationships; their absence is more likely to blur structure than their presence.
Answer D attributes the problem to monotone pitch and low volume. While these affect engagement and audibility, the classmate's specific complaint — losing the "thread" and unable to identify importance — points to a hierarchy problem, not a volume or energy problem.
Study tip: On questions about spoken language breakdowns, map the listener's complaint to a specific linguistic system. "Can't tell what's important" → prosody/stress. "Can't follow logic" → discourse markers. "Can't hear" → volume/articulation. Matching symptoms to systems will guide you to the right answer.
Question 10
A transcript excerpt from a student's recorded academic presentation reads as follows (capital letters indicate where the student placed the heaviest stress): 'The RESULTS showed that THE participants who RECEIVED the treatment DID improve significantly, but THE control group did NOT show any CHANGE.' A rater notes that the stress pattern 'obscures the contrast the speaker intends to draw.'
To make the intended contrast clearest in extended speech, on which words should the speaker place primary nuclear stress, and why?
- On 'RESULTS' and 'CONTROL,' because these are the two topic nouns that anchor each clause; stressing topic nouns ensures the listener can identify the grammatical subject of each assertion and track the overall organization of the argument.
- On 'IMPROVE' and 'CHANGE,' because verbs carry the core predication and always receive nuclear stress in English declarative sentences, making them the natural prosodic targets regardless of the surrounding discourse context.
- On 'TREATMENT' and 'CONTROL' in the first and second clauses respectively, because these are the contrastive elements — the two groups being directly compared — and contrastive stress signals to the listener that these items are being set against each other. (correct answer)
- On 'SIGNIFICANTLY' in the first clause and 'NOT' in the second clause, because degree adverbs and negation markers are the most semantically dense elements in each clause and therefore require the strongest prosodic prominence to convey the speaker's evaluative stance.
Explanation: When you hear a question about stress and intonation in academic speech, think about why a speaker stresses certain words. In English, nuclear stress — the heaviest emphasis in a phrase — typically falls on the element that carries the most communicative importance in context. When a speaker is drawing a contrast, the stressed words should be the items being directly compared, because that's what tells the listener what is different from what.
In this passage, the speaker is contrasting two groups: those who received the treatment versus the control group. The contrast isn't about whether improvement happened in the abstract — it's about which group showed which outcome. Stressing "TREATMENT" and "CONTROL" highlights the two opposing subjects of comparison, signaling to the listener that these items are being set against each other. That's why C is correct.
Choice A is tempting because "RESULTS" and "CONTROL" sound important, but "RESULTS" anchors the whole sentence rather than one side of a contrast — stressing it doesn't help the listener hear the comparison. Stressing topic nouns for organizational tracking is a different goal than signaling contrast.
Choice B applies a misleading rule. Verbs don't automatically receive nuclear stress — context determines stress placement. Stressing "IMPROVE" and "CHANGE" would highlight the actions rather than the groups performing them, blurring the intended comparison.
Choice D focuses on degree and negation, which matter emotionally but don't anchor the logical contrast. "NOT" emphasizes the absence of change, but without stressing "CONTROL," the listener doesn't know clearly who didn't change.
Study tip: On prosody questions, always ask: What is the speaker trying to contrast or highlight? Stress should land on the word that answers that question — not just the "most important" word in isolation.