Historical Context & Motivation
For most of human history, people made claims about the world in absolute terms — something was either true or false, certain or impossible. It wasn't until the development of probability theory and modern statistics that thinkers began to formalize how we talk about uncertainty. The story of statistical language is really the story of how humans learned to be honest about what they know — and what they don't.
Today, the ability to communicate conclusions with appropriate uncertainty is more important than ever. Headlines routinely claim that studies "prove" things, when the data actually only suggest them. How do you write a conclusion that's honest about what the data shows without overstating or understating the evidence? That is the question this lesson addresses.
Core Principles of Uncertainty Language
When you draw a conclusion from data, you are making an inference — a statement that goes beyond the raw numbers to say something meaningful about the world. Because data always involves variability, sampling error, and limitations, your conclusion should reflect how confident you are. This is where uncertainty language comes in: a set of carefully chosen words and phrases that signal the strength of the evidence behind your claim.
Correlation ≠ Causation
Sample vs. Population
Strength of Evidence
Context Matters
The Uncertainty Language Spectrum
Uncertainty language exists on a spectrum from very cautious to very confident. The diagram below organizes common phrases along that spectrum and connects them to the type of evidence that justifies each level of confidence. Study it carefully — you'll use these phrases throughout the rest of the lesson.
The key idea in the diagram is that the phrases you choose should be proportional to the strength of the evidence. A small survey of 20 students does not justify the same language as a controlled experiment with 2,000 participants. As the sample size grows, the study design improves, and the effect size becomes larger, you can move rightward on the spectrum toward more confident language — but you should almost never reach the word "proves."
The Mathematical Framework Behind Confidence
Uncertainty language isn't just a matter of taste — it's anchored in mathematical concepts you may already know. Two of the most important are confidence intervals and p-values. Understanding what these numbers mean helps you decide which uncertainty phrase to use in your conclusion.
When reporting a confidence interval, the appropriate language is: "We are 95% confident that the true mean lies between [lower bound] and [upper bound]." Notice that this does not say "there is a 95% probability" — the true mean is a fixed number, not a random variable. The confidence refers to the process, not the specific interval.
Matching Language to Evidence Strength
Now that you understand the math behind confidence, you need a practical guide for choosing the right words. The table below maps different statistical scenarios to the phrases that best communicate your findings. Use it as a reference whenever you write a statistical conclusion.
| Scenario | Appropriate Language | Example |
|---|---|---|
| Small sample, observational study | "It is possible…" / "The data hint…" | "Based on a survey of 15 students, it is possible that screen time is associated with lower sleep quality." |
| Moderate sample, correlation found | "Evidence suggests…" / "It is likely…" | "Evidence from 200 respondents suggests that students who exercise regularly tend to report higher focus in class." |
| Large sample, controlled experiment, p < 0.05 | "The data strongly support…" / "There is convincing evidence…" | "A randomized experiment with 1,500 patients provides convincing evidence that the new drug reduces symptoms compared to a placebo." |
| No significant result found, p > 0.05 | "There is insufficient evidence to conclude…" / "The data do not support…" | "Based on these results, there is insufficient evidence to conclude that the tutoring method improves test scores." |
| Confidence interval reported | "We are __% confident that…" | "We are 95% confident that the true average commute time for city residents is between 22 and 28 minutes." |
Worked Example: Writing a Statistical Conclusion
Let's walk through a realistic scenario step by step. Suppose a school administrator wants to know whether a new breakfast program improves students' math test scores. She conducts a study with 120 students randomly assigned to two groups: 60 who eat the school breakfast and 60 who don't. After 8 weeks, the breakfast group's average score is 78 and the non-breakfast group's average is 73. The p-value from a two-sample t-test is 0.02, and the 95% confidence interval for the difference in means is (1.8, 8.2).
Common Mistakes vs. Strong Conclusions
Even after learning the right phrases, students often fall into predictable traps. The table below compares weak or incorrect conclusion language with improved versions. Study each pair to develop a feel for what good statistical writing looks like.
| Weak / Incorrect | What's Wrong | Improved Version |
|---|---|---|
| "This proves that coffee improves grades." | Statistics never "prove" anything; also claims causation without specifying study type. | "Evidence suggests that coffee consumption is associated with higher grades in this sample." |
| "The data failed to show a relationship." | Implies data has a duty to "succeed"; absence of evidence ≠ evidence of absence. | "There is insufficient evidence to conclude that a relationship exists between the two variables." |
| "The p-value is 0.03, so there is a 3% chance the null is true." | Misinterprets the p-value — it measures the probability of the data, not the hypothesis. | "The result was statistically significant (p = 0.03), meaning data this extreme would be unlikely if the null hypothesis were true." |
| "Everyone should switch to this diet because our study showed weight loss." | Overgeneralizes from one sample; makes a recommendation beyond the scope of the data. | "Among the 80 participants in this study, those on the new diet lost an average of 4 lbs more, suggesting the diet may be effective for similar populations." |
Connecting to Advanced Statistical Communication
The uncertainty language skills you are building now form the foundation for more advanced statistical communication you'll encounter in AP Statistics, college research courses, and professional settings. The table below shows how the concepts from this lesson connect to their more advanced counterparts.
| What You Learn Now | Where It Leads |
|---|---|
| Choosing phrases like "evidence suggests" vs. "proves" | Writing formal research conclusions with effect sizes, confidence intervals, and Bayesian credible intervals |
| Distinguishing correlation from causation | Identifying confounding variables, designing controlled experiments, and evaluating causal inference methods |
| Interpreting p-values in plain language | Discussing Type I/II errors, statistical power, and the replication crisis in science |
| Noting limitations and sample scope | Writing full "limitations" sections in research papers and assessing external validity |
In college and professional settings, poor communication of statistical results can have real consequences. Medical researchers who overstate findings may lead to unsafe treatments. Journalists who confuse correlation with causation can mislead millions. By learning to communicate with appropriate uncertainty now, you're building a skill that matters far beyond your math class.
Practice Problems
Lesson Summary
Communicating statistical conclusions is about more than just reporting numbers — it requires choosing uncertainty language that honestly reflects the strength of your evidence. The words you choose exist on a spectrum: cautious phrases like "it is possible" suit small samples and observational data, moderate phrases like "evidence suggests" fit significant correlations and larger studies, and confident phrases like "convincing evidence" are reserved for well-designed experiments with significant results. Critically, you should never use the word "proves" in a statistical conclusion.
Every strong conclusion references the study type (observational vs. experimental), the sample size, and any limitations of the data. When a confidence interval or p-value is available, include it — but interpret it correctly. Remember: only randomized experiments support causal claims, and the distinction between correlation and causation is one of the most important ideas in all of statistics.