COLLEGE STATISTICS • PROBLEM-SOLVING & STATISTICAL REASONING

Writing Statistical Conclusions — Writing Conclusions with Proper Statistical Language

Master the precise language that transforms raw statistical output into defensible, publishable conclusions.

Historical Context & Motivation

The language used to communicate statistical results has evolved alongside inferential statistics itself, shaped by decades of debate among mathematicians, philosophers, and applied researchers. Early practitioners of probability theory, such as Pierre-Simon Laplace, wrote in conversational prose that blurred the line between subjective belief and objective evidence. As the discipline matured during the twentieth century, the need for standardized, unambiguous language became acute — misinterpretations of p-values, confidence intervals, and hypothesis tests led to costly errors in medicine, public policy, and the social sciences. The history of statistical conclusions is therefore inseparable from the history of the methods they describe, and understanding how this language crystallized helps modern students avoid the very pitfalls that prompted its formalization.

1925
Fisher's Framework for Significance
Ronald A. Fisher published Statistical Methods for Research Workers, introducing the p-value as a continuous measure of evidence against a null hypothesis. Fisher's language — 'significant,' 'highly significant' — shaped how researchers reported results for decades, though he never intended rigid cutoffs.
1933
Neyman–Pearson Decision Theory
Jerzy Neyman and Egon Pearson formalized hypothesis testing as a decision procedure with explicit Type I and Type II error rates. Their framework introduced the language of 'rejecting' or 'failing to reject' a null hypothesis, terms still central to proper statistical conclusions today.
1978
APA Style Codifies Reporting Standards
The American Psychological Association began mandating specific formats for reporting test statistics, degrees of freedom, and p-values in empirical papers, creating a template that other disciplines gradually adopted.
2016
ASA Statement on P-Values
The American Statistical Association released a landmark statement clarifying that p-values do not measure the probability that a hypothesis is true and that conclusions should never rest on whether a p-value crosses a fixed threshold alone — underscoring the importance of careful language.
2019
"Retire Statistical Significance"
Over 800 statisticians signed an open letter in Nature calling for the abandonment of the phrase 'statistically significant.' The debate renewed attention to how conclusions are worded and the dangers of binary language.

This century-long arc reveals a persistent question: how do we translate the nuanced output of a statistical test into a conclusion that is both accurate and comprehensible? The answer lies in mastering a specific vocabulary — phrases like "sufficient evidence", "fail to reject", and "at the α = 0.05 significance level" — that carry precise statistical meaning and guard against overstatement. This lesson equips you with that vocabulary and the reasoning behind it.

Core Principles of Statistical Conclusions

A well-crafted statistical conclusion does far more than announce a result; it communicates the strength of the evidence, the context of the claim, and the limits of the inference. Several foundational principles govern the language you should use. Violating any one of them can transform a valid analysis into a misleading or indefensible claim. Before examining specific templates, internalize these core ideas — they are the guardrails that keep your conclusions honest.

1

Evidence, Not Proof

Statistical tests yield evidence for or against a hypothesis; they never prove anything. Conclusions must use evidential language: 'There is sufficient evidence to suggest…' rather than 'This proves that…'
2

Fail to Reject ≠ Accept

When p > α, we fail to reject the null hypothesis. We do not 'accept' it, because the absence of evidence is not evidence of absence — the test may simply lack power.
3

State the Hypothesis in Context

A conclusion must restate the claim in plain language tied to the real-world question, not merely declare 'Reject H₀.' Readers should understand what the finding means without consulting the hypothesis section.
4

Report the Significance Level

Always specify the chosen significance level (α) and the p-value or test statistic. A conclusion without these details is unverifiable and incomplete.
5

Scope and Causation Caveats

Observational studies support association claims; only randomized experiments support causation claims. The conclusion must match the study design.
KEY TAKEAWAY
Think of a statistical conclusion like a courtroom verdict. A jury delivers 'not guilty' — not 'innocent.' The distinction matters: 'not guilty' means the prosecution failed to meet its burden of proof, just as 'fail to reject H₀' means the data did not provide sufficient evidence against the null hypothesis. In both settings, the language encodes the logic of the procedure and prevents overreach.

Visual Explanation — The Anatomy of a Statistical Conclusion

A properly written statistical conclusion contains several mandatory components arranged in a logical order. The diagram below dissects a model conclusion sentence, labeling each element by its function. Understanding this anatomy ensures that every conclusion you write is complete, defensible, and unambiguous.

A model conclusion sentence decomposes into four mandatory components: ① the significance level that anchors the decision, ② the evidential phrase that conveys the outcome without overstating it, ③ the contextual claim restating the alternative hypothesis in plain language, and ④ the supporting test statistics that let the reader verify the claim.

Notice that the model conclusion never uses the words 'prove,' 'accept,' or 'true.' It also avoids phrasing the conclusion solely in symbolic notation — a statement like 'Reject H₀' is technically correct but communicatively incomplete. The contextual claim (component ③) is what makes the conclusion meaningful to a reader who may not be a statistician. Every piece of quantitative evidence — the test statistic, degrees of freedom, and p-value — is included parenthetically so that the claim is independently verifiable.

The Decision Framework Behind Every Conclusion

Before you can write a conclusion, you must understand the logical machinery that produces it. Every hypothesis test follows a deterministic decision rule: compare the computed p-value (or test statistic) to the predetermined significance level α. The outcome of that comparison dictates which of exactly two conclusion templates you use. Although the formulas differ across z-tests, t-tests, χ²-tests, and F-tests, the decision logic and the resulting language are universal.

The Decision Rule

REJECTION CRITERION (P-VALUE APPROACH)
If p-value ≤ α, reject H₀.
Where p-value is the probability of observing a test statistic at least as extreme as the one obtained, assuming H₀ is true, and α is the pre-specified significance level (commonly 0.01, 0.05, or 0.10).
FAILURE-TO-REJECT CRITERION
If p-value > α, fail to reject H₀.
This does not mean H₀ is true — only that the data do not provide sufficient evidence against it at the chosen α level. Power (1 − β) determines the test's ability to detect a real effect.

The Two Conclusion Templates

Once the decision is made, the conclusion follows one of two templates. Memorizing these templates and adapting them to context is the single most effective strategy for writing correct conclusions.

TEMPLATE A — Reject H₀ (p ≤ α)
"At the α = [value] significance level, there is sufficient evidence to conclude that [restate H₁ in context]. ([test statistic] = [value], [df if applicable], p = [value])."
TEMPLATE B — Fail to Reject H₀ (p > α)
"At the α = [value] significance level, there is insufficient evidence to conclude that [restate H₁ in context]. ([test statistic] = [value], [df if applicable], p = [value])."

Several points deserve emphasis. First, notice that both templates reference the alternative hypothesis (H₁), not the null. This is deliberate: the conclusion addresses the research question, which is always framed by H₁. Second, Template B says 'insufficient evidence to conclude that [H₁]' — it does not say 'there is evidence that H₀ is true.' Third, both templates open with the significance level because the conclusion is conditional on that choice; a different α could yield a different outcome with the same data.

Detailed Language Guide — Correct vs. Incorrect Phrasing

The difference between a correct and an incorrect statistical conclusion often comes down to a single word. This section catalogs the most common errors, explains why they are wrong, and offers corrected alternatives. Refer to the diagram and table below as a reference when drafting your own conclusions.

Six common language errors (left, in red) paired with their corrected versions (right, in green). Each correction preserves the logical constraints of hypothesis testing: statistics provide evidence, not proof; we fail to reject rather than accept; and causal language requires experimental design.
Quick-reference table of forbidden phrases and their correct alternatives.
ScenarioForbidden PhraseCorrect Alternative
p ≤ α"This proves H₁ is true.""There is sufficient evidence to support [H₁ in context]."
p > α"We accept H₀.""There is insufficient evidence to conclude [H₁ in context]."
Observational study"X causes Y.""X is associated with Y."
Confidence interval"95% chance the true mean is in this interval.""We are 95% confident that the true mean lies between [a, b]."
p-value interpretation"There is a p% probability H₀ is true.""If H₀ were true, the probability of observing data this extreme is p."

Worked Example — From Test Output to Written Conclusion

A university researcher wants to determine whether a new tutoring program improves final exam scores. She collects a random sample of 35 students who used the program and 40 students who did not. The mean exam score for the tutoring group is 78.4 (s = 10.2), and the mean for the control group is 73.1 (s = 11.5). She conducts a two-sample t-test at the α = 0.05 significance level. The test yields t = 2.14, df = 71, p = 0.036. Let us walk through the process of converting this output into a properly worded statistical conclusion.

Writing a Conclusion from a Two-Sample t-Test
1
Step 1 — State the HypothesesBegin by recalling the null and alternative hypotheses. H₀: μ₁ = μ₂ (there is no difference in mean exam scores between the tutoring and control groups). H₁: μ₁ ≠ μ₂ (there is a difference in mean exam scores). Note: if the researcher specifically predicted the tutoring group would score higher, a one-tailed test with H₁: μ₁ > μ₂ would be appropriate — but the problem states a two-sample t-test without directional specification.
H₀: μ₁ = μ₂ ; H₁: μ₁ ≠ μ₂
2
Step 2 — Identify the Decision RuleThe significance level is α = 0.05. Since p = 0.036, we compare: is 0.036 ≤ 0.05? Yes. Therefore, we reject H₀.
p = 0.036 ≤ α = 0.05 → Reject H₀
3
Step 3 — Select the Correct TemplateBecause we rejected H₀, we use Template A: 'At the α = [value] significance level, there is sufficient evidence to conclude that [restate H₁ in context].' We must now translate H₁ (μ₁ ≠ μ₂) into plain language tied to the research context.
Template A selected (sufficient evidence)
4
Step 4 — Restate H₁ in ContextThe contextual restatement of H₁ should reference the specific populations and variable: 'the mean final exam scores differ between students who used the tutoring program and those who did not.' Avoid vague language like 'the groups are different.' Specificity matters.
Contextual claim: mean final exam scores differ between tutoring and control groups
5
Step 5 — Assemble the Final ConclusionCombine all components — significance level, evidential phrase, contextual claim, and test statistics — into a single, polished sentence. Also consider whether to add a study-design caveat. Since this is observational (no mention of random assignment to groups), we should note the limitation.
Final conclusion: "At the α = 0.05 significance level, there is sufficient evidence to conclude that the mean final exam scores differ between students who participated in the tutoring program and those who did not (t = 2.14, df = 71, p = 0.036). Because students self-selected into the program, this result demonstrates an association rather than a causal effect."
DESIGN CAVEAT
The final sentence of the conclusion above — 'this result demonstrates an association rather than a causal effect' — is not mere boilerplate. Because students were not randomly assigned to the tutoring program, confounding variables (e.g., motivation, prior GPA) could explain the difference. Including this caveat is expected in any college-level conclusion drawn from an observational study.

Conclusion Language Across Different Test Types

The conclusion template remains structurally identical regardless of the test you perform. What changes is the test statistic symbol, the degrees of freedom format, and the contextual phrasing of the claim. The table below provides model conclusions for the most common test types encountered in an introductory statistics course, illustrating how the template adapts while the core language stays constant.

Model conclusion phrasing for common hypothesis tests. Replace bracketed placeholders with your data.
Test TypeKey Statistics ReportedModel Conclusion Phrasing (Reject H₀)
One-sample z-testz, p"…sufficient evidence that the population mean differs from [μ₀] (z = [val], p = [val])."
One-sample t-testt, df, p"…sufficient evidence that the mean [variable] is [greater than / less than / different from] [μ₀] (t = [val], df = [val], p = [val])."
Two-sample t-testt, df, p"…sufficient evidence that the mean [variable] differs between [Group A] and [Group B] (t = [val], df = [val], p = [val])."
Paired t-testt, df, p"…sufficient evidence that the mean difference in [variable] between [conditions] is not zero (t = [val], df = [val], p = [val])."
χ² test of independenceχ², df, p"…sufficient evidence of an association between [Variable 1] and [Variable 2] (χ² = [val], df = [val], p = [val])."
One-way ANOVAF, df₁, df₂, p"…sufficient evidence that at least one group mean differs (F = [val], df = [df₁, df₂], p = [val])."
Linear regression (slope)t or F, df, p, r²"…sufficient evidence of a linear relationship between [X] and [Y] (t = [val], df = [val], p = [val])."
KEY TAKEAWAY
Think of the conclusion template as a reusable form letter: the structure and legal language stay the same, but the names, dates, and specifics change with each case. Whether you are running a z-test on proportions or an ANOVA on group means, the skeleton — significance level, evidential phrase, contextual claim, test statistics — remains identical. Mastering the template once means you can write correct conclusions for any test you will ever encounter.

Connecting to Confidence Intervals and Effect Sizes

Hypothesis test conclusions answer a binary question — is there sufficient evidence or not? — but modern statistical practice increasingly demands that conclusions go further. Two extensions are essential for advanced work: confidence intervals and effect sizes. Both enrich the conclusion by quantifying the magnitude and precision of the observed effect, not merely its statistical significance.

Hypothesis test conclusions vs. enhanced conclusions with confidence intervals and effect sizes.
FeatureHypothesis Test ConclusionEnhanced Conclusion (with CI & Effect Size)
Question answeredIs the effect different from zero (or the hypothesized value)?How large is the effect, and how precisely is it estimated?
Key language"sufficient / insufficient evidence""The 95% CI for the mean difference is [2.1, 8.5]" and "Cohen's d = 0.49"
Information conveyedDirection of evidence (for or against H₀)Magnitude, direction, and precision of the estimated effect
Sensitivity to sample sizeLarge n can make trivial effects significantEffect size is independent of n; CI width reflects precision
When to useIntro courses, binary decision contextsResearch reports, publications, applied settings

An enhanced conclusion might read: 'At the α = 0.05 significance level, there is sufficient evidence that the tutoring program is associated with higher exam scores (t = 2.14, df = 71, p = 0.036). The 95% confidence interval for the difference in means is (0.36, 10.24), and the effect size is Cohen's d = 0.49, indicating a medium practical effect.' This richer statement helps stakeholders evaluate not just whether an effect exists but whether it is large enough to matter — a distinction that hypothesis testing alone cannot make. As you move into upper-division courses and research methods, you should aim to include these elements in every conclusion you write.

📐 CONFIDENCE INTERVAL LANGUAGE
When reporting a confidence interval, say 'We are 95% confident that the true [parameter] lies between [a] and [b].' Do not say 'There is a 95% probability that [parameter] is in this interval.' The parameter is fixed; the interval is random. The confidence level refers to the long-run success rate of the procedure, not the probability for any single interval.

Practice Problems

PROBLEM 1CONCEPTUAL
A student writes the following conclusion: 'Since p = 0.12, we accept the null hypothesis and conclude that the new medication has no effect.' Identify two errors in this statement and rewrite it using proper statistical language. Assume α = 0.05.
PROBLEM 2BASIC CALCULATION
A one-sample t-test is performed to test whether the mean commute time for employees at a company differs from 30 minutes. The test yields t = −2.45, df = 49, p = 0.018, with α = 0.05. Write the complete statistical conclusion.
PROBLEM 3INTERMEDIATE
A researcher uses a chi-squared test of independence to investigate whether there is an association between exercise frequency (low, moderate, high) and sleep quality (poor, fair, good) in a sample of 240 adults. The test yields χ² = 5.83, df = 4, p = 0.212. The study is observational. Write a proper conclusion at the α = 0.05 level, and explain why the word 'cause' would be inappropriate here even if p had been less than 0.05.
PROBLEM 4APPLIED
A pharmaceutical company conducts a randomized controlled trial comparing a new blood-pressure medication to a placebo. Among 120 patients randomized to the drug, the mean reduction in systolic BP is 8.3 mmHg (s = 6.1); among 115 placebo patients, the mean reduction is 3.1 mmHg (s = 5.8). A two-sample t-test yields t = 6.52, df = 232, p < 0.001. The 95% CI for the difference in means is (3.6, 6.8). Write a comprehensive conclusion that includes both the hypothesis test result and the confidence interval, using language appropriate for the randomized experimental design.
PROBLEM 5CRITICAL THINKING
A researcher tests 20 independent hypotheses in an exploratory study, each at α = 0.05. One test yields p = 0.038, and the remaining 19 yield p > 0.05. The researcher writes: 'There is sufficient evidence at the 0.05 level that Variable 7 is significantly associated with the outcome.' Critically evaluate this conclusion. What statistical concept is being violated, and how should the researcher modify both the analysis and the conclusion?

Lesson Summary

Writing a statistical conclusion requires more than announcing a result — it demands precise language that reflects the logic of inference. Every conclusion should open with the significance level (α), include an evidential phrase ('sufficient evidence' when p ≤ α or 'insufficient evidence' when p > α), restate the alternative hypothesis in real-world context, and report the test statistic, degrees of freedom, and p-value. These four components form the non-negotiable skeleton of every defensible conclusion.

Equally important is what you must never say: do not claim that results 'prove' a hypothesis, do not 'accept H₀' when p > α, do not equate a p-value with the probability that H₀ is true, and do not use causal language for observational studies. For advanced reporting, supplement the hypothesis test conclusion with a confidence interval and an effect size to communicate both the magnitude and precision of the finding. Mastering this language ensures that your conclusions are accurate, transparent, and worthy of publication.

Varsity Tutors • College Statistics • Writing Statistical Conclusions — Writing Conclusions with Proper Statistical Language