IB CHEMISTRY • SKILLS IN THE STUDY OF CHEMISTRY

Concluding & Evaluating Investigations — Concluding and evaluating investigations

Learn to draw valid conclusions from data and critically evaluate the quality of your experimental investigations.

Historical Context & Motivation

Science has always depended on the ability to draw reliable conclusions from observations. In the early days of chemistry, alchemists often failed to distinguish between genuine patterns in their data and wishful thinking, leading to centuries of fruitless pursuits like transforming lead into gold. The development of the scientific method introduced a structured approach: form a hypothesis, collect data, and then critically assess whether your results actually support your claim. This final stage—concluding and evaluating—is where the real intellectual rigor of science lives.

1620
Bacon's Novum Organum
Francis Bacon published his framework for inductive reasoning in science, arguing that conclusions must arise from systematic observation rather than speculation.
1789
Lavoisier's Quantitative Chemistry
Antoine Lavoisier demonstrated that careful measurement and evaluation of experimental mass data could overthrow the phlogiston theory and establish the law of conservation of mass.
1935
Fisher's Statistical Methods
Ronald Fisher formalized statistical hypothesis testing, giving scientists quantitative tools to evaluate whether experimental results are significant or due to chance.
1968
IB Programme Founded
The International Baccalaureate programme was established, placing explicit emphasis on experimental skills including data evaluation and conclusion-writing as core competencies in science education.

The central question this lesson addresses is straightforward but surprisingly tricky in practice: How do you know if your experiment actually worked? Drawing a conclusion is more than just restating your results. It requires you to connect your data back to your hypothesis, acknowledge the limitations of your method, and suggest how the investigation could be improved. In the IB Chemistry programme, this skill is assessed directly in your Internal Assessment, making it essential to master.

Core Principles & Definitions

Concluding and evaluating an investigation involves several interconnected skills. You need to interpret processed data, relate your findings to a scientific context, identify errors and limitations, and propose realistic improvements. Let's break these down into their essential components.

1

Conclusion

A statement that directly addresses the research question or hypothesis, supported by reference to processed data (graphs, calculations, trends). A good conclusion interprets what the data shows and why it makes chemical sense.
2

Systematic Errors

Errors that shift all measurements consistently in one direction (too high or too low). They affect accuracy and cannot be reduced by repeating trials. Examples include a miscalibrated balance or heat loss in calorimetry.
3

Random Errors

Unpredictable fluctuations that cause measurements to scatter around the true value. They affect precision and can be reduced by taking more trials and averaging. Examples include reading a burette or slight temperature fluctuations.
4

Accuracy vs. Precision

Accuracy describes how close a result is to the accepted or true value. Precision describes how closely repeated measurements agree with each other. An experiment can be precise but inaccurate, or accurate but imprecise.
5

Improvements

Specific, realistic modifications to the experimental design that would reduce identified errors. Improvements must be linked to particular weaknesses—vague statements like 'be more careful' are insufficient.
KEY TAKEAWAY
Think of evaluating an experiment like reviewing a restaurant meal. You don't just say "the food was good" (that's a vague conclusion). A proper review explains what was good, acknowledges what could be better, and suggests how the chef could improve. Systematic errors are like a recipe that's fundamentally flawed (always too salty), while random errors are like slight variations from plate to plate.

Visual Explanation — The Evaluation Cycle

The process of concluding and evaluating follows a logical cycle. You begin with your processed data, draw a conclusion that references your hypothesis, then systematically identify weaknesses in the method before suggesting improvements. The diagram below illustrates this flow and highlights the key questions you should ask at each stage.

The evaluation cycle shows how processed data leads to a conclusion, which is then scrutinized for systematic and random errors. Each identified error is paired with a specific improvement, ultimately informing a redesigned investigation.

Notice how the diagram separates the two types of error before linking each one to a targeted improvement. This is exactly how the IB expects you to structure your evaluation: identify the weakness first, classify it, and then propose a solution that directly addresses it. A vague improvement like "use better equipment" is meaningless unless you specify which piece of equipment and how it would reduce a particular error.

Mathematical Framework — Percentage Error & Uncertainty

Quantifying how well your experiment performed is a core part of evaluation. The two most commonly used calculations in IB Chemistry are percentage error (which assesses accuracy) and percentage uncertainty (which assesses precision). These numbers tell you whether your errors are significant enough to invalidate your conclusion.

PERCENTAGE ERROR
% Error = |Experimental Value − Accepted Value| ÷ Accepted Value × 100%
The vertical bars indicate absolute value (always positive). A small percentage error suggests high accuracy. If percentage error exceeds the total percentage uncertainty, systematic errors are likely present.
PERCENTAGE UNCERTAINTY (SINGLE MEASUREMENT)
% Uncertainty = (Absolute Uncertainty ÷ Measured Value) × 100%
Absolute uncertainty is typically ±½ the smallest division of an analogue instrument, or ±1 of the last digit for digital instruments. This measures precision.
PROPAGATION — ADDITION/SUBTRACTION
Total Absolute Uncertainty = Δa + Δb + Δc + …
When adding or subtracting measured values, add the absolute uncertainties. For example, a temperature change ΔT = T₂ − T₁ has uncertainty ΔT = ΔT₂ + ΔT₁.
PROPAGATION — MULTIPLICATION/DIVISION
Total % Uncertainty = %a + %b + %c + …
When multiplying or dividing, add the percentage uncertainties. This total percentage uncertainty sets the threshold: if your percentage error is smaller, random errors may account for the discrepancy.
⚠️ Key Decision Rule
Compare your percentage error to your total percentage uncertainty. If % error < total % uncertainty, your result is consistent with the accepted value within experimental limitations. If % error > total % uncertainty, systematic errors are likely present and must be identified.

Detailed Breakdown — Classifying Errors & Limitations

When evaluating an investigation, you need to do more than simply say "there were errors." The IB expects you to identify specific weaknesses, classify them as systematic or random, explain their impact on your results, and propose targeted improvements. The diagram below and the table that follows give you a framework for doing this effectively.

Four target diagrams illustrating the combinations of accuracy and precision. The red dot at the center represents the true/accepted value. Tightly clustered dots (high precision) result from reducing random errors; dots centered on the bullseye (high accuracy) result from eliminating systematic errors.
Comparing systematic and random errors in experimental chemistry
FeatureSystematic ErrorRandom Error
DefinitionConsistent deviation in one directionUnpredictable fluctuations in both directions
Effect on resultsReduces accuracy (shifts mean away from true value)Reduces precision (increases scatter of data)
Can repeating trials help?No — the same bias affects every trialYes — averaging reduces the effect
Chemistry examplesHeat loss in calorimetry, impure reagent, miscalibrated balanceJudging colour change endpoint, reading a burette meniscus, minor temperature fluctuations
How to fixRedesign apparatus or method (insulation, calibration, purer reagents)Increase trials, use digital sensors, use data-logging software

Worked Example — Evaluating a Calorimetry Experiment

Suppose you performed a calorimetry experiment to determine the enthalpy of combustion of ethanol (C₂H₅OH). The accepted value is −1367 kJ mol⁻¹. Your experimental result was −1180 kJ mol⁻¹. You measured a temperature change of 14.2 °C using a thermometer with an uncertainty of ±0.5 °C, and you used a balance with an uncertainty of ±0.01 g to measure 0.46 g of ethanol. Let's walk through the full conclusion and evaluation.

Evaluating a Combustion Calorimetry Experiment
1
Step 1 — State the conclusionThe experimental enthalpy of combustion of ethanol was determined to be −1180 kJ mol⁻¹. This is lower in magnitude than the accepted value of −1367 kJ mol⁻¹, suggesting that not all the heat released by the combustion was absorbed by the water. The data supports the hypothesis that burning ethanol is exothermic, but the measured value is significantly lower than expected.
2
Step 2 — Calculate percentage error% Error = |−1180 − (−1367)| ÷ 1367 × 100% = 187 ÷ 1367 × 100%
% Error = 13.7%
3
Step 3 — Calculate percentage uncertaintiesTemperature change: ΔT = 14.2 °C with uncertainty ±1.0 °C (since ΔT = T₂ − T₁, uncertainties add: 0.5 + 0.5 = 1.0 °C). % uncertainty in ΔT = 1.0 ÷ 14.2 × 100% = 7.0%. Mass of ethanol: % uncertainty = 0.01 ÷ 0.46 × 100% = 2.2%. Total % uncertainty ≈ 7.0 + 2.2 = 9.2%.
Total % uncertainty ≈ 9.2%
4
Step 4 — Compare % error to % uncertaintyThe percentage error (13.7%) exceeds the total percentage uncertainty (9.2%). This means the discrepancy cannot be explained by random errors alone, and systematic errors must be present in the experiment.
13.7% > 9.2% → systematic errors are significant
5
Step 5 — Identify errors and propose improvementsSystematic error 1: Heat loss to the surroundings — the open calorimeter allowed significant heat to escape, causing the temperature rise (and therefore ΔH) to be too low. Improvement: Use a bomb calorimeter or add polystyrene insulation and a lid. Systematic error 2: Incomplete combustion — soot on the bottom of the can suggests not all ethanol burned completely. Improvement: Ensure a well-ventilated draught shield to promote complete combustion. Random error: Difficulty reading the thermometer precisely during rapid temperature change. Improvement: Use a digital temperature probe connected to data-logging software.

Strengths, Limitations & Common Pitfalls

Writing a strong evaluation requires avoiding several common mistakes that students make. The table below contrasts what the IB considers strong evaluation responses versus weak ones that would receive limited credit.

Comparison of weak vs. strong evaluation responses for IB Chemistry
AspectWeak Response ✗Strong Response ✓
Conclusion"The experiment worked and supported the hypothesis.""The data shows a proportional relationship between concentration and rate, with R² = 0.97, supporting the predicted first-order kinetics."
Error identification"There were some errors in the experiment.""Heat loss to surroundings is a systematic error that consistently reduced the measured ΔT, giving an enthalpy value lower than expected."
Improvement"We should be more careful next time.""Using a polystyrene cup with a lid and conducting the reaction in a draught-free environment would reduce heat loss."
Data reference"The graph shows a trend.""As shown in Figure 2, the gradient of the best-fit line is 0.034 mol⁻¹ dm³ s⁻¹, which gives a rate constant consistent with literature values."
Uncertainty analysis"The uncertainty was small.""The % error of 13.7% exceeds the total % uncertainty of 9.2%, indicating that systematic errors dominate over random measurement limitations."
KEY TAKEAWAY
In science, admitting what went wrong is not a weakness—it's a strength. The best evaluations demonstrate that you understand why your results deviate from expected values. Think of it like debugging code: you wouldn't just say "the program has bugs." You'd identify the specific line causing the issue and explain how to fix it. That's exactly what a good evaluation does for an experiment.

Connection to Advanced Theory — Validity & Reliability

Beyond the IB Chemistry IA, the concepts of concluding and evaluating investigations connect to broader ideas in scientific methodology. Two terms that often appear in advanced discussions are validity and reliability. Understanding these will help you think more critically about experimental design and is essential preparation if you continue into university-level sciences.

IB-level evaluation compared to advanced scientific methodology
ConceptIB Chemistry LevelUniversity / Advanced Level
ConclusionState whether data supports hypothesis; reference processed dataInclude statistical tests (t-tests, chi-squared) to determine if results are significant at a given confidence level
Error analysisIdentify systematic and random errors; calculate % error and % uncertaintyUse standard deviation, standard error of the mean, and error propagation through calculus-based partial derivatives
ValidityConsider whether method measures what it claims; control variablesInternal validity (confounding variables controlled) and external validity (generalizability to other contexts)
ReliabilityRepeat trials to show consistency; note outliersReproducibility studies, inter-laboratory comparisons, peer replication

As you progress in science, you'll find that the same logical structure applies: draw a conclusion from evidence, quantify confidence in your results, identify weaknesses, and improve. The difference at higher levels is the mathematical sophistication of the tools. Mastering the fundamentals now—classifying errors, calculating percentage error, and linking improvements to specific weaknesses—gives you the foundation for all future experimental work.

Practice Problems

PROBLEM 1CONCEPTUAL
A student conducts three trials of a titration and obtains volumes of 24.80 cm³, 24.75 cm³, and 24.85 cm³. The accepted value is 25.20 cm³. Are these results accurate, precise, both, or neither? Justify your answer.
PROBLEM 2BASIC CALCULATION
A student measures the molar mass of magnesium experimentally and obtains a value of 25.8 g mol⁻¹. The accepted value is 24.31 g mol⁻¹. Calculate the percentage error.
PROBLEM 3INTERMEDIATE
In an enthalpy of neutralisation experiment, a student uses a thermometer (±0.5 °C) to measure T₁ = 21.5 °C and T₂ = 28.3 °C, and a graduated cylinder (±0.5 cm³) to measure 50.0 cm³ of solution. Calculate the total percentage uncertainty in the enthalpy value. (Assume the specific heat capacity and density have negligible uncertainty.)
PROBLEM 4APPLIED
A student investigates the rate of reaction between magnesium ribbon and hydrochloric acid by measuring gas volume over time. Their results show a percentage error of 18% compared to the theoretical yield, while the total percentage uncertainty from their apparatus is only 4.5%. Identify two specific systematic errors that could account for the discrepancy, explain their effects, and propose one improvement for each.
PROBLEM 5CRITICAL THINKING
Two students perform the same calorimetry experiment. Student A obtains a percentage error of 5% with a total percentage uncertainty of 8%. Student B obtains a percentage error of 22% with a total percentage uncertainty of 6%. Compare and contrast their evaluations. Which student's result is more scientifically meaningful, and what does each result tell you about the nature of the errors in their experiments?

Summary — Concluding & Evaluating Investigations

A strong conclusion directly addresses the research question, references processed data (graphs, calculated values, trends), and explains the results using relevant chemical theory. You should clearly state whether the data supports or refutes your hypothesis and discuss any patterns or anomalies. The quantitative backbone of evaluation involves comparing percentage error (a measure of accuracy) against total percentage uncertainty (a measure of precision) to determine whether systematic errors are significant.

Effective evaluation requires identifying specific systematic errors (which bias results in one direction and cannot be fixed by repeating trials) and random errors (which cause scatter and can be reduced by averaging more trials). Each identified weakness must be paired with a realistic, specific improvement that explains exactly how the modification would reduce that particular error. Remember: vague statements earn limited credit. Always connect your conclusion back to the science, your errors to quantitative evidence, and your improvements to the specific weaknesses they address.

Varsity Tutors • IB Chemistry • Concluding & Evaluating Investigations