IB MATHEMATICS: APPLICATIONS AND INTERPRETATION • STATISTICS AND PROBABILITY

Correlation & Regression — SL 4.2 Correlation and linear regression; interpreting slope/intercept and residuals (technology-supported)

Discover how to measure relationships between variables and use regression lines to make predictions.

Historical Context & Motivation

Humans have always looked for patterns in data, but for most of history we relied on intuition rather than mathematics. Is there a connection between how many hours you study and the grade you earn? Does temperature affect ice-cream sales? These questions feel simple, yet answering them rigorously required centuries of mathematical development. The story of correlation and regression begins in the 1800s, when scientists first tried to quantify how two measurements move together.

1805
Legendre's Method of Least Squares
French mathematician Adrien-Marie Legendre published the method of least squares, a technique for fitting the best straight line through a set of data points by minimizing the sum of squared differences.
1886
Galton Coins 'Regression'
Sir Francis Galton studied the heights of parents and children, discovering that children's heights tended to regress toward the mean. He called this phenomenon 'regression,' giving the technique its modern name.
1896
Pearson's Correlation Coefficient
Karl Pearson formalized the Pearson product-moment correlation coefficient (r), providing a single number between −1 and 1 that measures the strength and direction of a linear relationship.
1970s–Today
Technology-Supported Analysis
With graphing calculators and software like GDCs, spreadsheets, and Desmos, students and researchers can compute regression equations and correlation coefficients in seconds, shifting the focus from calculation to interpretation.

The central question this topic addresses is deceptively straightforward: When two variables seem related, how do we measure the strength of that relationship and use it to make predictions? In IB Mathematics: Applications and Interpretation SL 4.2, you will learn how to answer this question using scatter plots, the correlation coefficient r, the regression line ŷ = ax + b, and residual analysis — all supported by technology.

Core Principles & Definitions

Before diving into calculations, it is essential to build a clear vocabulary. The ideas below form the foundation of everything you will do with bivariate data — data sets that pair up two variables for each individual or observation.

1

Bivariate Data & Scatter Plots

Bivariate data consists of paired observations (x, y). When plotted on coordinate axes, each pair becomes a dot, and the collection of dots is called a scatter plot. The independent (explanatory) variable goes on the x-axis and the dependent (response) variable on the y-axis.
2

Correlation Coefficient (r)

The Pearson correlation coefficient (r) is a number from −1 to 1 that describes both the direction and strength of a linear relationship. Values close to 1 or −1 indicate a strong linear pattern, while values near 0 suggest little or no linear association.
3

Line of Best Fit (Regression Line)

The regression line ŷ = ax + b is the unique straight line that minimizes the sum of the squared residuals. It passes through the point (x̄, ȳ) and is calculated using the least-squares method.
4

Residuals

A residual is the vertical distance between an observed y-value and the y-value predicted by the regression line: residual = y − ŷ. Positive residuals sit above the line; negative residuals sit below it.
5

Interpolation vs. Extrapolation

Interpolation is predicting within the range of the given data and is generally reliable. Extrapolation is predicting outside that range and can be unreliable because the linear trend may not continue.
KEY TAKEAWAY
Think of correlation like a friendship rating. If two variables are strongly positively correlated (r close to 1), they are 'best friends' — when one goes up, the other goes up too. If they are strongly negatively correlated (r close to −1), they are 'opposites' — when one goes up, the other goes down. An r near 0 means they are strangers who act independently of each other. The regression line is your best prediction tool: it draws the tightest-fitting line through the data so you can estimate one variable from the other.

Visual Explanation — The Scatter Plot & Regression Line

A scatter plot is the starting point for every correlation and regression analysis. The diagram below shows a bivariate data set with a regression line drawn through it. Notice how the data points cluster around the line, with some sitting above it (positive residuals) and others below (negative residuals). The vertical dashed segments illustrate the residuals themselves.

Each cyan dot represents an (hours, score) pair. The violet line is the least-squares regression line ŷ = 6.1x + 26. The pink and green dashed segments show residuals — the vertical distances between observed points and the line. Points above the line have positive residuals; points below have negative residuals.

In the diagram, notice that the data points generally trend upward from left to right. This indicates a positive correlation: as hours of study increase, test scores tend to increase as well. The regression line captures this upward trend and provides a formula you can use for predictions. The residuals show you where the model is imperfect — no real-world data falls perfectly on a line. Examining whether residuals are randomly scattered (good) or show a curved pattern (bad) tells you whether a linear model is appropriate.

Mathematical Framework

In IB Applications and Interpretation SL, you are expected to use your GDC (graphing display calculator) or approved technology to compute the values of r, a, and b. However, understanding the formulas helps you know what your calculator is doing and why the outputs make sense.

REGRESSION LINE (LEAST-SQUARES)
ŷ = ax + b
ŷ = predicted value of the dependent variable; a = slope (gradient) of the line; b = y-intercept (the value of ŷ when x = 0); x = value of the independent variable.
SLOPE OF REGRESSION LINE
a = Sₓᵧ / Sₓₓ
Sₓᵧ = Σ(xᵢ − x̄)(yᵢ − ȳ) is the sum of the products of deviations; Sₓₓ = Σ(xᵢ − x̄)² is the sum of squared deviations of x. Your GDC computes these automatically.
PEARSON CORRELATION COEFFICIENT
r = Sₓᵧ / √(Sₓₓ × Sᵧᵧ)
Sᵧᵧ = Σ(yᵢ − ȳ)². The value of r always satisfies −1 ≤ r ≤ 1. A positive r means a positive linear trend; a negative r means a negative linear trend.
RESIDUAL
eᵢ = yᵢ − ŷᵢ
yᵢ = observed (actual) value; ŷᵢ = value predicted by the regression line at the same x. A residual of zero means the point sits exactly on the line.
📝 IB Exam Tip
On the IB exam, you will not be asked to compute r or the regression equation by hand. You will use your GDC. The exam focuses on interpreting the slope, intercept, correlation coefficient, and residuals in context. Always write your interpretation in the context of the problem (e.g., 'For every additional hour of study, the predicted test score increases by 6.1 points').

Interpreting the Correlation Coefficient

One of the most common exam tasks is to describe the correlation between two variables. You need to state both the direction (positive or negative) and the strength (weak, moderate, or strong) of the correlation. The spectrum bar below gives you a visual reference for how to classify values of r.

Strength of Linear Correlation (r)
Strong −
Moderate −
Weak −
None
Weak +
Moderate +
Strong +
−1
−0.7
−0.4
0
0.4
0.7
1
Perfect negativePerfect positive
Four scatter plots illustrate how different values of r look visually. As |r| approaches 1, data points cluster more tightly around the regression line. When r ≈ 0, data points appear as a shapeless cloud. The information boxes below summarize how to interpret the slope and intercept in context.
⚠️ Correlation ≠ Causation
A strong correlation does not prove that one variable causes changes in the other. For example, ice-cream sales and drowning rates are strongly positively correlated — but buying ice cream does not cause drowning. Both increase because of a lurking variable: hot weather. Always be cautious about claiming cause and effect based on correlation alone.

Worked Example

A teacher records the number of absences (x) and the final exam score (y) for eight students. The data are shown below. Let's walk through a complete analysis using technology.

Absences vs. Exam Score data
StudentAbsences (x)Exam Score (y)
A288
B574
C382
D860
E192
F668
G478
H765
Finding and Interpreting the Regression Line
1
Step 1 — Enter Data into Your GDCEnter the absences (x) into List 1 and the exam scores (y) into List 2 on your GDC. Make sure both lists have the same number of entries (8 in this case).
2
Step 2 — Run Linear RegressionUse the statistics menu to perform linear regression (LinReg or equivalent). The GDC outputs: a = −4.37 (3 s.f.), b = 96.5 (3 s.f.), and r = −0.987 (3 s.f.).
ŷ = −4.37x + 96.5, r = −0.987
3
Step 3 — Interpret the SlopeThe slope a = −4.37 means that for each additional absence, the predicted exam score decreases by approximately 4.37 points. The negative slope reflects the negative correlation between absences and scores.
4
Step 4 — Interpret the y-InterceptThe y-intercept b = 96.5 means that a student with zero absences would be predicted to score about 96.5 on the exam. This is within a reasonable range, so the intercept has a sensible real-world interpretation here.
5
Step 5 — Interpret the Correlation CoefficientSince r = −0.987, there is a very strong negative linear correlation between absences and exam scores. The value is very close to −1, which tells us the data points lie almost exactly along the regression line.
6
Step 6 — Make a Prediction and Calculate the ResidualPredict the score for a student with 3 absences: ŷ = −4.37(3) + 96.5 = −13.11 + 96.5 = 83.4 (3 s.f.). The actual score for Student C (x = 3) is 82. The residual is: e = 82 − 83.4 = −1.4. The negative residual tells us the student scored slightly below the prediction.
Predicted score = 83.4; Residual = −1.4
7
Step 7 — Assess Reliability of a PredictionIf someone asks you to predict the score for a student with 15 absences, be cautious! The data only goes up to x = 8. Predicting at x = 15 would be extrapolation and the result (ŷ = −4.37 × 15 + 96.5 = 31) may be unreliable because the linear trend might not hold that far.

Strengths & Limitations of Linear Regression

Linear regression is one of the most widely used statistical tools in the world, but it is not perfect. Understanding when it works well and when it fails is just as important as knowing how to compute the regression line.

Strengths vs. Limitations of Linear Regression
StrengthsLimitations
Simple to compute and interpret — a single equation summarizes the relationship.Only models linear relationships; curved patterns will be poorly described.
The correlation coefficient r gives a quick measure of strength and direction.Sensitive to outliers — a single extreme point can dramatically shift the line.
Enables predictions (interpolation) within the data range with reasonable accuracy.Extrapolation beyond the data range is unreliable and can produce nonsensical predictions.
Residual analysis helps diagnose whether the model is appropriate.Correlation does not imply causation — lurking variables may create misleading associations.
Technology makes the computation nearly instantaneous on a GDC or spreadsheet.Requires bivariate quantitative data — cannot be applied to categorical variables directly.
KEY TAKEAWAY
Think of the regression line as a GPS route. It gives you the best estimated path from A to B based on the roads (data) it knows about. If you stay on roads it has mapped (interpolation), the directions are reliable. But if you drive off into unmapped territory (extrapolation), the GPS might send you into a lake. Also, just because two cities are connected by a road doesn't mean one city caused the other to be built — correlation is not causation.

Connection to Advanced Theory

The linear regression you learn in SL 4.2 is the foundation for more powerful techniques you may encounter in HL Mathematics or university statistics courses. The table below shows how the concepts you have learned relate to their advanced counterparts.

SL 4.2 vs. Advanced Statistics
SL 4.2 ConceptAdvanced Extension
Simple linear regression (one x variable)Multiple regression: ŷ = a₁x₁ + a₂x₂ + … + b, using several predictor variables simultaneously
Pearson's r for linear correlationCoefficient of determination r² (the proportion of variance explained by the model); Spearman's rank correlation for non-linear monotonic relationships
Residuals (y − ŷ) checked visuallyFormal residual analysis including tests for normality, homoscedasticity (constant variance), and independence
Line of best fit (linear model)Non-linear regression models: quadratic, exponential, logarithmic, logistic, and polynomial curves fitted to data

The key idea connecting SL to higher-level work is that r² (the coefficient of determination) tells you what fraction of the variability in y is 'explained' by the model. For instance, if r = −0.987, then r² ≈ 0.974, meaning about 97.4% of the variation in exam scores can be explained by the number of absences. This powerful interpretation extends directly into university-level statistics and data science.

Practice Problems

PROBLEM 1CONCEPTUAL
A researcher finds that the correlation between daily temperature and the number of hot drinks sold at a café is r = −0.82. Describe the strength and direction of this correlation in context, and explain whether this means that high temperatures cause fewer hot drinks to be sold.
PROBLEM 2BASIC CALCULATION
The regression line for the relationship between hours of exercise per week (x) and resting heart rate in beats per minute (y) is ŷ = −1.8x + 78. (a) Interpret the slope in context. (b) Interpret the y-intercept in context. (c) Predict the resting heart rate for someone who exercises 5 hours per week.
PROBLEM 3INTERMEDIATE
Using the regression equation ŷ = −1.8x + 78 from Problem 2, a person who exercises 4 hours per week actually has a resting heart rate of 74 bpm. (a) Calculate the residual. (b) Explain what the sign of the residual tells you. (c) If the data set included exercise values from x = 1 to x = 10, would it be appropriate to predict the heart rate for someone exercising 20 hours per week? Justify your answer.
PROBLEM 4APPLIED
A car dealership records the age (in years) and selling price (in thousands of dollars) of 10 used cars of the same model. Using a GDC, the regression line is found to be ŷ = −2.15x + 28.3 with r = −0.946. (a) Describe the correlation. (b) Interpret the slope and y-intercept in context. (c) Estimate the price of a 6-year-old car. (d) A 4-year-old car sold for $22,000. Calculate the residual and explain what it suggests about this car.
PROBLEM 5CRITICAL THINKING
Two students analyse the same data set. Student A plots 'advertising spend' on the x-axis and 'revenue' on the y-axis, obtaining ŷ = 3.2x + 50 with r = 0.91. Student B plots 'revenue' on the x-axis and 'advertising spend' on the y-axis, obtaining a different regression equation. (a) Will Student B get the same value of r? Explain. (b) Will Student B get the same regression line? Explain why or why not. (c) Discuss which student's model is more appropriate for predicting revenue from advertising spend, and what assumption must be checked about the residuals before trusting either model.

Lesson Summary

In SL 4.2, you learned to analyse bivariate data using scatter plots to visualize relationships. The Pearson correlation coefficient (r) quantifies the direction and strength of a linear association on a scale from −1 to 1. Using technology (your GDC), you can find the least-squares regression line ŷ = ax + b, where the slope (a) tells you how much ŷ changes for each one-unit increase in x, and the y-intercept (b) gives the predicted value when x = 0.

Residuals (y − ŷ) measure how far each data point falls from the regression line and should be randomly scattered if a linear model is appropriate. Always interpret your results in context. Use interpolation for reliable predictions within the data range, and be cautious with extrapolation outside it. Finally, remember that correlation does not imply causation — lurking variables may create associations that do not reflect a direct cause-and-effect relationship.

Varsity Tutors • IB Mathematics: Applications and Interpretation • Correlation & Regression — SL 4.2