Where Did Statistics Come From?
People have been collecting data for thousands of years. Governments counted their citizens, farmers tracked their harvests, and traders kept records of prices. But it took a long time before anyone figured out smart ways to summarize all those numbers. Here's a quick look at how the ideas of center and spread developed over time.
So here's the big question this lesson answers: when you collect a bunch of numbers, how do you describe what those numbers look like as a group? You do it by talking about the center (where the data clusters) and the spread (how far the data stretches out).
Core Ideas: Distribution, Center & Spread
Before we dive in, let's nail down some vocabulary. A statistical question is a question where you expect the answers to vary. For example, "How tall are the students in my class?" is a statistical question because not everyone is the same height. When you collect those heights, you get a data set. The way those values are arranged—where they pile up and where they thin out—is called the distribution.
Distribution
Center
Spread
Shape
Seeing a Distribution: Dot Plots
One of the easiest ways to see a distribution is with a dot plot. Each data value gets a dot, and dots stack up above a number line. Let's look at an example. Suppose you asked 20 classmates, "How many books did you read last month?" and got these results:
Here's what that data looks like as a dot plot. Notice how the dots pile up around 3 and 4—that's the center of this distribution. The data spreads from 1 to 7.
In this dot plot you can instantly see the center—the tallest stack is at 4, and most values are between 2 and 5. You can also see the spread—the data goes from 1 on the left all the way to 7 on the right. The shape is roughly symmetric, meaning it looks about the same on the left and right of the tallest stack.
Measuring Center & Spread with Numbers
Graphs are great, but sometimes you need a single number to describe a data set. Here are the most important formulas.
Measures of Center
The mean is like the balance point of the data. If you put all the dots on a seesaw number line, the mean is where the seesaw would balance perfectly.
The median splits the data in half—half the values are below it and half are above it. It's very useful when your data has extreme values (called outliers) that could pull the mean in one direction.
Measures of Spread
The range gives you a quick idea of spread, but it only uses two values (the biggest and smallest). A more reliable measure is the interquartile range (IQR), which focuses on the middle 50% of the data. You'll explore the IQR more in later lessons.
Shapes of Distributions
Not all data sets look the same when you graph them. The shape of a distribution helps you decide which measure of center and spread to use. Let's look at three common shapes.
When a distribution is symmetric, the mean and median are close together, so either one is a good description of center. When data is skewed (one tail is longer than the other), the mean gets pulled toward the tail. In that case, the median is a better choice because it stays in the middle of the data no matter what.
For spread, the range is quick to calculate but can be misleading if there's a single outlier. That's why statisticians often prefer the interquartile range, which ignores the extremes and focuses on the middle half.
Worked Example
A teacher asked 12 students, "How many minutes did you spend on homework last night?" Here are the responses:
Comparing Mean, Median & Mode
Each measure of center has strengths and weaknesses. Here's a handy comparison table.
| Measure | Strengths | Limitations |
|---|---|---|
| Mean | Uses every data value; great for symmetric data | Pulled by outliers; can be misleading for skewed data |
| Median | Not affected by outliers; works well for skewed data | Ignores how far apart values are from each other |
| Mode | Easy to find; works for non-number data (like favorite color) | May not exist, or there may be several; doesn't describe spread at all |
| Range | Super quick to calculate; gives an overall idea of spread | Only uses two values; heavily affected by outliers |
Looking Ahead: Beyond Center & Spread
In this lesson you learned two big ideas: center and spread. But statisticians use even more tools to describe data as you move into 7th and 8th grade and beyond. Here's a quick peek at what's coming.
| What You Know Now | What You'll Learn Later |
|---|---|
| Range (max − min) | Interquartile Range (IQR) — measures the spread of just the middle 50% of data |
| Mean (average) | Mean Absolute Deviation (MAD) — the average distance each data point is from the mean |
| Dot plots | Histograms & Box Plots — more powerful ways to see shape, center, and spread |
| Describing one data set | Comparing two data sets — using center & spread side by side to draw conclusions |
Everything you learn later builds on the ideas of center, spread, and shape. Once you're comfortable with these, you'll be ready to tackle more advanced topics like standard deviation and even probability distributions in high school. For now, focus on understanding what center and spread mean and when to use each measure.
Practice Problems
Lesson Summary
Every time you collect data to answer a statistical question, your data set has a distribution—a pattern that shows how the values are arranged. You can describe that distribution using two big ideas: center and spread. The center tells you the typical or middle value. The three main measures of center are the mean (the average), the median (the middle value when data is ordered), and the mode (the most frequent value). The spread tells you how far apart the values are, and the simplest measure is the range (maximum minus minimum).
The shape of the distribution—symmetric, skewed right, or skewed left—helps you pick the best measure. For symmetric data, the mean works great. For skewed data with outliers, the median is more reliable. Understanding center, spread, and shape gives you the power to summarize any data set with just a few numbers, making it easier to see patterns, compare groups, and make decisions.