PRE-ALGEBRA • STATISTICS & PROBABILITY

Describing Distribution Shape — I can describe shape, clusters, gaps, and outliers and explain what they mean in context.

Learn to read the story that data tells through its shape, clusters, gaps, and outliers.

Why Do We Describe Data Shapes?

Imagine you collected the test scores of every student in your school. You'd have hundreds of numbers! Just staring at a long list of numbers doesn't help much. People needed a way to see patterns in data quickly. That's why, over hundreds of years, mathematicians invented graphs and learned to describe the shape of data.

1660s
Early Data Tables
John Graunt in London used tables to organize birth and death records. This was one of the first times anyone tried to find patterns in large amounts of data.
1786
First Bar Charts
William Playfair invented the bar chart and line graph. These visuals let people see data patterns instead of just reading numbers.
1800s
The Bell Curve
Carl Friedrich Gauss noticed that many measurements (like heights) form a symmetric, bell-shaped curve. This shape became one of the most important ideas in statistics.
1900s
Describing Distributions
Statisticians developed a common vocabulary — shape, center, spread, clusters, gaps, and outliers — so everyone could describe data the same way.

Today, when you look at a dot plot or histogram, you're using these same ideas. The big question is: What story does the shape of the data tell us? That's exactly what this lesson will teach you.

Core Vocabulary & Definitions

Before we can describe data, we need to learn four key words. Think of these as your toolkit for talking about any data set.

1

Shape

Shape describes the overall outline of the data when you graph it. Is it symmetric (a mirror image), skewed left (tail stretches left), or skewed right (tail stretches right)?
2

Clusters

Clusters are groups of data points that bunch together. They show values that are common or popular. For example, most kids in a class might score between 75 and 85.
3

Gaps

Gaps are spaces in the data where no values (or very few values) appear. A gap might mean something unusual happened, like nobody scored between 60 and 70 on a test.
4

Outliers

Outliers are data points that are far away from the rest. If most students ran a mile in 8–10 minutes but one student ran it in 5 minutes, that 5 is an outlier.
KEY TAKEAWAY
Think of a data distribution like a crowd of people at a concert. The shape is the overall outline of the crowd. Clusters are groups of friends standing together. Gaps are empty spaces where nobody is standing. And outliers are the people standing way off by themselves, far from everyone else.

Seeing Distribution Shapes

The best way to understand distribution shape is to look at examples. The diagram below shows three common shapes you'll see when data is graphed as a dot plot or histogram.

A symmetric distribution looks the same on both sides. A skewed-left distribution has a long tail stretching to the left. A skewed-right distribution has a long tail stretching to the right.

To decide if data is skewed, look for the tail. The direction the tail points tells you which way it is skewed. If the tail stretches to the right, we call it skewed right. If it stretches to the left, it's skewed left. A handy trick: the tail points in the direction of the skew.

💡 Quick Tip
Not sure which way the data is skewed? Ask yourself: "Where is the long, thin tail?" The tail points the same direction as the skew name.

How Shape Connects to Center and Spread

Shape isn't just about how data looks — it actually changes how we measure the center. Two measures of center you know are the mean (the average) and the median (the middle value). Here's how shape affects them.

MEAN (AVERAGE)
Mean = (Sum of all values) ÷ (Number of values)
Add up every data value, then divide by how many values there are. The mean is pulled toward outliers and tails.
MEDIAN (MIDDLE VALUE)
Median = the middle value when data is in order
Line up your data from smallest to largest. The value right in the center is the median. The median is not pulled by outliers.
How shape affects the mean and median
ShapeMean vs. MedianBest Measure of Center
SymmetricMean ≈ MedianEither one works well
Skewed RightMean > Median (pulled right by tail)Median is usually better
Skewed LeftMean < Median (pulled left by tail)Median is usually better

This is why shape matters so much. If you don't know the shape, you might pick the wrong measure of center. When data is skewed or has outliers, the median usually tells a truer story than the mean.

Spotting Clusters, Gaps, and Outliers

Now let's zoom in on three important features you should always look for in a data display. The dot plot below shows the number of books 20 students read over the summer.

This dot plot shows a cluster of students who read 2–7 books, a gap at 8 books (nobody read exactly 8), and an outlier at 15 books — one super reader far from the rest!

What Do These Features Mean in Context?

  • Cluster (2–7 books): Most students read a moderate number of books. This is the "typical" range for this group.
  • Gap (8 books): No one read exactly 8 books. This gap separates the main group from a few heavier readers.
  • Outlier (15 books): One student read far more than everyone else. Maybe they love reading, or maybe they counted audiobooks too. Outliers often have an interesting story behind them.
📝 Always Explain in Context!
Don't just say "there's an outlier at 15." Say "One student read 15 books, which is much more than the rest of the class." Using real-world words makes your description meaningful.

Worked Example: Describing a Distribution

Here's a real scenario. A teacher recorded how many minutes each of her 15 students spent on homework last night:

0, 10, 15, 20, 20, 25, 25, 25, 30, 30, 35, 35, 40, 60, 90

Describe the distribution's shape, clusters, gaps, and outliers.
1
Step 1 — Organize & Graph the DataThe data is already sorted from smallest to largest. Imagine plotting each value as a dot on a number line from 0 to 90.
2
Step 2 — Describe the ShapeMost values are between 10 and 40, with a few large values stretching out to the right (60 and 90). Since the tail stretches right, the distribution is skewed right.
Shape: skewed right
3
Step 3 — Identify ClustersThere is a cluster of data from about 15 to 40 minutes. This tells us that most students spent between 15 and 40 minutes on homework.
Cluster: 15–40 minutes (most students)
4
Step 4 — Look for GapsThere is a gap between 40 and 60 — nobody spent around 45 or 50 minutes. There is also a gap between 60 and 90.
Gaps: between 40 and 60 minutes, and between 60 and 90 minutes
5
Step 5 — Spot OutliersThe value 90 is far from all other values. It could be an outlier. The student who spent 90 minutes might have had a big project due.
Outlier: 90 minutes — much higher than the rest
6
Step 6 — Write a Full DescriptionPutting it all together: "The distribution of homework times is skewed right. Most students spent between 15 and 40 minutes, forming a cluster. There are gaps between 40 and 60 minutes and between 60 and 90 minutes. One student spent 90 minutes, which is an outlier far above the rest, possibly due to a special assignment."

Strengths and Limits of Each Feature

Each feature — shape, clusters, gaps, and outliers — tells you something different about the data. Let's compare what each one is good at revealing.

What each distribution feature tells you — and what it doesn't
FeatureWhat It RevealsPossible Limitation
ShapeOverall pattern: is data balanced or lopsided?Doesn't tell you specific values or details.
ClustersWhere most values group together — the 'popular' range.Two people may disagree on exact cluster boundaries.
GapsRanges where no data exists — may signal something unusual.Small data sets may have gaps just by chance.
OutliersExtreme values — possible errors, special cases, or interesting stories.Not every extreme value is a mistake — you need context to decide.
KEY TAKEAWAY
Think of describing a distribution like being a detective. Shape gives you the big picture (like looking at a room from the doorway). Clusters and gaps are clues you find by zooming in. Outliers are the surprising evidence that might crack the case. A good detective uses all the clues together!

Connecting to More Advanced Ideas

The vocabulary you're learning now is the same vocabulary used in high school and college statistics. As you move forward, you'll learn even more precise ways to describe data. Here's a preview.

How today's skills lead to tomorrow's concepts
What You Know NowWhat Comes Next
Describing shape as symmetric, skewed left, or skewed rightLearning about the normal distribution (bell curve) and its exact formula
Spotting outliers by eyeUsing the IQR rule (1.5 × IQR) to mathematically identify outliers
Noticing clusters in dot plotsUsing histograms with different bin widths and box plots to see clusters
Describing gaps with wordsAnalyzing bimodal distributions (two peaks) that create natural gaps

The important thing is that the skills you build now — looking at graphs, using the right words, and explaining what you see in context — are the foundation for everything that comes later in statistics.

Practice Problems

PROBLEM 1CONCEPTUAL
A histogram has most of its bars on the left side, with a long tail stretching to the right. Is this distribution symmetric, skewed left, or skewed right? Explain how you know.
PROBLEM 2BASIC CALCULATION
Here are the ages (in years) of people at a family party: 2, 5, 7, 8, 9, 10, 10, 11, 35, 38, 40, 42, 70. Identify any clusters you see and explain what they might mean.
PROBLEM 3INTERMEDIATE
A dot plot shows the number of goals scored per game by a soccer team over 12 games: 0, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3, 8. Describe the shape, any clusters, gaps, and outliers. Explain what the outlier might mean in context.
PROBLEM 4APPLIED
A store manager collected the prices (in dollars) of all 18 items on a sale rack: 3, 5, 5, 5, 7, 7, 8, 8, 8, 8, 10, 10, 10, 12, 12, 15, 15, 45. The manager says the 'typical' sale price is about $12 because that's the mean. Do you agree? Use the shape and outlier to explain your reasoning.
PROBLEM 5CRITICAL THINKING
Two classes took the same test. Class A's scores are: 70, 72, 74, 75, 76, 78, 80. Class B's scores are: 40, 55, 72, 75, 78, 92, 98. Both classes have a median of 75. Describe the shape and spread of each class's scores and explain why the median alone doesn't give you the full picture.

Lesson Summary

When you describe a distribution, always talk about four things. The shape tells you the overall pattern — is it symmetric, skewed left, or skewed right? Clusters show where data groups together, telling you the most common range. Gaps are empty spaces where no data appears. Outliers are values far away from the rest that deserve a closer look.

Always explain these features in context — use the real-world meaning of the data. Shape also affects which measure of center to use: the median is usually better for skewed data, while the mean works well for symmetric data. Think of yourself as a data detective — use shape, clusters, gaps, and outliers as your clues to tell the full story!

Varsity Tutors • Pre-Algebra • Describing Distribution Shape