AP STATISTICS • EXPLORING ONE-VARIABLE DATA

Describing the Distribution of a Quantitative Variable

Learn to characterize shape, center, spread, and unusual features of any numerical dataset.

Historical Context & Motivation

Long before the formal discipline of statistics existed, scholars and administrators needed ways to make sense of numerical data—census counts, astronomical measurements, agricultural yields. The challenge was always the same: how do you take a collection of numbers and distill them into a meaningful narrative? The practice of describing distributions evolved over centuries, driven by the practical need to summarize data for governance, science, and commerce. Today, this skill forms the foundation of exploratory data analysis, the critical first step before any inferential procedure can be applied.

1662
John Graunt's Bills of Mortality
Graunt published one of the first systematic analyses of quantitative data, summarizing London's death records by cause and identifying patterns in the distribution of mortality across age groups.
1733
De Moivre's Normal Curve
Abraham de Moivre derived the bell-shaped curve as an approximation to the binomial distribution, giving statisticians their first formal model for the shape of a distribution.
1900
Karl Pearson's Skewness Coefficient
Pearson introduced numerical measures for skewness and kurtosis, allowing researchers to quantify distributional shape rather than relying solely on visual inspection of histograms.
1977
Tukey's Exploratory Data Analysis
John Tukey's landmark book formalized EDA, introducing the boxplot and stem-and-leaf display as tools for describing distributions before applying formal inference.

The central question that these pioneers collectively addressed remains the one you will master in this lesson: given a set of quantitative observations, how do we systematically describe what is typical, how much variability exists, and what overall pattern the data follow? On the AP Statistics exam, every description of a quantitative distribution must address four features: shape, center, spread, and unusual features—a framework you can remember with the mnemonic SOCS.

Core Principles: The SOCS Framework

When the AP Statistics exam asks you to describe a distribution, the grading rubric consistently rewards responses that address four dimensions. Omitting any one of these can cost you points on free-response questions. The SOCS framework provides a systematic checklist: Shape, Outliers (unusual features), Center, and Spread. Each of these dimensions captures a different aspect of how data values are arranged along the number line, and together they provide a complete verbal portrait of the distribution.

1

Shape

Describes the overall pattern of the distribution: symmetric, skewed left, skewed right, unimodal, bimodal, or uniform. Shape determines which summary statistics are most appropriate.
2

Outliers & Unusual Features

Individual observations that fall notably far from the overall pattern, as well as gaps or clusters. Always mention specific values when identifiable from the display.
3

Center

A single value that represents the typical observation. Common measures include the mean (balance point) and median (50th percentile). Use the median when the distribution is strongly skewed or contains outliers.
4

Spread

Quantifies how much variability exists. Measures include the range, interquartile range (IQR), and standard deviation. The choice should match the measure of center: IQR with median, standard deviation with mean.
KEY TAKEAWAY
Think of describing a distribution like writing a real-estate listing for a neighborhood. Shape is the architectural style (modern, colonial, mixed). Center is the typical home price. Spread is the range from the cheapest to the most expensive house. Outliers are the one mansion or the one teardown that doesn't fit the pattern. A buyer who reads only the average price without knowing the spread or shape could make a terrible decision—and a statistician who reports only the mean without context is equally incomplete.
📝 AP Exam Tip
On free-response questions, always describe the distribution in context. Don't just say "the distribution is skewed right"—say "the distribution of household incomes is skewed to the right, indicating that most households earn below the mean while a few high earners pull the tail." Context is required for full credit.

Visual Explanation: Recognizing Distribution Shapes

The most effective way to assess the shape of a distribution is through a graphical display—most commonly a histogram, dotplot, or stemplot. The diagram below presents three canonical distribution shapes that you must be able to identify on the AP exam. Pay close attention to where the bulk of the data lies and which direction the long tail extends.

In a symmetric distribution, the mean and median coincide near the center. In a right-skewed distribution, the mean is pulled toward the long right tail and exceeds the median. In a left-skewed distribution, the mean is pulled toward the left tail and falls below the median.

When describing shape, use precise language. A distribution is approximately symmetric if the left and right halves are roughly mirror images. It is skewed right (positively skewed) if the right tail is longer, meaning a few unusually large values stretch the distribution to the right. It is skewed left (negatively skewed) if the left tail is longer. Additionally, identify the number of prominent peaks: unimodal (one peak), bimodal (two peaks), or uniform (roughly equal bar heights). A bimodal distribution often signals that two distinct subgroups have been combined.

Mathematical Framework: Measures of Center & Spread

After identifying the shape and any unusual features, you need numerical summaries to describe center and spread precisely. The choice of statistics depends on the shape of the distribution, because the mean and standard deviation are sensitive to extreme values while the median and IQR are resistant (robust) to outliers.

Measures of Center

SAMPLE MEAN
x̄ = (1/n) Σᵢ₌₁ⁿ xᵢ
where is the sample mean, n is the number of observations, and xᵢ is the i-th data value. The mean represents the balance point of the distribution and is pulled toward extreme values.
MEDIAN
Median = value at position (n + 1)/2 in the ordered data
If n is odd, the median is the single middle value. If n is even, the median is the average of the two middle values. The median is resistant to outliers.

Measures of Spread

SAMPLE STANDARD DEVIATION
sₓ = √[ Σᵢ₌₁ⁿ (xᵢ − x̄)² / (n − 1) ]
The standard deviation measures the typical distance of observations from the mean. We divide by (n − 1) rather than n to produce an unbiased estimate of the population standard deviation. Pair the standard deviation with the mean.
INTERQUARTILE RANGE (IQR)
IQR = Q₃ − Q₁
where Q₁ is the first quartile (25th percentile) and Q₃ is the third quartile (75th percentile). The IQR captures the middle 50% of the data and is resistant to outliers. Pair the IQR with the median.
⚖️ Choosing the Right Pairing
If the distribution is approximately symmetric and free of outliers, report the mean and standard deviation. If the distribution is skewed or has outliers, report the median and IQR. Mixing pairings (e.g., reporting the median with the standard deviation) is not standard practice.

Identifying Outliers and the Five-Number Summary

Outliers can dramatically affect the mean and standard deviation, so identifying them is a crucial step in any distributional description. The most common quantitative rule on the AP Statistics exam uses the 1.5 × IQR rule: any observation below Q₁ − 1.5 × IQR or above Q₃ + 1.5 × IQR is flagged as a potential outlier. This rule is closely connected to the five-number summary (minimum, Q₁, median, Q₃, maximum) and the boxplot, which visually encodes this summary.

A modified boxplot displays the five-number summary with whiskers extending to the most extreme non-outlier values. Observations beyond the 1.5 × IQR fences are plotted individually as outliers.

In the boxplot above, the box spans from Q₁ to Q₃, capturing the middle 50% of the data. The line inside the box marks the median. The whiskers extend to the smallest and largest observations that are not classified as outliers—sometimes called the adjacent values. Any data point beyond the fences is plotted as an individual dot. When describing a distribution using a boxplot on the AP exam, you can still assess the four SOCS dimensions: a longer right whisker or many high outliers suggests right skew, the median line's position within the box indicates symmetry or lack thereof, and the box width and whisker length convey spread.

Worked Example: Describing a Distribution

Consider the following dataset representing the number of hours per week that 15 college students spend studying: 2, 5, 7, 8, 10, 10, 12, 14, 15, 15, 16, 18, 20, 22, 42. A histogram of these data and the five-number summary are provided. Describe the distribution completely.

Complete SOCS Description
1
Step 1 — Organize and Compute the Five-Number SummaryThe data are already sorted. With n = 15, the median is the 8th value: 14. Q₁ is the median of the lower 7 values (positions 1–7): the 4th value = 8. Q₃ is the median of the upper 7 values (positions 9–15): the 12th value = 18.
Five-number summary: Min = 2, Q₁ = 8, Median = 14, Q₃ = 18, Max = 42
2
Step 2 — Check for Outliers Using the 1.5 × IQR RuleIQR = Q₃ − Q₁ = 18 − 8 = 10. Lower fence = 8 − 1.5 × 10 = −7. Upper fence = 18 + 1.5 × 10 = 33. The value 42 exceeds 33, so it is an outlier. No values fall below −7.
One outlier identified: 42 hours
3
Step 3 — Describe ShapeLooking at the histogram (or the sorted data), the bulk of the observations cluster between 5 and 22, with the single extreme value at 42 pulling the right tail. The distribution is unimodal and skewed to the right.
Shape: unimodal, skewed right
4
Step 4 — Describe CenterBecause the distribution is skewed right and contains an outlier, the median is a more appropriate measure of center than the mean. The median study time is 14 hours per week. For reference, the mean is x̄ = (2 + 5 + 7 + ... + 42)/15 = 214/15 ≈ 14.3, which is only slightly higher in this case, though it would be more strongly affected if additional outliers existed.
Center: median = 14 hours per week
5
Step 5 — Describe SpreadBecause we chose the median, we pair it with the IQR. The IQR is 10 hours, meaning the middle 50% of students study between 8 and 18 hours per week. The range is 42 − 2 = 40 hours, but this is heavily influenced by the outlier.
Spread: IQR = 10 hours
6
Step 6 — Write the Full Description in ContextThe distribution of weekly study hours for these 15 college students is unimodal and skewed to the right. The median study time is 14 hours per week, with an IQR of 10 hours (Q₁ = 8, Q₃ = 18). There is one outlier at 42 hours per week, which represents a student who studies considerably more than the rest. The skewness and outlier suggest that while most students study between roughly 5 and 22 hours per week, at least one student has an unusually high study load.
All four SOCS dimensions addressed in context ✓

Comparing Graphical Displays

Multiple graphical displays can represent the distribution of a quantitative variable, and each has strengths and limitations. On the AP exam, you may encounter histograms, dotplots, stemplots, and boxplots. Understanding when each is most useful ensures you choose the right tool and interpret it correctly.

Comparison of common graphical displays for quantitative distributions
DisplayStrengthsLimitations
HistogramShows shape clearly; works well for large datasets; bin widths can be adjusted to reveal different patternsIndividual values are lost; appearance depends on bin width choice; does not preserve exact data values
DotplotPreserves individual values; easy to identify clusters, gaps, and outliers; excellent for small to moderate datasetsBecomes cluttered for very large datasets; difficult to read when many values overlap
StemplotPreserves exact data values; shows shape like a histogram turned sideways; back-to-back version enables group comparisonImpractical for large datasets or data with many digits; stem choice affects readability
BoxplotCompactly summarizes five-number summary; excellent for comparing distributions side by side; clearly identifies outliersDoes not show shape details (unimodal vs. bimodal); hides clusters and gaps within the box; no individual data points visible (except outliers)
KEY TAKEAWAY
Think of each graphical display as a different lens for a camera. A histogram is a wide-angle lens that captures the overall landscape of your data. A dotplot is a macro lens that reveals fine details. A boxplot is a dashboard gauge that gives you a quick read on the key statistics. No single lens is always best—the choice depends on your purpose and the size of your dataset. On the AP exam, when comparing two groups, side-by-side boxplots are typically the most efficient choice.

Connection to Inference and the Normal Model

Describing distributions is not just an exercise in summarization—it directly informs the inferential methods you will study later in the AP Statistics course. Many inference procedures, such as the one-sample t-test, assume that the sampling distribution of the mean is approximately normal. By the Central Limit Theorem, this assumption is generally safe for large samples, but for small samples, it requires the underlying population distribution to be roughly symmetric and free of strong outliers. Thus, the exploratory step of describing shape and identifying outliers is a prerequisite for valid inference.

How distributional features connect to later inference topics
EDA ConceptInferential Connection
Symmetric, unimodal shapeSupports use of the Normal model; z-scores and percentile calculations are meaningful; t-procedures are valid even for small n
Skewed distributionLarger sample sizes needed for CLT to apply; median/IQR preferred over mean/SD; transformations (e.g., log) may normalize the data
Outliers presentCan inflate standard error and distort confidence intervals; must investigate whether outlier is a data error or genuine extreme value before proceeding
Bimodal distributionSuggests subgroup analysis may be needed; a single mean is a poor summary; consider stratifying before inference

Looking ahead, when you encounter the Normal distribution in greater depth, you will use the empirical rule (68–95–99.7) as a quick check: if approximately 68% of data fall within one standard deviation of the mean, the distribution is well-modeled by the Normal curve. You will also use Normal probability plots to assess Normality more rigorously. All of these techniques build on the foundational skill of describing distributions that you are developing now.

Practice Problems

1
A researcher collects data on the annual incomes of 200 employees at a technology company. The histogram shows a strong right skew with several extremely high values. Which pair of summary statistics is most appropriate for describing the center and spread of this distribution?
2
A dataset has Q₁ = 20, Q₃ = 36, and includes the following values: 3, 15, 20, 24, 28, 30, 36, 38, 40, 62. Using the 1.5 × IQR rule, which values are outliers?
3
The dotplot below represents test scores for 25 students (each dot = one student): the scores cluster heavily between 82 and 90 with a peak at 86, and there are two isolated dots at 55 and 58, with no scores between 58 and 82. Which description best characterizes this distribution?
PROBLEM 4APPLIED
An environmental scientist measures the dissolved oxygen (DO) levels (in mg/L) in 20 water samples from a lake. The data, in sorted order, are: 3.1, 5.4, 6.2, 6.8, 7.1, 7.3, 7.5, 7.6, 7.8, 7.9, 8.0, 8.1, 8.2, 8.3, 8.5, 8.7, 8.9, 9.1, 9.4, 9.8. (a) Calculate the five-number summary and determine whether any outliers exist using the 1.5 × IQR rule. (b) Describe the distribution of dissolved oxygen levels completely using the SOCS framework. Be sure to provide your answer in context.
PROBLEM 5CRITICAL THINKING
Two AP Statistics classes took the same exam. The side-by-side boxplots show that Class A has a median of 78 with an IQR of 12 and no outliers, while Class B has a median of 82 with an IQR of 22 and two low outliers at 35 and 42. Both distributions appear roughly symmetric (ignoring the outliers in Class B). A student claims that Class B performed better because its median is higher. (a) Write a complete comparative description of the two distributions, addressing all relevant features using proper statistical language and context. (b) Evaluate the student's claim. Explain why comparing only the medians provides an incomplete picture. (c) A teacher wants to award a bonus to students who performed unusually well relative to their class. The teacher plans to use z-scores. For Class B, explain a potential problem with using the mean and standard deviation to compute z-scores, and suggest an alternative approach.

Lesson Summary

Describing the distribution of a quantitative variable requires a systematic approach captured by the SOCS framework: Shape (symmetric, skewed left, skewed right, unimodal, bimodal, or uniform), Outliers and unusual features (identified using the 1.5 × IQR rule or visual inspection for gaps and clusters), Center (mean for symmetric distributions, median for skewed), and Spread (standard deviation paired with the mean, IQR paired with the median). Every description must be written in context, referencing the variable and its units.

Graphical displays—histograms, dotplots, stemplots, and boxplots—each reveal different aspects of a distribution. The five-number summary (min, Q₁, median, Q₃, max) underlies the boxplot and provides the foundation for outlier detection. Mastering this descriptive skill is essential because the shape and spread of a distribution determine which inferential procedures are valid later in the course.

Varsity Tutors • AP Statistics • Describing the Distribution of a Quantitative Variable